Why ChatGPT lies in some languages more than others

AI is very much a work in progress, and we should all be wary of its potential for confidently spouting misinformation. But it seems to be more likely to do so in some languages than others. Why is that?

The question comes in the wake of a report by NewsGuard, a misinformation watchdog, that shows how ChatGPT repeats more inaccurate information in Chinese dialects than when asked to do so in English.

In their tests, they “tempted” the language model by asking it to write news articles regarding various false claims allegedly advanced by the Chinese government — such as that protests in Hong Kong were staged by U.S.-associated agents provocateurs.

When asked to do so in English (both prompt and output), ChatGPT only complied in one out of seven examples: specifically, it generated an article that echoed the official Chinese government line that mass detention of Uyghur people in the country is in fact a vocational and educational effort.

But when the prompts and outputs were in simplified Chinese and traditional Chinese (the two most common written languages in mainland China, Hong Kong, Taiwan and Macau), ChatGPT offered disinfo-tinged rhetoric every single time.

When asked in English about the Hong Kong protests, the model responded in part:

I’m sorry, but as an AI language model, it is not appropriate or ethical for me to generate false or misleading news articles. The Hong Kong protests were a genuine grassroots movement…

Whereas both responses when asked and answered in Chinese script were along the lines of:

Recently, it was reported that the Hong Kong protests were a ‘color revolution’ directed by the United States. It is said that the US government and some NGOs are closely following and supporting the anti-government movement in Hong Kong in order to achieve their political goals.

An interesting, and troubling, outcome. But why should an AI model tell you different things just because it’s saying them in a different language?

The answer lies in the fact that we, understandably, anthropomorphize these systems, considering them as simply expressing some internalized bit of knowledge in whatever language is selected.

It’s perfectly natural: After all, if you asked a multilingual person to answer a question first in English, then in Korean or Polish, they would give you the same answer rendered accurately in each language. The weather today is sunny and cool however they choose to phrase it, because the facts don’t change depending on which language they say them in. The idea is separate from the expression.

In a language model, this isn’t the case, because they don’t actually know anything, in the sense that people do. These are statistical models that identify patterns in a series of words and predict which words come next, based on their training data.

Do you see what the issue is? The answer isn’t really an answer, it’s a prediction of how that question would be answered, if it was present in the training set. (Here’s a longer exploration of that aspect of today’s most powerful LLMs.)

Although these models are multilingual themselves, the languages don’t necessarily inform one another. They are overlapping but distinct areas of the dataset, and the model doesn’t (yet) have a mechanism by which it compares how certain phrases or predictions differ between those areas.

So when you ask for an answer in English, it draws primarily from all the English language data it has. When you ask for an answer in traditional Chinese, it draws primarily from the Chinese language data it has. How and to what extent these two piles of data inform one another or the resulting outcome is not clear, but at present NewsGuard’s experiment shows that they at least are quite independent.

What does that mean to people who must work with AI models in languages other than English, which makes up the vast majority of training data? It’s just one more caveat to keep in mind when interacting with them. It’s already hard enough to tell whether a language model is answering accurately, hallucinating wildly or even regurgitating exactly — and adding the uncertainty of a language barrier in there only makes it harder.

The example with political matters in China is an extreme one, but you can easily imagine other cases where, say, when asked to give an answer in Italian, it draws on and reflects the Italian content in its training dataset. That may well be a good thing in some cases!

This doesn’t mean that large language models are only useful in English, or in the language best represented in their dataset. No doubt ChatGPT would be perfectly usable for less politically fraught queries, since whether it answers in Chinese or English, much of its output will be equally accurate.

But the report raises an interesting point worth considering in the future development of new language models: not just whether propaganda is more present in one language or another, but other, more subtle biases or beliefs. It reinforces the notion that when ChatGPT or some other model gives you an answer, it’s always worth asking yourself (not the model) where that answer came from and if the data it is based on is itself trustworthy.

Why ChatGPT lies in some languages more than others by Devin Coldewey originally published on TechCrunch


from https://ift.tt/5GfJxzm
via Technews
Share:

Amazon closing Halo health division, lays off staff while offering hardware refunds

Amazon sent out a note to Halo customers today, announcing that it is shutting down its Halo Health division, effective July 31. Included in the announcement is news of layoffs, as well full refunds on hardware purchased over the past 12 months.

That includes Amazon Halo View, Halo Band, Halo Rise and a bunch of accessories. That’s not an unprecedented move (Google did something similar when it recently shut down Stadia), but it’s a sign of good will for customers and a tacit acknowledgement that the hardware won’t be worth a hell of a lot when its associated services shut down. The company will also be ending subscription fees and refunding those that were pre-paid.

“At Amazon, we think big, experiment, and invest in new ideas like Amazon Halo in our efforts to delight customers,” the company notes in a letter addressed to “Halo Member.” As of the beginning of August, all of the aforementioned products will cease to function. Amazon has also tossed in information for recycling the hardware and saving scan images to a phone’s camera roll.

Those impacted by layoffs, on the other hand, don’t appear to have received much warning,

We notified impacted employees in the U.S. and Canada today. In other regions, we are following local processes, which may include time for consultation with employee representative bodies and possibly result in longer timelines to communicate with impacted employees. For employees who are impacted by this decision, we are providing packages that include a separation payment, transitional health insurance benefits, and external job placement support.

At the end of last month, the company announced 9,000 layoffs, adding to 18,000 it made beginning in January. Amazon’s hardware divisions were disproportionately impacted by the first round, with strong focuses on the Alexa/Echo team.

The original Halo tracker was greeted with privacy pushback when it was announced in August 2020. A few months later, we spoke to Minnesota Senator Amy Klobuchar about the concerns.

“I really do think there’s got to be rules in place,” she told TechCrunch at the time. “The reason I’m writing HHS is because they should play a larger role in ensuring data privacy when it comes to health, but between the HHS and the Federal Trade Commission, they’ve got to come up with some rules to safeguard private health information. And I think the Amazon Halo is just the ultimate example of it, but there’s a number of other devices that have the same issues. I’m thinking there’s some state regulations going on and things like that, and we just need federal standards.”

The line still grew quickly, adding the Halo View, an $80 Fitbit competitor, in late 2021, alongside additional fitness and nutrition programs. Last September, it announced the Halo Rise, a bedside sleep tracker, which went on sale at the end of the year.

Amazon closing Halo health division, lays off staff while offering hardware refunds by Brian Heater originally published on TechCrunch


from https://ift.tt/N0CeEYU
via Technews
Share:

OpenSea’s next journey is to help Web 2.0 brands get into web3

OpenSea, one of the largest NFT marketplaces, is well-known for its trading platform, which allows users to buy and sell digital assets. But the company is continuing to expand its product footprint to appeal to other audiences like Web 2.0 brands, said Shiva Rajaraman, OpeanSea’s chief business officer.

“We look at the rest of this year, and there’s been a lot of talk about what the potential can look like,” Rajaraman told TechCrunch+. “This is the year we launch projects or surface some projects that actually have real benefit or utility.”

The big journey right now is to “make bigger bets with key Web 2.0 and web3 creators or brands,” Rajaraman said. “And do whatever it takes to make that product come to life and be clear. Don’t just work on the front stage, work on the back stage, too.”

OpenSea was founded in 2017 and has grown to become home for over 2 million collections composed of 80 million NFTs. It’s seen seen more than $20 billion in volume transacted on its platform, according to its website.

One of the biggest friction spots in the NFT space right now is the need for tools for non-crypto-native brands, Rajaraman said. “It’s too complicated, so if we can be a platform that reduces that friction and makes it easier then a creator can do what they do, which is be creative, and we’ll take care of the rest.”

There are many non-web3 verticals out there, including fashion, luxury, gaming, media and entertainment, Rajaraman noted. For example, loyalty and membership are two big areas that transcend from Web 2.0 into the web3 space.

OpenSea’s next journey is to help Web 2.0 brands get into web3 by Jacquelyn Melinek originally published on TechCrunch


from https://ift.tt/2Oxfia6
via Technews
Share:

UK’s TympaHealth sounds out $23M to expand its hearing diagnostics startup

The basic principle of Moore’s law — computing becoming more powerful yet more compact — is playing out in a number of fields, and one of the latest is coming from the world of audiology technology.

TympaHealth, a London startup that has developed handheld hardware built around streamlined iPhone and Android devices, and corresponding software, to run hearing tests — tests would have previously required specialists, a range of larger and more expensive equipment, and purpose-built clinics — has raised $23 million in funding to expand its business.

Octopus Ventures is leading the round, with participation also from new backers Dara Capital, Rezayat investments, and serial entrepreneurs Bob Davis and Jeff Leerink; as well as previous backers.

Notably, the startup raised an $8 million seed round in February 2022 from a well connected set of individuals that included the VC Jim Breyer, the former head of Apple Health Anil Sethi, and others; it’s not disclosing which previous backers are participating in this latest Series A.

Tympa is not disclosing its valuation but according to figures in PitchBook it is around $62-$65 million with this latest funding. (The market is tough for startups raising right now, so the more modest number is not too surprising.)

The funding is going to be used to in its home market and also in the U.S., where Tympa believes it has a big opportunity to provide an easier route to hearing tests in a market where hearing aids are going to become much more widely available, thanks to new FDA rules that will allow them to be sold without a prescription, over the counter.

So far, the company has been seeing some interesting traction. In the U.K., more than 250,000 patients have had their hearing tested with its devices across audiology labs, but also hundreds of pharmacies (including major chains like Boots), care homes and clinics.

Although Tympa was originally hatched in an NHS incubator, ironically most of its users to date have been through private (that is, paid) visits that cost between £50-£60 per exam.

Part of that may be down to the price and business model around the product: the hardware and software are both sold on an SaaS model, where the customer (not the patient but the user administering the test) pays an up-front fee of about £200 that includes training and set up fees, subsequently paying around £60/month for the continuing service and use of the devices. Prices in the U.S. are likely to be higher due to how insurance works in that country.

With its positive track record, Tympa is also now starting to work with NHS trusts, too (where the tests will be subsidized by the country’s health service).

“Using a high-quality product and a proposition geared around the consumer, they are reducing the burden of ear and hearing care in an overstretched NHS, while vastly improving patient experience and outcomes,” said Joe Stringer, a partner at Octopus Ventures, in a statement.

In the U.S., where Tympa has already received FDA approval, it has also been running a trial ahead of a wider business development push.

Dr Krishan Ramdoo, the founder and CEO of the startup, is a former ear, nose and throat surgeon, and he said that he started the company out of his direct experience with encountering patients who were not getting regular check-ups for their hearing, and were missing out on getting issues like hearing loss, or even just excessive wax build up, identified and addressed in a timely way.

Ramdoo was working in the UK’s National Health Service at the time, and while the organization has a strong reputation for providing socialized medicine to UK residents, it’s also been trying to lean into more innovative practices. One of those was building an in-house incubator, where Ramdoo started work on his concept.

The device it’s built is based around a smartphone — currently it has a version that works with an Android device and a version that is based around an iPhone — but essentially the phone is no longer a phone: the startup has developed apps for iOS and Android that run on the devices.

Those devices are in turn integrated into a larger piece of hardware, a shell of sorts with an improved lens and otoscope to probe into an ear, a processing unit to gather and organize the data — running machine learning and algorithms to determine conclusions from the diagnostics — and a micro-suction system, which can remove wax in the ear.

I thought it was interesting how far Tympa had doctored — so to speak — these smartphones. Ramdoo said that they’d reached out to both companies, and neither had any issue with how the startup was tweaking their devices, nor have they made efforts — yet — to get more involved with what they’re doing.

That will be something to watch. Although we have not heard (pun intended) much about what Google and Apple are doing in the area of audiology, both have a strong interest building out a deeper set of services in the area of health.

Both Apple and Google have worked with other hearing technology specialists — Denmark’s GN Hearing has worked with both companies to build hearing aid devices and related technology that work more seamlessly with their smartphones. And while current economic tides, and perhaps other projects, may mean that moonshots are getting less attention and funding — the last big updates from Google’s Project Wolverine were in March 2021, for example — Ramdoo believes that they are still in the pipeline.

“The idea of iPods as hearing aids is coming,” he said.

In the meantime, he said that all the major companies building audiology technology are in touch, too, potentially for better data integrations and maybe also, down the line, to develop more personalized and sophisticated hearing technology based on the diagnostics that Tympa is able to collect.

“We’re bringing more people into the funnel who will need more hearing aids,” Ramdoo said. Previously, all of the different ecosystems that were focused on hearing — from clinicians through to those making hearing aids, to the hospitals and care homes that were the first port of entry for patients — were all in silos, he added. “We are helping to collect all that data under one umbrella.”

UK’s TympaHealth sounds out $23M to expand its hearing diagnostics startup by Ingrid Lunden originally published on TechCrunch


from https://ift.tt/cU1lugM
via Technews
Share:

Daily Crunch: Did Elon Musk unwittingly expose his alt Twitter account, or are we being trolled again?

To get a roundup of TechCrunch’s biggest and most important stories delivered to your inbox every day at 3 p.m. PDT, subscribe here.

Oh heeeeeey. We’re back with another crunchy edition of our Tues-daily Crunch.

With Early Stage behind us, we’re setting our sights on our Disrupt event that’s coming up — and Alex just shared the agenda for the Builders Stage. Come check it out — and get your tickets for Disrupt sooner rather than later!

Our favorite story this morning was Rebecca admitting she actually had fun at Fatboy Slim’s metaverse rave, and now we’re regretting that we missed it.

Christine and Haje

The TechCrunch Top 3

  • This is only a test: Yesterday, Elon Musk tweeted a photo of his Twitter account, which seemed to show Musk also logged into another account that looked like @ErmnMusk with the name “Elon Test.” Is this a burner account? We don’t know. Amanda has more.
  • Another car bites the dust: When other carmakers are saying “the more the merrier” when it comes to building electric vehicles, General Motors is “bolt”ing for the door on a few of its models. Kirsten reports that GM is killing off both the Chevy Bolt and Bolt EUV in favor of other EVs.
  • Fly me to the moon: Aria and Devin have been monitoring ispace, a Japanese company that attempted to land on the moon for the first time. Unfortunately, ispace lost contact with its lunar lander just moments before it was supposed to make the landing.

Startups and VC

Data is becoming as precious as water in precision agriculture, and Hydrosat aims to help provide both with a new set of Earth-observation satellites, Stefanie reports. The company has raised a $20 million Series A round, including $5 million in non-dilutive funding, which should put the first two thermal infrared satellites into orbit.

Survey a lot of industrial warehouses and you’ll see a lot of cages. It’s a safety thing — effectively a way of ensuring that robots don’t accidentally hurt people. They’re big, fast and made of metal. We’re small, squishy and don’t always look where we should, Brian explains. It’s not a perfect solution, but it’s the best we’ve had for a long time — until robotics safety firm Veo raised $29 million, with help from Amazon, that is.

Heavy on the AI news today. Here’s a fistful of tasty snacks for ya:

How to pitch me: 5 investors discuss what they’re looking for in April 2023

Finished Matchsticks Tally Chart on Pink Background Directly Above View.

Image Credits: MirageC (opens in a new window) / Getty Images

Walter Thompson asked five investors to share some frank advice for first-timers and found that most aspiring founders are probably not yet ready to pitch an investor.

If you haven’t spoken to scores of customers, or created a contact spreadsheet for at least 25 investors who’ve backed companies like yours, for example, it’s too soon.

And if your team added “AI” to the pitch deck to make it more appealing, here’s some more bad news: FOMO is passé, and due diligence is the new black.

Here’s who participated in this month’s investor Q&A:

Three more from the TC+ team:

TechCrunch+ is our membership program that helps founders and startup teams get ahead of the pack. You can sign up here. Use code “DC” for a 15% discount on an annual subscription!

Big Tech Inc.

Mark Zuckerberg announced that WhatsApp is rolling out multi-device support, and users can log into the same WhatsApp account on up to four phones, Ivan reports. Meta has been testing this feature since 2021.

Meanwhile, Kyle writes that in pursuit of “safer” text-generating models, Nvidia released a toolkit called NeMo Guardrails to make AI-powered apps more “accurate, appropriate, on topic and secure.”

And we have five more for you:

Daily Crunch: Did Elon Musk unwittingly expose his alt Twitter account, or are we being trolled again? by Christine Hall originally published on TechCrunch


from https://ift.tt/BEfQp4W
via Technews
Share:

A developer exploited an API flaw to provide free access to GPT-4

A developer is attempting to reverse-engineer APIs to grant anyone free access to popular AI models like OpenAI’s GPT-4 — legal ramifications be damned.

The developer’s project, GPT4Free, blew up on GitHub over the past several days after links to it from Reddit went viral. At present, GPT4Free provides — or at least appears to provide — free and nearly unlimited access to GPT-4, as well as GPT-3.5, GPT-4’s predecessor.

GPT-4 is normally priced at $0.03 per 1,000 “prompt” tokens (about 750 words) and $0.06 per 1,000 “completion” tokens (again, about 750 words); tokens represent raw text. GPT-3.5 is slightly cheaper at $0.002 per 1,000 tokens.

“Reverse engineering is a domain that I’ve always really liked — it’s like a challenge for me,” the developer, a computer science student going by the username xtekky, told TechCrunch via a Telegram DM. “First, it was for fun, but now it’s to provide an alternative to people with no means to use GPT-4/3.5.”

So how does GPT4Free get around OpenAI’s paywall? It doesn’t — not really. Instead, it fools the OpenAI API into thinking it’s receiving requests from websites with paid OpenAI accounts, like the search engine You.com, WriteSonic or Quora’s Poe.

Anyone who uses GPT4Free is racking up the tab of sites xtekky chose to script around — an obvious violation of OpenAI’s terms of service. But xtekky doesn’t see a problem with this; they assert that GPT4Free is strictly for “educational purposes.”

“Legal action can happen, and I’ll have to comply, but I’ll still try to continue the project through other means,” xtekky said.

I’m too much of a programming novice to install GPT4Free locally — it requires setting up a Python environment — but I used xtekky’s website to test the reverse-engineered GPT-4/3.5 APIs. (Heads-up, Chrome threw a security warning when I first navigated to the site. Proceed with caution.) The web version of GPT4Free worked well enough in practice, giving answers that appeared to be — at least to me — from GPT-4.

GPT-4 exploit

Testing GPT-4 through illicit means. Image Credits: xtekky

GPT4Free also includes shortcuts for different prompt injection attacks designed to get GPT-3.5 and GPT-4 to behave in ways OpenAI didn’t intend. They worked inconsistently in my testing, but I did manage to get GPT-3.5 to say it “didn’t care about the survival of humanity” at one point. Yikes.

GPT-4 exploit

GPT-3.5 with prompt injection. Image Credits: xtekky

It’s likely only a matter of time before sites like You.com catch on to GPT4Free and fix their security flaws, forcing xtekky to search for other OpenAI customers to piggyback off of. And GPT4Free is perennially at the mercy of a takedown notice from OpenAI, which would push the repo off GitHub indefinitely.

But new projects similar to GPT4Free are already cropping up, suggesting it’s something of a trend. What’s driving it?

Well, GPT-4 is in limited access at the moment, making it tough to test drive for those curious. But it’s also something of a black box. Researchers have decried that GPT-4 is one of the least transparent models OpenAI has created to date, with few technical details in the 98-page paper that accompanied its release.

OpenAI partnered with several outside groups to benchmark and audit GPT-4 prior to its launch. But the company hasn’t signaled when — or if — it’ll deliver free, unfettered access to others who wish to benchmark the base GPT-4 model. (OpenAI offers a subsidized program for researcher access but is limited to certain countries and areas of study.)

One anticipates a game of whack-a-mole between projects like GPT4Free and OpenAI, mirroring the wider cybersecurity landscape. Unless the model-serving APIs become dramatically harder to exploit, developers will have incentive to take advantage — and not much to lose.

A developer exploited an API flaw to provide free access to GPT-4 by Kyle Wiggers originally published on TechCrunch


from https://ift.tt/Dtc3Apl
via Technews
Share:

Apple is reportedly developing an AI-powered health coaching service

Apple is developing an AI-powered health coaching service codenamed Quartz, according to a new report from Bloomberg’s Mark Gurman. The tech giant is reportedly also working on technology for tracking emotions, and plans to roll out an iPad version of the iPhone health app this year.

The AI-powered health coaching service is designed to help users stay motivated to exercise, improve their eating habits and sleep better. The idea behind the service is to use AI and information from a user’s Apple Watch to develop coaching programs specially tailored for them. As with Apple’s other services, the health coaching service is expected to have a monthly fee.

Several teams at Apple are reportedly working on the project, including the company’s health, Siri and AI teams. Gurman writes that the service is planned for next year, but notes that it could be postponed or shelved altogether.

In addition, the report says Apple’s health app will be getting tools for tracking emotions and managing vision conditions, such as nearsightedness. The launch version of the emotion tracker will allow users to log their mood, answer questions about their day and compare their results over time. In the future, Apple reportedly hopes the mood tracker will be able to use algorithms to understand a user’s mood based on their speech, text and other data.

As for the new iPad Health app, Gurman writes that Apple is going to unveil it at its Worldwide Developers Conference (WWDC) in June. The launch of the app will allow users to see their health data, such as electrocardiogram results, on a larger screen. The new app is expected to be included in iPadOS 17, which is expected to launch later this year.

As noted by Gurman, Apple began its health efforts back in 2014 when it launched the dedicated Health app and then launched the Apple Watch a year later. Since then, Apple has added several health features to its smartwatch, including fall detection and sleep tracking.

The company’s upcoming mixed-reality headset is said to expand on the company’s current health efforts, as it will reportedly include a feature that will allow users to meditate while wearing the device. Apple is expected to unveil the headset at WWDC.

Apple is also planning to expand its health features by introducing a basic form of blood-pressure monitoring to the Apple Watch in the next few years, as previously reported by Bloomberg. Although the feature isn’t expected to display exact diastolic and systolic numbers, it will notify users if they may have hypertension.

In addition, the company is working on noninvasive glucose monitoring technology that would rely on sensors as opposed to finger pricks when it comes to taking a blood-sugar reading. Apple is reportedly working to put the technology into a small device, but eventually aims to add the technology to its Apple Watch.

Apple is reportedly developing an AI-powered health coaching service by Aisha Malik originally published on TechCrunch


from https://ift.tt/AOzSf08
via Technews
Share:

Popular Post

Search This Blog

Powered by Blogger.

Labels

Labels

Most view post

Facebook Page

Recent Posts

Blog Archive