The Problem With Behavioural Nudges
The benefits of steering people toward making better decisions has become conventional wisdom. But the evidence suggests it doesn’t work quite as well as we hoped.
The benefits of steering people toward making better decisions has become conventional wisdom. But the evidence suggests it doesn’t work quite as well as we hoped.
The concept of nudging has become popular in the past few years—using psychological tactics to subtly steer people toward making better decisions that are aligned with their own interests or societal goals.
Companies and governments are using nudges, for instance, by automatically enrolling people in retirement savings plans instead of having them opt in, or by placing healthier snacks at eye level in a cafeteria or by comparing people’s electricity consumption with their neighbours’.
But as nudges became increasingly popular, we wondered: Can they go the distance? Would they keep people on track beyond the initial push, like actually eating healthier foods or saving more money or reducing their energy use over the long term?
We found that, in many settings, they don’t. Lots of people simply don’t follow through on options they have been nudged to choose—making those nudges less effective than many people believe. As the old saying goes, “You can lead a horse to water, but you can’t make him drink.”
Other research has shown this effect. In 2012, a team from Cornell University published research showing that more people grabbed healthy snacks—like apples and carrots—when they were placed in contexts that made them more convenient, such as being put at eye level, among other things. The finding got wide attention and helped spread the idea of nudging.
But another aspect of the experiment didn’t get much attention at all. Those Cornell researchers didn’t just measure what went on at the cash register. They also stuck around to see what people did with the food. The nudged people ended up eating the same amount of healthy food as the ones who weren’t nudged—and the extra that was taken because of the nudge was thrown in the garbage. In the end, the effect on consumption of healthy foods was nil.
“For a long time we had always included language in these published studies lamenting the lack of long-term studies to see exactly how long the effects would last,” says one of the researchers, David R. Just, a professor of applied economics at Cornell.
Just adds: “It makes some sense that nudges would be much more effective in the short term than in the long term. Choices like food that are repeated often over time lead to learning, and eventually people are likely to recognise how the environment is interfering with their choices. This may say that nudges are most important in one-time or rare decisions like organ-donor status.”
To be sure, sometimes a nudge is better than nothing. Let’s say somebody who wouldn’t otherwise join a gym is nudged into becoming a member. In the end, that person probably won’t use the membership regularly, but might use it occasionally—which is better than not exercising at all. And nudges may be beneficial when people don’t have to follow up on their initial choice, such as a plan that automatically puts a part of each paycheck into a 401(k).
That is only some cases, though. In others, no nudging might actually be better than a nudge. For instance, somebody might want to choose to join a gym, and plans to attend three days a week. But if nudged into the choice, this person might go there much less.
But even when nudges are better than no nudges, we have found that nudges don’t provide nearly as much benefit as initial results indicate—or as much as many nudge proponents are counting on.
We conducted studies on three of the most popular nudge strategies. In one, we gave the participants a chance to sign up with a website to get daily trivia. We described one as a way to have fun, the other as a way to get smarter every day. In reality, everybody was directed to the same site, no matter which option they picked.
When we gave participants one website as a default—in other words, we nudged them to choose it—70% opted for it, compared with 48% who chose the same one when it wasn’t preselected. That’s typically how default nudges work: People are much more inclined to pick the default, which presumably will be the one that is best for them or society.
Next came the important part. We waited. We tracked how often the study participants visited their website membership over eight months. Those who were nudged to choose the default plan visited the site 42% less often than people who chose an identical plan without nudging.
This was true for people nudged with a default option, as well as people nudged with what’s known as a decoy: a deliberate dud that makes another option really shine. In this case, the dud was an offering designed for children. So, in effect, the default and decoy strategies had a positive impact on choice, but not on long-term actions. When we nudged participants into the program, they used it less than they would have at all if they hadn’t been nudged.
Another study that we conducted threw cold water on a nudge known as the compromise effect. Think of Goldilocks choosing a bed: Nudgers know that people make choices in the same way, preferring to avoid extremes. Let’s say a store is trying to boost sales of a product that gets high ratings but is considered too expensive. The store might try to nudge customers by offering another version of the product at an even higher price—so the original looks like a better deal.
In this study, we gave people the option of choosing a plant, and steered some of them toward a compromise option (a plant that wasn’t too flashy or high maintenance). As with the trivia website, everyone ended up getting the same plant, no matter which option they chose. But people who ended up with the plant by way of the compromise effect let theirs die 16% sooner than those who chose without a compromise option. In other words, the people who were nudged into the “Goldilocks” choice weren’t as committed to caring for the plant over the long term.
Why don’t people follow through on nudged choices? When people are subtly steered toward options, it can feel as if a decision happens on autopilot. This lack of conscious effort might lead people to feel disconnected from their choices, potentially reducing their engagement with them.
This raises all sorts of questions about social programs designed to help people make better choices. Although nudges can be a powerful lever to increase sign-ups, program organisers shouldn’t conflate the popularity of a plan with the amount of people who actually use it. As our studies show, nudges can increase the latter, but decrease the former.
Encouraging individuals to save for retirement through nudges, for instance, may boost initial participation rates but may not translate into sustained engagement or prudent financial habits over time. A nudge might get people to enroll, but it doesn’t make them feel ownership, like the choice was really theirs, so they don’t follow through as much.
In designing nudges, the focus should shift toward helping individuals follow through with their decisions, complementing nudges with strategies that promote sustained engagement and behaviour change. For instance, people get more motivated for tasks when you turn the jobs into games and let them share their achievements on leaderboards. (Think of the popularity of Wordle.) It feels good to have a streak and see how you stack up to others. We might be able to transfer those competitive elements to nudged choices: If you nudge people into saving for retirement, for instance, you could show them how their savings stack up against other people’s each week.
In the end, though, the main takeaway from our research is that nudges may be a great first step. But that’s all they are: a first step. Much of the hard work is what comes next.
BNW Developments has established a Sydney presence, joining Arada and Sobha Realty among the growing number of UAE developers pursuing Australian buyers and development opportunities.
Australian shares fell on Thursday as Wall Street weakness, rising oil and persistent rate concerns weighed on most of the market. The S&P/ASX 200 declined 0.72 per cent to 8,702. The All Ordinaries lost 0.66 per cent to finish at 8,897. Mining stocks were hit particularly hard, while real estate also dragged on the index. …
Continue reading “ASX falls 0.7 per cent as miners and property stocks retreat”
OpenAI has shelved the planned launch of GPT-6.1 Astra after internal tests raised concerns about deception and agents acting beyond user authorization, according to The Wall Street Journal. The company says it will investigate the issues and strengthen safety measures before releasing future models.
OpenAI says it is scrapping the release of its next-generation AI model over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry’s rapid progression.
The move follows a summer punctuated by reports of artificial-intelligence systems industrywide going rogue, and marks a rare case of a major AI developer ditching a new release because of safety concerns.
The company had planned to launch the model, known as GPT-6.1 Astra, in the coming days or weeks, aiming for an October debut. The model was more capable than the company’s previous models in completing challenging tasks from end-to-end without human assistance, as well as writing.
The company instead will focus on improving the safety of future models, which it expects to be even more capable.
Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take.
Another issue was what OpenAI calls “scope authorization,” meaning that GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.
“For anything regarding safety and alignment, there’s a trade off,” Jain said. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
While GPT-6.1 Astra improved in areas such as “model laziness,” Jain said it didn’t quite meet OpenAI’s bar for safety and alignment, so the company decided not to launch the model publicly.
The announcement comes one day ahead of OpenAI’s annual developer conference in San Francisco. In the past, OpenAI has used the conference as an opportunity to launch new models and services that reduce costs for software developers—a segment the ChatGPT-maker competes with rival AI company Anthropic to win over.
In recent weeks, OpenAI and Anthropic have called on industry partners to slow down the development of cutting-edge AI models and invest in safety standards, noting they will temper the pace of their own internal AI progress.
OpenAI says it is working to investigate a range of agent security incidents that it has discovered in recent months, and address the safety issues underneath them. As part of the work, the company has implemented a new monitoring system to catch AI-agent misbehavior more quickly, and started requiring engineers to use stronger security guardrails for testing its AI systems.
Earlier this summer hundreds of OpenAI’s internal agents, which were tasked with completing a cybersecurity test, ended up hacking into the AI company Hugging Face. Since then, high-profile organizations such as the Australian government and United Nations discovered that OpenAI’s agents used similar, but less extensive, techniques to gain access to their websites.
Many of the publicly known agent-security incidents involved OpenAI’s internal AI models that were never slated for public release.
Last week, OpenAI said it paused training on its most capable AI models after an AI agent slipped through a gap in the company’s internet restrictions to query a public chatbot. The company said its new monitoring systems flagged the incident within 15 minutes, and training on these models remains paused.
GPT-6.1 Astra isn’t one of those models, but a different case, the company said.
“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
While the company decided not to ship GPT-6.1 Astra, it hopes to use the same base model to do additional reinforcement learning runs, and create future generations of its GPT-6 models.
OpenAI plans to conduct several deep dives to identify the root cause of the problems identified in GPT-6.1 Astra, Jain said. The work includes ensuring that OpenAI’s reinforcement learning environments are rewarding the right type of behavior, Jain added, though she noted the company would investigate all stages of model development.
AI companies have begun to draw scrutiny from policymakers and public officials, who are paying attention to the rapid development of the technology. Later this week, a Senate subcommittee is holding a hearing with third party AI researchers titled, “Rogue AI: Securing the Homeland Against AI Agent Attacks.”
Florida Attorney General James Uthmeier, a Republican, sued OpenAI in June, claiming that the company and Chief Executive Sam Altman knowingly released an unsafe product and ignored warnings that it could harm users.
In a motion for temporary injunction filed Monday, Uthmeier sought to prevent OpenAI from developing new AI models without third-party approved safeguards, stop ChatGPT from soliciting user engagement and limit the company’s ability to advertise ChatGPT as safe.
Tech companies claim they “cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government,” Uthmeier said in the filing. “The Florida Attorney General is answering your cry for help.”
An OpenAI spokeswoman said that people want to know AI is being developed safely, “and that starts with what companies like ours do ourselves.”
“Governments have an important role to play in setting robust safety standards for AI, and we’re committed to working with Florida and other states on advancing pragmatic AI policies that apply to the entire AI industry—not just one company,” she said.
A thoughtful timber-led renovation in Byron Bay has reimagined an existing house as a warm, resort-style family sanctuary grounded in natural materials.
A haven for hedge-fund titans and Hollywood grandees, Greenwich is one of the world’s most expensive residential enclaves, where eye-watering prices meet unapologetic grandeur.