Insights · Article · August 3, 2026

Berkeley RDI Agentic AI Summit 2026.

Field notes from UC Berkeley — two days, five thousand in person, and the recurring argument that the model is no longer what separates outcomes.

UC Berkeley Campanile at golden hour with abstract agentic AI network overlay

This past weekend on the University of California, Berkeley campus, Berkeley RDI, the Center for Responsible, Decentralized Intelligence at UC Berkeley, together with the Responsible Decentralized Intelligence Foundation hosted the Agentic AI Summit 2026. RDI is co-directed by Dawn Song, whose open agentic AI course has drawn close to forty thousand learners worldwide, and the summit grew out of that community. The first edition ran in 2025 as a single day, with two thousand people in the room and forty thousand more on the livestream. This was the second, and it roughly doubled: two days, about five thousand in person, fifteen hundred industry organizations, two hundred and fifty universities, two hundred posters, and a much larger global stream.

The format shaped what anyone could absorb. Four stages ran in parallel across both days, keynotes were fifteen minutes and featured talks five to ten, so a single stage could move through eight speakers before lunch. Four stages at once means nobody saw the summit. Everyone saw a quarter of it and chose which quarter.

I attended across both days, and a fair amount of what I sat through went over my head. A Caltech professor described progress on a group theory conjecture open for sixty years and I followed maybe the first third. A talk from SK hynix on disaggregated memory pools was clearly important to the people around me and mostly opaque to me. That is the right experience at a conference like this.

What follows is the part I could use.

The gap is not the model

The most repeated claim across every stage, from people with no connection to each other, was that the model is no longer what separates outcomes.

Faraz Shafiq, who leads product and solutions at Wells Fargo, put it most directly. Their AI assistant has crossed a billion customer interactions with thirty-three million active users, and where the tools are in use they see roughly twenty-five percent higher account openings. Then: "The differentiator is actually not the model. We all have access to the same frontier models. The differentiator is the stuff around the model."

Surojit Chatterjee of Ema put a number on what that stuff is. Asked why enterprise adoption lags, he said the blocker is "not so much technical anymore. It used to be technical a few years back. It's seventy percent organizational."

The production data supports this in an uncomfortable way. Jun Yang, a senior director at NVIDIA, showed a dashboard from their own bug-fixing agent running against their own production codebase. Roughly eight hundred bugs, three hundred and thirteen agent-generated fixes, seventy-three accepted into the main branch. More than a hundred were rejected by human reviewers as incomplete or superficial. About a twenty-three percent acceptance rate, from a company with every incentive to report better.

Zelin Wan, Ph.D. at Postman showed the pattern that explains a lot of failed pilots. On single API tasks, most models score between eighty-eight and ninety-seven percent. Chain those same tasks so each depends on the last and the score falls to between forty-four and seventy-three percent. "The task itself didn't get harder. That's the difference between answering and executing."

The framing I keep returning to came from Vincent Vanhoucke at Waymo, who has been operating autonomous agents commercially for years. Reliability is a game of nines. Ninety-nine percent is fine at small scale, then you need 99.9, then 99.999. "Every nine typically requires a different solution. You don't earn a nine by improving your system a bit or getting a better model. You have to redesign the system."

Where the durable asset sits

If the model is not the differentiator, the obvious question is what is.

Yu Su, who teaches at The Ohio State University and runs NeoCognition, gave the answer I found most useful. He separated intelligence, the capacity to solve a problem given a statement and some context, from expertise, which is accumulated and situated competence in a particular job in a particular environment. Modern society, he argued, is not one world. It is millions of micro-worlds. "Every company is special. That's why they exist." Without a way to accumulate, a very capable model is "the world's smartest novice," brute-forcing every problem from scratch. "That's why everyone's token bill is exploding."

Chuan Li of Lambda gave it the most memorable form. The model weights and today's solution are the fish, and they depreciate. "Investing in how you get there, the tools, the environment and the system of record, those are the things that have long term value that compounds."

Jianfeng Gao of Microsoft Research supplied the mechanism. Agent harnesses generate trajectories, and those trajectories become training data that gets folded back into the model. His formulation: "The harness is like the training data to agentic modeling, just like the Internet data to model pre-training."

Autonomy as something earned

The governance material was better than I expected, largely because it came from practitioners rather than theorists.

Malgorzata (Gosia) Steinder, an IBM Fellow, explained in one sentence why agent security is structurally different rather than merely harder. Zero trust depends on understanding interaction patterns between applications a priori, at configuration time. Agents have no a priori patterns. There is nothing to configure against.

Credo AI framed the response as a ladder. Every intern engineer starts with limited production access and earns more. Every trader earns their limits. "Organizations have built these ladders of earned authority over years of trial and error. But for agents, that ladder largely doesn't exist yet. Many agents receive full authority the moment they're deployed. And that is a design error." Capability comes from the model, they argued, but autonomy must be earned from the enterprise.

Neil Lawrence of Trent AI gave the version I have quoted most since. "Accounting is in the numbers. Accountability is in the human authority and the judgment. Computers do accounting very, very well, but they don't do accountability well. They are not socially accountable. They can't be sent to jail. They can't be embarrassed. They can't lose their job. And our society is entirely based on that form of accountability."

One finding reframed the threat model for me. Huan Sun at Ohio State showed that severe safety failures in computer-use agents emerge with no adversarial attack at all, from slight and entirely benign rephrasings of ordinary instructions.

The contested parts

Several of the most interesting moments were arguments rather than findings, and I will report them as such.

Andrew Ng was blunt about the open-weight debate. "I was in the room when a number of executives from a number of companies were saying to government regulators frankly misleading, hyperbolic things about AI safety in order to try to drive regulatory capture." He added that he hopes both Anthropic and OpenAI succeed. Speakers on other stages made the opposite case with equal conviction.

On capital, Jasjeet Sekhon of Google DeepMind offered the most honest framing I heard. Current AI capex dwarfs Apollo, the internet buildout, and the Manhattan Project. The only larger capital expenditure in history is the railroads, "and the railroads were not a scientific bet. We knew how to make railroads." Then the concession: "The revenues don't sustain the capital expenditures we're making so far. That's the definition of a scientific bet."

Ali Ghodsi of Databricks pushed back on compressed timelines. The electric motor existed in the 1840s and took until roughly 1920 to show up in GDP. The internet existed by 2000 and Airbnb arrived in 2009, with no physical infrastructure required. "Most of the interesting things that are going to happen haven't happened yet."

What happens to the work

Andrew Ng's position was the clearest treatment I heard: "There will be no AI job apocalypse." His mechanism is what makes it more than reassurance. A coding agent frees perhaps thirty to fifty percent of a developer's time, and "that remaining work, which is a complement to the coding, therefore becomes even more valuable." Scope expands. "Which is why these days I don't hire front-end or back-end developers. Almost all of my engineers are full-stack developers." He did not present it as painless: "If thirty percent of your job goes away, how do you rise up and do that broader set of tasks and also learn to use AI?"

Shafiq described the organizational consequence in terms I had not heard before. "There are not going to be any individual contributors. Out of our bank of two hundred thousand employees, largely they're individual contributors. That means the ICs now will be managers of agents. And people who are individual contributors today are typically not used to delegating work." His conclusion: "You're not used to a world of no ICs and all managers. The playbooks really don't exist."

The question that has stayed with me longest came from a panelist at a quantitative trading firm, and nobody on the stage answered it. "Traditionally we try something, we make a mistake, we learn from that mistake, and that's how judgment is developed. Now that AI does most of the experimentation for us, how is that judgment going to be developed?"

The line I keep coming back to

Nikhil Chandhok, chief product and technology officer at Circle, closed his panel by describing himself as AGI-pilled and then saying this:

"It may be a data center full of geniuses. But the humans are Byzantine. They are idiosyncratic. They are very hard to coordinate. So it may be very lonely for those data centers full of geniuses while the rest of us try to coordinate and figure this out. I am surprised by how much effort it takes just to convince people to do things even when those things are objectively good for them."

Two days of talks about capability, and the constraint that came up most often was not capability. It was reliability, accountability, and the ordinary difficulty of getting an organization to change how it works. That is a useful conclusion, and it is where the actual work is.

Thanks to Dawn Song and the Berkeley RDI team, who ran something genuinely ambitious. Recordings for all four stages are on the RDI YouTube channel.

#AgenticAI · #ResponsibleAI · #BerkeleyRDI · #AI · #Deeptech