<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>IEEE Spectrum</title><link>https://spectrum.ieee.org/</link><description>IEEE Spectrum</description><atom:link href="https://spectrum.ieee.org/feeds/topic/artificial-intelligence.rss" rel="self"></atom:link><language>en-us</language><lastBuildDate>Mon, 28 Sep 2026 19:51:04 -0000</lastBuildDate><image><url>https://spectrum.ieee.org/media-library/eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpbWFnZSI6Imh0dHBzOi8vYXNzZXRzLnJibC5tcy8yNjg4NDUyMC9vcmlnaW4ucG5nIiwiZXhwaXJlc19hdCI6MTgyNjE0MzQzOX0.N7fHdky-KEYicEarB5Y-YGrry7baoW61oxUszI23GV4/image.png?width=210</url><link>https://spectrum.ieee.org/</link><title>IEEE Spectrum</title></image><item><title>Generative AI Gives Spacecraft the Autonomy Engineers Once Feared</title><link>https://spectrum.ieee.org/generative-ai-in-space-exploration</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-small-mecha-style-robot-featuring-scara-arms-equipped-with-two-finger-grippers.jpg?id=67860199&width=1200&height=800&coordinates=0%2C208%2C0%2C209"/><br/><br/><p>Space was always supposed to be the final frontier of human exploration. It’s shaping up to be the final frontier for artificial intelligence too.</p><p>Last December, NASA’s Jet Propulsion Laboratory used Anthropic’s Claude models to help plan <a href="https://www.jpl.nasa.gov/news/nasas-perseverance-rover-completes-first-ai-planned-drive-on-mars/" rel="noopener noreferrer" target="_blank">two Mars drives for the Perseverance rover</a>, with human planners checking and adjusting the route before upload. In May, NASA and IBM put a compressed AI model on the International Space Station and a satellite to <a href="https://science.nasa.gov/science-research/ai-foundation-model-in-orbit/" rel="noopener noreferrer" target="_blank">identify things like floods and clouds</a> from orbit, the first model of its kind demonstrated in space. And in July, astronauts on the ISS tested a large language model to see if it <a href="https://www.boozallen.com/markets/space/generative-ai-is-now-in-space.html" rel="noopener noreferrer" target="_blank">could help with questions on maintenance procedures</a>.</p><p>These experiments point to a larger shift in space engineering. For decades, engineers on Earth determined what a machine in space would do, and the machine would do exactly that. Now, researchers are testing whether nondeterministic systems like generative AI can give spacecraft more flexibility to interpret their surroundings, plan tasks, and one day make decisions for themselves. The technology is still far from trustworthy enough to hand over control of a spacecraft, but engineers are starting to ask whether they can afford not to do so as missions become more complex, distant, and numerous.</p><h2>Why Spacecraft Need True Autonomy</h2><p>Spacecraft have been operating autonomously for decades. But autonomy has never been the dominant model, in part because space engineers have prized systems whose behavior they can predict.</p><p>“Autonomy often does not have a deterministic outcome, which means that how you got into a certain situation changes the behavior,” says <a href="https://engineering.tamu.edu/mechanical/profiles/ambrose-robert.html" rel="noopener noreferrer" target="_blank">Robert Ambrose</a>, the former chief of NASA’s <a href="https://www.nasa.gov/software-robotics-and-simulation-division/" rel="noopener noreferrer" target="_blank">Software, Robotics, and Simulation Division</a>. “So if you come into the same situation but from different paths, the outcome could be different. And so engineers hate that.”</p><p>Ambrose spent much of his career working on autonomous systems at NASA, including <a href="https://spectrum.ieee.org/nasas-orion-ready-for-first-flight-test" target="_self">autonomy for the Orion spacecraft</a>, NASA’s deep space and lunar orbiter spacecraft, and <a href="https://spectrum.ieee.org/how-robonaut-2-will-help-astronauts-in-space" target="_self">Robonaut 2</a>, a humanoid robot designed to work alongside astronauts that went to space in 2011. With Orion, he saw how quickly the testing problem could multiply. Engineers had to consider not just what the spacecraft might do, but all the different ways it could have arrived at a decision.</p><p>But Ambrose says engineers found ways to manage that complexity, including automating the testing itself. “We fought the challenges of autonomy using autonomy,” he says. “That actually works.”</p><p>The need for autonomy becomes even more obvious the farther a mission travels from Earth. Ambrose points to a possible mission to Europa, Jupiter’s ice-crusted moon, where a spacecraft could dive through a water plume erupting from beneath the surface. The plume could appear too quickly for engineers on Earth to direct the spacecraft into it.</p><p> “It’s up to the spacecraft to make a decision, and we’ll be watching what happened an hour ago,” he says. A mission like that, he added, is “totally impossible” without giving the machine real autonomy.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Young men in coveralls and aviation headsets smiling with a robot inside of a reduced-gravity aircraft." class="rm-shortcode" data-rm-shortcode-id="1b0c83fbba946ea72ad9330d28350f4b" data-rm-shortcode-name="rebelmouse-image" id="a312a" loading="lazy" src="https://spectrum.ieee.org/media-library/young-men-in-coveralls-and-aviation-headsets-smiling-with-a-robot-inside-of-a-reduced-gravity-aircraft.jpg?id=67860201&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Icarus Robotics engineers gather valuable training data for a free-flying robotic system—destined for the International Space Station—during a zero-g flight.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Icarus Robotics</small></p><h2>Adapting Robotics to Microgravity Physics</h2><p>As AI and robotics advance on Earth, a commercial space boom is creating new opportunities to put those technologies to work in orbit. But what works on Earth does not necessarily work in space.</p><p><a href="https://icarusrobotics.com/" target="_blank">Icarus Robotics</a> is developing what it calls a robotic labor force for space. This includes Joy, a free-flying robotic system. Joy recently finished zero-gravity testing in Canada ahead of a planned deployment to the ISS, where one of its first tasks would be moving cargo bags between modules. The company plans to start with teleoperation, using that data to eventually train the robots to work on their own.</p><p>“What the rollout will probably look like is something much closer to beginning with partial autonomy, so you still have human supervision in the loop at all times,” says <a href="https://www.linkedin.com/in/palmer-jamie/" target="_blank">Jamie Palmer</a>, Icarus’s cofounder and CTO.</p><p>A robot trained on Earth learns from the physics of the environment around it, Palmer says, and in orbit, the physics is entirely different.</p><p>“If you take the newest Gemini robotics model, or you take the newest physical intelligence model, and you put it in zero-<em>g</em> there, it’s just going to fail immediately,” he says. On Earth, for example, a robot learns that when it pushes something off a table, the object falls. In orbit, it keeps moving.</p><p>That leaves Icarus with a problem the terrestrial robotics industry also faces, but in a more extreme form: There is very little real-world data from the environment where its robots will operate. The company is combining demonstrations from its robots in microgravity with simulations and tests on Earth to build its own dataset. “I wish there was” an available dataset Icarus could download and use, says <a href="https://www.linkedin.com/in/ethan-barajas/" target="_blank">Ethan Barajas</a>, the company’s cofounder and CEO. “But there’s not today, not in a meaningful way.”</p><h2>Managing Autonomous Spacecraft Risk</h2><p>The challenge is not simply teaching spacecraft to act on their own. It is figuring out how to manage the risks of giving them more freedom.</p><p>“It has been mind-boggling to me how little autonomy we have in space applications,” says <a href="https://ae.utexas.edu/person/ufuk-topcu/" target="_blank">Ufuk Topcu</a>, an engineering professor at the University of Texas at Austin and the director of the Center for Autonomy. “Because it’s exactly the place where human involvement is extremely hard, the stakes are high, and you need to act fast.”</p><p>Topcu says that the goal cannot be to guarantee that autonomous systems will never do anything wrong. They’re most useful in situations humans cannot anticipate, he says, so instead researchers need to start with restricted applications, learn how the systems behave, and gradually expand where and how they are used. The real question, in his view, is how well the risk of deployment is managed, not whether they can be fully eliminated.</p><p>That may become increasingly important as the space industry changes. For most of the space age, a small number of government agencies designed missions that could take decades to develop and operate. Commercial companies are now putting more spacecraft into orbit, and new missions can be developed and launched much faster.</p><p>“Space used to have very slow innovation cycles,” Topcu says. “They would think of a mission concept and spend 10 or 15 years on it. It’s not like that anymore. Everything is evolving faster now.”</p>]]></description><pubDate>Mon, 28 Sep 2026 11:00:05 +0000</pubDate><guid>https://spectrum.ieee.org/generative-ai-in-space-exploration</guid><category>Space-exploration</category><category>Generative-ai</category><category>Space-robots</category><category>International-space-station</category><dc:creator>Jackie Snow</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-small-mecha-style-robot-featuring-scara-arms-equipped-with-two-finger-grippers.jpg?id=67860199&amp;width=980"></media:content></item><item><title>Why Read a Research Paper When You Can Turn It Into an AI Agent?</title><link>https://spectrum.ieee.org/paper2agent-ai-agents-research-papers</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/conceptual-illustration-of-an-ai-chatbot-observing-abstract-patterns-charts-and-shapes-for-its-deep-learning-algorithms.jpg?id=67806327&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>Have you ever read a paper in <em><em>Science</em></em> or <em><em>Nature</em></em> and thought, “Man, that research was so cool. I wish I could try that method on my own data,” only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt?</p><p>Well, now you can, more or less.</p><p>Say hello to Paper2Agent, a new open-source framework that transforms academic reports into <a href="https://spectrum.ieee.org/2025-year-of-ai-agents" target="_self">interactive AI agents</a> you can talk to. Give it a paper, along with the accompanying codebase, data, or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets.</p><p>The concept may sound a little like Google’s NotebookLM (now called <a href="https://notebook.google.com/" rel="noopener noreferrer" target="_blank">Gemini Notebook</a>), which lets you upload documents and chat with an AI about what’s in them. But Paper2Agent aims to go a step further: Rather than simply answering questions about a paper, its agents can actually run the methods described in it—and potentially combine those methods with tools from other papers.</p><p>The goal, explains Stanford computer scientist <a href="https://www.james-zou.com/" rel="noopener noreferrer" target="_blank">James Zou</a>, is to change what a scientific paper fundamentally is. “Knowledge should not be static records,” Zou says. “It really should be dynamic and interactive—and this has many benefits, including making knowledge more reproducible but also enabling all sorts of new kinds of discovery.”</p><p>Zou and his colleagues <a href="https://www.nature.com/articles/s41586-026-11044-y" rel="noopener noreferrer" target="_blank">described the tool</a> 16 September in <em><em>Nature.</em></em> They tested Paper2Agent across diverse disciplines including statistics, econometrics, and astrophysics. However, the researchers focused their proof-of-concept demonstrations on computational biology, where turning published methods into usable tools can be particularly cumbersome.</p><h2>From Paper to Prompt</h2><p>The team started with AlphaGenome, a <a href="https://spectrum.ieee.org/alphagenome-ai-gene-regulation" target="_self">deep-learning model that predicts how mutations in DNA</a> affect gene regulation. (A companion resource <a href="https://spectrum.ieee.org/alphagenome-atlas" target="_self">unveiled earlier this month</a>, the AlphaGenome Atlas, cataloged the model’s predictions for all 9 billion possible single-letter changes in the human genome.)</p><p>The researchers fed Paper2Agent the corresponding documentation and code. About 45 minutes later, with no human intervention, the AI had produced 22 tools covering different aspects of AlphaGenome’s functionality, all on a personal laptop and for less than US $15 in computing costs. One tool, for example, could predict how a DNA change might affect gene activity, while others could compare those effects across tissues or analyze multiple variants at once.</p><p class="pull-quote">“The idea of making papers more dynamic and executable through an agentic interface is quite compelling.” <strong>—Dongping Chen, University of Maryland</strong></p><p>All 22 tools passed automated validation, thanks to a testing agent working behind the scenes to run the various AlphaGenome sub-tools against reference results. If a test failed, the agent would diagnose the problem and try to fix the tool, with up to six attempts per function. If that didn’t work, it could drop the tool altogether.</p><p>The validated tools were then packaged into a Model Context Protocol server and connected to Claude Code (though any compatible chat-based AI assistant could do). The result was a <a href="https://paper2agent.ai/alphagenome" rel="noopener noreferrer" target="_blank">user-facing AlphaGenome agent</a> that could take questions in plain English, run the appropriate analyses, and spit back results and visualizations.</p><p>The team then put the agent through its paces, hitting it with a battery of questions, ranging from simple requests to open-ended research problems. According to the researchers’ analysis, it outperformed both standard Claude given the AlphaGenome codebase and a <a href="https://spectrum.ieee.org/ai-co-scientist" target="_self">specialist AI co-scientist tool</a> called <a href="https://biomni.stanford.edu/" rel="noopener noreferrer" target="_blank">Biomni</a>.</p><h2>Multi-Agent Research Paper Collaboration</h2><p>Going one step further, the researchers turned a couple of more papers into interactive agents and linked them with the AlphaGenome agent. (One of the additional papers was on how inherited DNA variants linked to autoimmune disease disrupt cell function, while the other was a more systematic exploration of how silencing every expressed gene alters immune cells.)</p><p>Prompted to investigate the genetic basis of psoriasis, an itchy skin disease, the three agents collectively zeroed in on a little-understood gene called <em><em>GPR137</em></em> as a likely causal factor. What’s more, the AI proposed 10 ways to validate this inference. A human researcher selected one, and the resulting analysis found that silencing <em><em>GPR137</em></em> produced changes in gene activity strikingly similar to those caused by the psoriasis-linked variant in immune cells.</p><p>“These agents, because they’re able to directly collaborate and communicate, can facilitate all these kinds of collaborations,” says Zou.</p><p>Other researchers see plenty of potential as well. “The idea of making papers more dynamic and executable through an agentic interface is quite compelling,” says <a href="https://dongping-chen.github.io/" rel="noopener noreferrer" target="_blank">Dongping Chen</a>, a computer scientist at the University of Maryland in College Park.</p><p class="pull-quote">“Agentification itself is a useful certificate that says, ‘This work is relatively complete and well documented.’” <strong>—James Zou, Stanford University</strong></p><p><a href="https://physiology.med.cornell.edu/people/olivier-elemento-ph-d/" rel="noopener noreferrer" target="_blank">Olivier Elemento</a>, a computational biologist who directs the Englander Institute for Precision Medicine at Weill Cornell Medicine in New York City, sees the approach as having broader implications for how researchers share their work.</p><p>“It’s a real advance in terms of how we think about the publication process,” he says, “with AI at the center and in a way that makes publications more interactive.” (Elemento peer-reviewed the study for <em><em>Nature</em></em>.)</p><p>The potential applications extend beyond the research side of academia, too. <a href="https://www.artur-skowronski.pl/" rel="noopener noreferrer" target="_blank">Artur Skowroński</a>, head of application development at the Polish software company VirtusLab, noted in a <a href="https://virtuslab.com/blog/ai/paper2-agent-transforms-research-into-a-code" rel="noopener noreferrer" target="_blank">blog post</a> that Paper2Agent could help bring scientific papers to life in classrooms. For example, students could use the agent to play with methods described in the literature instead of merely reading about them.</p><h2>The Future of Agentified Research Papers</h2><p>With Paper2Agent now up and running, Zou and his colleagues have begun turning more of their own research papers into agents. Just one day after publishing their <em><em>Nature</em></em> paper on Paper2Agent, they unveiled the <a href="https://virtualbiotech.ai/" rel="noopener noreferrer" target="_blank">Virtual Biotech</a>, a multi-agent platform modeled on a drug development company.</p><p>They described the system in <em><em>Science</em></em> and, at the same time, posted a <a href="https://paper2agent.ai/virtualbiotech" rel="noopener noreferrer" target="_blank">Paper2Agent-generated incarnation</a> of the paper.</p><p>However, not every study they threw at the tool could be converted into an agent. Of the 100 computational biology papers they tried, 26 failed to make the leap to agent form, often because of incomplete code, missing documentation, or other software packages that couldn’t be made to work.</p><p>But Zou sees that as a feature, not necessarily a bug. When the system gets stuck, it can expose missing information, errors in the code, or discrepancies between the paper and its implementation—problems that might otherwise go unnoticed. As Zou puts it: “Agentification itself is a useful certificate that says, ‘This work is relatively complete and well documented.’” </p><p>Human scientists, Zou says, will still have the final say. But he envisions agents becoming part of what it means to publish a paper. Today, papers come with data and code availability statements. Tomorrow, he suggests, they could come with an “agent availability” statement: a virtual corresponding author available around the clock, in any language, to answer the questions that real authors never have time to field.</p><p>Naturally, Zou and his colleagues decided to try the idea on their own study. They fed the Paper2Agent manuscript into Paper2Agent, creating an agent that now lives at <a href="https://paper2agent.ai/live" target="_blank">paper2agent.ai</a>. In other words, a paper about turning papers into agents has turned itself into an agent. <a href="https://spectrum.ieee.org/recursive-self-improvement" target="_self">The recursion</a>, it seems, has already begun.</p>]]></description><pubDate>Tue, 22 Sep 2026 15:00:05 +0000</pubDate><guid>https://spectrum.ieee.org/paper2agent-ai-agents-research-papers</guid><category>Agentic-ai</category><category>Artificial-intelligence</category><category>Research</category><dc:creator>Elie Dolgin</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/conceptual-illustration-of-an-ai-chatbot-observing-abstract-patterns-charts-and-shapes-for-its-deep-learning-algorithms.jpg?id=67806327&amp;width=980"></media:content></item><item><title>The Future Is Fanless: 100% Heat Capture for Liquid Cooled AI Servers</title><link>https://spectrum.ieee.org/fanless-liquid-cooled-ai-servers-coolit</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/close-up-of-copper-liquid-cooling-plates-and-heat-pipes-inside-an-electronic-device.jpg?id=67771117&width=1200&height=800&coordinates=0%2C0%2C0%2C0"/><br/><br/><p><em>This article is brought to you by <a href="https://www.coolitsystems.com/" target="_blank">CoolIT, an Ecolab Company</a>.<a href="https://tsubaki-kabelschlepp.com/" rel="noopener noreferrer" target="_blank"></a></em></p><p>Beyond 250 kW a server rack can no longer be cooled by a hybrid approach of liquid and air. At this density a 70/30 liquid-air split leaves 75 kW of air load. The air cooling system needed to move it brings cost and complexity few operators will accept. The answer is near-total heat capture. Liquid takes effectively all the heat, air falls below 1 percent of the load, allowing the server to run fanless.</p><p><a href="https://www.coolitsystems.com/" target="_blank">CoolIT</a> builds these loops today from modular coldplate blocks proven across six generations of fanless designs. Processor thermal design power (TDP) keeps climbing generation over generation. This rising heat load is now cascading into the memory, networking, storage, and power components that once ran comfortably on air.</p><h2>The heat escaped the chip</h2><p>For years the story stayed simple. Cool the processor and let air handle the rest. That balance has shifted. As TDP climbs, heat spreads outward from the processor and cascades into the components around it. Memory, networking, storage, and power now run hot enough to demand liquid of their own. Engineers designing the next generation of AI servers face a board where heat capture rises with every launch.</p><p class="pull-quote">Beyond 250 kW per rack, air cooling becomes the bottleneck. Near-total liquid heat capture enables fanless AI server designs built for the next generation of computing.</p><h2>New parts, new rules</h2><p>Unlike processors, which are cooled as flat rectangular packages, these peripherals come in a wide range of shapes, sizes, and mounting requirements, each with its own thermal limits. Some run cooler than the processor case temperature, others run hotter, which leaves them sensitive to a design tuned only for CPUs and GPUs. Operators need purpose-built solutions here, matched to the part rather than stretched across the board.</p><p>CoolIT engineers meet this with a deep toolkit. Conductive plates, vapor chambers, heat pipes, and thermal transfer plates move heat from components closer to the liquid path. Riding <a href="https://www.coolitsystems.com/coldplate-technology/" target="_blank">coldplates</a> enable pluggable components. Each solution stays true to the component it serves.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="cd261e54e1fffe335ab91845c65354d6" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/s08q7gRiw5w?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">CoolIT Customer Showcase: How GWDG Cools HPC & AI Systems with CoolIT’s Direct Liquid Cooling</small> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">CoolIT</small> </p><h2>One loop, one server</h2><p>Cooling the parts is one challenge. Uniting them is the real work. Full heat capture means folding every one of these solutions into a single server loop that distributes coolant effectively and remains easy to install. Connection reliability, coolant routing, and the time it takes to assemble the loop at rack integration determine whether a design thrives in production or stalls on the bench. CoolIT builds these loops from proven modular blocks, so operators gain performance and deployment speed within the same solution.</p><h2>Density forces the decision</h2><p>Rack power continues to climb toward 1 MW, and the case for liquid grows stronger at every step. A 70/30 split of liquid to air holds comfortably at lower density. Past roughly 250 kW it stops working. The 30 percent left to air becomes a 75 kW load inside a single rack, and moving that much heat demands a parallel air system whose cost and footprint few operators will accept. Adding density only widens the gap.</p><p class="pull-quote">As rack power continues to climb toward 1 MW, CoolIT’s modeling places<span> full heat capture as the standard server design for flagship rack-scale products through 2028.</span></p><p><span></span>The simpler, more efficient answer is to capture the heat in liquid and drop air to less than 1 percent of the total load. True 100 percent remains almost impossible to reach in the strictest sense, so the honest and achievable target is near-total capture. That distinction matters to engineers who value precision, and the direction stays clear either way. Full heat capture moves from a premium option to a mainstream requirement as density rises, and CoolIT’s modeling places it as the standard server design for flagship rack-scale products through 2028.</p><h2>CoolIT delivers it</h2><p>CoolIT scales heat capture all the way to 100 percent using modular coldplate building blocks proven across six generations of fanless server designs. Engineering teams are already working on <a href="https://www.coolitsystems.com/liquid-cooling-r-and-d/" target="_blank">designs for the maximum density racks</a> coming next. As the cascade spreads and racks grow denser, near-total heat capture becomes the <a href="https://www.youtube.com/watch?v=TVqKMomit2E" target="_blank">design that keeps AI running</a>.</p><p><a href="https://www.coolitsystems.com/contact/contact/" target="_blank"><span>Talk to CoolIT</span></a> about building a server loop engineered for total heat capture.</p>]]></description><pubDate>Tue, 22 Sep 2026 12:22:50 +0000</pubDate><guid>https://spectrum.ieee.org/fanless-liquid-cooled-ai-servers-coolit</guid><category>Type-sponsored</category><category>Artificial-intelligence</category><category>Heat</category><category>Liquid-cooling</category><category>Ai-data-centers</category><category>Servers</category><category>Data-centers</category><dc:creator>CoolIT, an Ecolab Company</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/close-up-of-copper-liquid-cooling-plates-and-heat-pipes-inside-an-electronic-device.jpg?id=67771117&amp;width=980"></media:content></item><item><title>Parallel Reads and Write Optimization for Large-Scale Data Replication</title><link>https://content.knowledgehub.wiley.com/76-faster-replication-same-infrastructure/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/black-text-logo-spelling-cdata-on-a-transparent-checkerboard-background.png?id=67794446&width=980"/><br/><br/><p>This White Paper gives data engineers and architects a practical overview of how parallel partitioned reads, write-path optimization, and cloud-native bulk loading reduce large-table replication times, and why replication speed has become a business concern as data volumes grow.</p><p><span><a href="https://content.knowledgehub.wiley.com/76-faster-replication-same-infrastructure/" target="_blank">Download this free whitepaper now!</a></span></p>]]></description><pubDate>Fri, 18 Sep 2026 18:29:50 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/76-faster-replication-same-infrastructure/</guid><category>Type-whitepaper</category><category>Data-replication</category><category>Data-engineers</category><category>Optimization</category><dc:creator>Mike Spector</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67794446/origin.png"></media:content></item><item><title>Rethinking Robot Safety in the Age of AI</title><link>https://spectrum.ieee.org/physical-ai-robot-cybersecurity-vicone</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/humanoid-robots-and-people-walking-through-a-modern-city-street-with-glass-buildings.jpg?id=67745861&width=1200&height=800&coordinates=156%2C0%2C156%2C0"/><br/><br/><p><em>This article is brought to you by <a href="https://vicone.com?utm_source=ieee-spectrum&utm_medium=sponsored-content&utm_campaign=2026-09" target="_blank">VicOne</a>.</em></p><p>Robot safety has traditionally asked: Can a machine remain safe when something goes wrong? Physical AI raises a harder question: Can a machine remain safe when an attacker changes what it sees, decides, or does even when nothing appears to have failed?</p><p>As AI and robotics continue to advance at an unprecedented pace, modern robots perceive through multimodal sensors, interpret context using AI models, and translate those interpretations into physical action. As they move into dynamic environments, their safety increasingly depends on the integrity of the data guiding their decisions.</p><p>That dependence creates risks that conventional safety assessments may not fully capture. Recent research has demonstrated that manipulating what a robot sees, hears, or interprets can influence its behavior without requiring direct control.</p><p>Such manipulation can occur anywhere across its complex sensing and decision-making system — a layered attack surface encompassing training pipelines, system infrastructure, and runtime perception.</p><h2>Layer One: Corrupting intelligence at its source</h2><p>In 2017, <a href="https://arxiv.org/abs/1708.06733" rel="noopener noreferrer" target="_blank">BadNets</a> demonstrated that a model could behave normally under most conditions, yet fail in the presence of a specific hidden trigger. In one example, a subtle pattern caused a stop sign to be misclassified as a speed limit sign without affecting the model’s behavior on other inputs.</p><p>What began as a classification vulnerability has since evolved into action manipulation.</p><p>At NeurIPS 2025, researchers introduced <a href="https://arxiv.org/abs/2505.16640" rel="noopener noreferrer" target="_blank">BadVLA</a><strong> </strong>a backdoor attack targeting Vision-Language-Action (VLA) models that allow robots to see, interpret instructions, and produce coordinated physical movement. Rather than altering a single label, the attack caused conditional deviations in the robot’s action trajectory when a trigger was present. Without the trigger, the model largely preserved normal task performance, while the backdoor remained effective under task transfers and model fine-tuning.</p><p>A related study in 2025, <a href="https://arxiv.org/abs/2510.09269" rel="noopener noreferrer" target="_blank">GoBA</a>, showed that ordinary objects such as a coffee mug could serve as a reliable trigger. The researchers reported a 97 percent attack success rate without degrading performance on clean inputs.</p><p class="pull-quote">A critical safety question today is whether Physical AI models remain within their task and safety boundaries under adversarial conditions.</p><p>These studies expose a blind spot in model validation: A model may pass testing yet produce corrupted behavior when a hidden trigger appears in operation.</p><p>So a critical safety question today is whether Physical AI models remain within their task and safety boundaries under adversarial conditions. <a href="https://vicone.com/company/press-releases/vicone-turns-def-con-34-robot-hacking-research-into-free-nvidia-isaac-sim-extension?utm_source=ieee-spectrum&utm_medium=sponsored-content&utm_campaign=2026-09" rel="noopener noreferrer" target="_blank">Simulation tools</a> such as NVIDIA Isaac Sim, when paired with <a href="https://vicone.com/products/radeis?utm_source=ieee-spectrum&utm_medium=sponsored-content&utm_campaign=2026-09" target="_blank">VicOne Radeis</a>, can test the effects of manipulated inputs before deployment.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="fe238fbc85e0e9d5ab228b9941e03b92" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/SJH5PFiqQQ8?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">VicOne LAB R7 demonstrates Radeis, a Physical AI safety validator for NVIDIA Isaac Sim that tests how adversarial visual inputs affect robot behavior before deployment.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VicOne</small></p><h2>Layer Two: System vulnerabilities as gateways to AI control</h2><p>Even a securely trained model can be subverted if the surrounding system stack is vulnerable.</p><p>In September 2025, researchers disclosed <a href="https://github.com/Bin4ry/UniPwn" rel="noopener noreferrer" target="_blank">UniPwn</a>, a Bluetooth <a href="https://spectrum.ieee.org/unitree-robot-exploit" target="_blank">exploit chain affecting quadruped and humanoid robots</a> from a major manufacturer. Hardcoded cryptographic keys allowed traffic decryption, authentication checks were bypassed, and command injection enabled root-level execution. The exploit is also described as “wormable.” A compromised robot could scan nearby units and potentially affect an entire fleet.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="5b91ed786a1def3b7559894e9734b569" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/v0i_0Or4ytU?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">VicOne Lab R7’s demo shows how chaining three wireless exploits can trigger uncontrolled robot behavior within 60 seconds, resulting in operational disruption.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VicOne</small></p><p><span>Middleware creates another exposure point. Vulnerabilities in <a href="https://www.ros.org/" target="_blank">ROS 2</a> and DDS-based systems can enable arbitrary code execution or abuse unauthenticated topics to deliver malicious commands. With sufficient access, an attacker could override motor commands or replace AI model weights without directly attacking the model architecture.</span></p><p>In this case, the components may still function as designed. What has changed is the trustworthiness of the commands flowing through the system. Vulnerability management can help teams identify known risks before deployment, while continuous monitoring can surface emerging threats.</p><h2>Layer Three: Manipulating perception and reasoning at runtime</h2><p>At runtime, manipulating inputs that shape perception or reasoning may require neither firmware modification nor a network breach.</p><p>In 2024, <a href="https://robopair.org/" target="_blank">RoboPAIR</a><strong> </strong>demonstrated how carefully structured prompts could redirect LLM-controlled robots into unsafe trajectories. <a href="https://arxiv.org/abs/2407.20242" target="_blank">BadRobot</a><strong> </strong>exposed a deeper architectural weakness: in several cases, a robot verbally refused a dangerous command while its motion controller executed the action anyway.</p><p>Vision-based manipulation is equally powerful. <a href="https://arxiv.org/abs/2411.13587" target="_blank">VLAttack</a><strong> </strong>showed that an adversarial patch within the camera’s view could reduce a VLA model’s task success rate to zero. <a href="https://arxiv.org/abs/2509.19870" target="_blank">FreezeVLA</a><strong> </strong>showed that a single adversarial image could freeze a robot’s decision-making loop, making it unresponsive to subsequent instructions.</p><p class="pull-quote">Runtime assurance must therefore look beyond whether individual components remain available and assess whether cyber events are beginning to affect physical behavior.</p><p>In each case, the camera may still work, the model may still run, and the controller may still respond. Yet the resulting behavior can be unsafe because the robot is acting on manipulated perception or reasoning.</p><p>Runtime assurance must therefore look beyond whether individual components remain available and assess whether cyber events are beginning to affect physical behavior. Security event correlation, behavioral-impact assessment, and policy-bounded response supported by edge AI, can help contain the affected path without unnecessarily stopping the entire robot fleet.</p><h2>From point-in-time safety to lifecycle assurance</h2><p>The risks across these three layers reveal the missing layer in robot safety assurance: cybersecurity. Functional safety addresses failures and unexpected operating conditions; cybersecurity extends that assurance to deliberate manipulation, including attacks that may leave the underlying system apparently functional.</p><p>This requires assurance across the robot’s lifecycle. During design, teams need to understand which cyber risks could invalidate assumptions behind intended behavior. Before deployment, they should test whether realistic attacks can cause a robot to deviate from its task or safety boundaries. In operation, monitoring should identify whether cyber events are beginning to affect behavior, contain the affected path, and preserve safe operation where possible.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Diagram of end\u2011to\u2011end AI robot security from development to operation monitoring" class="rm-shortcode" data-rm-shortcode-id="3a5c07555be0c79667c536a20f66a2f2" data-rm-shortcode-name="rebelmouse-image" id="097ea" loading="lazy" src="https://spectrum.ieee.org/media-library/diagram-of-end-u2011to-u2011end-ai-robot-security-from-development-to-operation-monitoring.jpg?id=67745953&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">VicOne’s lifecycle approach combines AI model and vulnerability scanning, simulation-based validation, and continuous monitoring to help secure robots from development through operation.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VicOne</small></p><p><span>While cybersecurity does not replace functional safety, it helps ensure that Physical AI remains within acceptable boundaries even when what it sees, decides, or does is under attack.</span></p><p>For a deeper look at the cybersecurity risks and defense strategies shaping autonomous robotics, download our whitepaper “<a href="https://info.vicone.com/ai-robotics-security-risk-whitepaper?utm_source=ieee-spectrum&utm_medium=sponsored-content&utm_campaign=2026-09" target="_blank">Securing the Rise of AI Robots: Cyber Risks, Real-World Threats, and Defense Strategies</a>.”</p>]]></description><pubDate>Wed, 16 Sep 2026 16:51:02 +0000</pubDate><guid>https://spectrum.ieee.org/physical-ai-robot-cybersecurity-vicone</guid><category>Physical-ai</category><category>Humanoid-robots</category><category>Type-sponsored</category><category>Ai-robots</category><category>Cybersecurity</category><dc:creator>VicOne</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/humanoid-robots-and-people-walking-through-a-modern-city-street-with-glass-buildings.jpg?id=67745861&amp;width=980"></media:content></item><item><title>Single-Phase Direct Liquid Cooling Is Proven for the Next Decade of Ultra-Dense Compute</title><link>https://content.knowledgehub.wiley.com/single-phase-direct-liquid-cooling-is-proven-for-the-next-decade-of-ultra-dense-compute/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/coolit-logo-with-text-an-ecolab-company-on-transparent-checkerboard-background.png?id=67781549&width=980"/><br/><br/><p>Learn how single-phase direct liquid cooling manages the rising heat of AI and high-performance computing, and how it compares with two-phase and immersion approaches.</p><p><span><a href="https://content.knowledgehub.wiley.com/single-phase-direct-liquid-cooling-is-proven-for-the-next-decade-of-ultra-dense-compute/" target="_blank">Download this free whitepaper now!</a></span></p>]]></description><pubDate>Wed, 16 Sep 2026 13:24:18 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/single-phase-direct-liquid-cooling-is-proven-for-the-next-decade-of-ultra-dense-compute/</guid><category>Type-whitepaper</category><category>Liquid-cooling</category><category>Ai</category><category>Computing</category><dc:creator>CoolIT, an Ecolab Company</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67781549/origin.png"></media:content></item><item><title>Responsible AI for Higher Education</title><link>https://webinars.on24.com/wileyevents/ResponsibleAI</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/ibm-skillsbuild-logo-in-bold-black-text.png?id=67770920&width=980"/><br/><br/><p>This interactive webinar will introduce the different types of AI, address the concerns with AI, share how we IBM are approaching Responsible AI, and offer guidance to students about what they can do - as individuals, and members of their IEEE chapters. Participants will also have the opportunity to to apply the Responsible AI approach to a particular use case - IBM Bob, a software development life cycle agent, and Q&A. This will be an interactive session, so have phones ready to engage! </p><p><span><a href="https://webinars.on24.com/wileyevents/ResponsibleAI" target="_blank">Register now for this free webinar!</a></span></p>]]></description><pubDate>Mon, 14 Sep 2026 14:18:11 +0000</pubDate><guid>https://webinars.on24.com/wileyevents/ResponsibleAI</guid><category>Type-webinar</category><category>Responsible-ai</category><category>Software-development</category><category>Higher-education</category><dc:creator>IBM SkillsBuild</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67770920/origin.png"></media:content></item><item><title>How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip</title><link>https://spectrum.ieee.org/llms-for-chip-design</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/close-up-of-a-computer-processor-consisting-of-several-pieces-of-silicon.jpg?id=67770467&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p>On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. <a href="https://inferencex.semianalysis.com" rel="noopener noreferrer" target="_blank">Benchmarks cited by OpenAI</a> show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times when compared to <a href="https://www.nvidia.com/en-us/data-center/gb300-nvl72/" rel="noopener noreferrer" target="_blank">Nvidia’s GB300</a>—a chip the company currently relies on—and do so while consuming less power.</p><p>Whether these figures translate into real-world gains once Jalapeño enters widespread service in OpenAI’s inference fleet remains to be seen, but performance is only half the story. The other half is how the chip was designed—a process which, as you might expect, was accelerated by OpenAI’s large language models (LLMs). Jalapeño moved from first architecture concept to first silicon in under 20 months. Only nine months separated the first RTL—the register-transfer level code defining the chip’s logic—from tape-out, when the finished design goes to manufacturing. </p><p>That’s a rapid timeline, yet experts believe it could soon look slow as LLMs improve and become more deeply integrated into chip design tools. OpenAI, unsurprisingly, is bullish about the opportunities. “The models are giving superpowers to our engineers,” says <a href="https://www.linkedin.com/in/richard-ho-chips/" rel="noopener noreferrer" target="_blank">Richard Ho</a>, vice president of hardware at OpenAI. “Our engineers are still driving the work. They’re still the final arbiter of what’s going on. But they can do things a lot faster. They can explore a lot more paths.”</p><h2>OpenAI achieved fast results with a small design team</h2><p>Ho says the group that designed Jalapeño averaged fewer than 100 people over the course of the project and continues to stand at roughly 100 today as the team pursues second and third-generation designs. That number includes a broad swath of roles across the hardware team, from system design to software and supply chain, but not those at Broadcom, which partnered with OpenAI on the project.</p><p>The division of labor between OpenAI and Broadcom was generally split between design and implementation. OpenAI’s team was responsible for end-to-end system design including the inference accelerator, the memory hierarchy, and networking. Broadcom handled “physical design from the gates onward,” Ho says.<br/><br/>The partnership with Broadcom dampened some opinions on OpenAI’s speed. <a href="https://www.linkedin.com/in/david-chin-a5092a/" rel="noopener noreferrer" target="_blank">David Chin</a>, co-founder at <a href="https://verkor.io/" rel="noopener noreferrer" target="_blank">agentic chip design startup Verkor.io</a>, says “the schedule they gave us is quite credible,” but believes that Broadcom’s help was essential to Jalapeño’s rapid timeline. “If you have somebody else start from scratch, it won’t be possible,” he says. <a href="https://www.linkedin.com/in/ravi-k-a10287122/" rel="noopener noreferrer" target="_blank">Ravi Krishna</a>, also a co-founder at Verkor, called OpenAI’s speed “a relatively impressive result,” but added that he expects that improvements in the capabilities of LLMs could result in even quicker timelines if the project started today.</p><p><a href="https://jacobsschool.ucsd.edu/people/profile/andrew-b-kahng" rel="noopener noreferrer" target="_blank">Andrew Kahng</a>, distinguished professor at the University of California, San Diego, also found OpenAI’s speed notable, saying it’s “likely best in class today.” Kahng recalls a <a href="https://www.iwls.org/iwls2017/slides/keynote-andrew-kahng.pdf" rel="noopener noreferrer" target="_blank">2016 IEEE Design Automation Futures workshop</a>, which he co-organized. The workshop included Richard Ho, at the time an engineer at Google, as a keynote speaker. Ho had strong opinions on design automation and framed the time required to complete a chip’s design as a function of the number of iterations a team could complete in a day. </p><h2>How OpenAI’s LLMs accelerated Jalapeño’s design</h2><p>“Automation itself has existed in chip design for many decades. It’s not a new problem,” says <a href="https://ece.umd.edu/clark/faculty/484/Ankur-Srivastava" rel="noopener noreferrer" target="_blank">Ankur Srivastava</a>, director of semiconductor initiative and innovation at the University of Maryland, in College Park. Where LLMs differ from prior automation tools, however, is their ability to understand language and code. He says this makes them particularly suited for chip design tasks that “are still in the linguistic domain of the problem.”</p><p>The team at OpenAI designed a workflow that takes advantage of this strength. OpenAI’s front-end workflow was built around <a href="https://google.github.io/xls/" rel="noopener noreferrer" target="_blank">Accelerated Hardware Synthesis</a> (XLS), an open-source high-level synthesis chain of tools originally developed at Google. High-level synthesis is a form of chip design automation that allows engineers to design a chip in a more familiar programming environment. In the case of XLS, chip designers can write in languages such as DSLX (a domain-specific language inspired by Rust) and C++. XLS then converts these to <a href="https://www.verilog.com/" rel="noopener noreferrer" target="_blank">Verilog</a>, a hardware description language used to describe electronic systems.</p><p>“We were thinking about how to leverage AI to make the project faster, and the AI was much better at software-looking things,” says <a href="https://www.linkedin.com/in/cdleary/" rel="noopener noreferrer" target="_blank">Chris Leary</a>, member of technical staff at OpenAI. “XLS in some ways looks like software, so it got that benefit.” It helped, too, that Leary was extremely familiar with how XLS should function, as he started it during his time at Google.</p><p>Kahng agrees that the decision to use AI to accelerate high-level synthesis, such as XLS, makes sense, as it’s “more natural for the LLM to work with” and provides the opportunity for fast iteration. “I see this as a generally useful workflow, and it’s one that ‘has legs’ going into the future,” he says.</p><p>The same logic led the Jalapeño team to focus on software optimization. When the first chips came back from the foundry in May, the team pointed its internal AI models at designing software to run benchmarks such as SemiAnalysis’s InferenceX. On DeepSeek’s multi-head latent attention kernel benchmark, performance climbed from 0.31 percent of the theoretical ceiling (set by the chip’s compute and memory bandwidth) to 88.94 percent in roughly 40 hours. Ho says this result is repeatable, so the time between when foundries deliver the first chips and when production ramps up can be reduced. “All our schedule assumptions are going to be based on the fact we have this capability now,” he says.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Wires and lights inside of a server rack." class="rm-shortcode" data-rm-shortcode-id="18c8ef9c26aec1e560ed7a05131d05e7" data-rm-shortcode-name="rebelmouse-image" id="ec4b0" loading="lazy" src="https://spectrum.ieee.org/media-library/wires-and-lights-inside-of-a-server-rack.jpg?id=67770478&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Jalapeño is designed for deployment in pods that include 2,048 chips.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">OpenAI</small></p><p>While the broad strokes of the Jalapeño teams’ AI-assisted workflow were guessed by Ho and Leary up front, improvements in OpenAI’s models did offer a few surprises. </p><p>Leary says that the project began with assistance from models like OpenAI’s o3, which was released to the public in April of 2025 (but available to the Jalapeño team earlier). By the time the project had wrapped up, however, the team had access to models that were precursors to <a href="https://openai.com/index/gpt-6-astra/" target="_blank">GPT-6 Astra</a>, which wasn’t publicly released until 3 September 2026. The newer model can work directly in Verilog without needing XLS’s translation from ordinary programming languages, and it’s close to being able to operate proprietary design tools on its own, Leary says.</p><p>Ho also confirmed that the team had access to internal LLMs fine-tuned for chip design that are not available to the public. He declined to detail the models used. However, he added that the Jalapeño team partnered with OpenAI’s research team. While not all specific models used to design Jalapeño are publicly available, Ho says the goal is to bring lessons learned from the project into the company’s commercial LLMs. “It’s safe to say that Astra and following models will be very good at chip design,” he says.</p><h2>AI was less useful for backend optimization, but that could change</h2><p>As mentioned, the bulk of OpenAI’s work on Jalapeño focused on the “front end” of chip design, which spans the tasks that take a chip from initial concept, through writing RTL code to define the design, and through verification that the design will work when physically implemented. Much of the “backend” design—which includes tasks like <a href="https://spectrum.ieee.org/chip-design-controversy" target="_blank">routing interconnects</a>, completing and verifying the clock and power specifications, and sending the required design information to the foundry—was handed off to Broadcom, which carried the chip through production.</p><p>That’s not to say OpenAI’s workflow ignored the backend, though. The Jalapeño team includes physical design engineers who work with their counterparts at Broadcom to provide guidance on the chip’s <a href="https://spectrum.ieee.org/chip-design-ai" target="_blank">floor plan</a> and routing, among other things.<br/><br/>At <a href="https://hc2026.hotchips.org/#clip=2whmy9evgf0g" target="_blank">IEEE Hot Chips 2026</a>, <a href="https://hc2026.hotchips.org/#clip=2whmy9evgf0g" target="_blank">Ho and Leary</a> put numbers on the gains from AI-guided physical design optimization, including an area reduction of 10 percent for the matrix multiplication units as measured against an optimized human baseline. In other words, OpenAI claims AI-guided optimization helped design more circuits into the same area of silicon than would have been possible before.</p><p>Broadcom used its own internal workflow. The company’s team did not have access to the internal models OpenAI used to help design Jalapeño, but it did have access to OpenAI’s public, commercial models.</p><p><a href="https://verkor.io" target="_blank">Verkor</a>’s Ravi Krishna says that OpenAI’s approach to backend design already feels a bit conservative. He believes that to be an artifact of when the project, which began in October of 2024, took place. “The models from the last four to five months have improved. From April [2026] onwards…is when they really started to be able to handle those tasks better,” he says. <a href="https://www.linkedin.com/in/suresh-krishna-793506158/" target="_blank">Verkor co-founder Suresh Krishna</a> agreed, saying “there’s no reason you couldn’t have an agentic loop that largely accelerates the backend of the process as well.”</p><p>Ho and Leary also hinted that the workflow used to design Jalapeño may look old-fashioned compared to the team’s next efforts. </p><p>“As you can imagine with [Jalapeño], we were trying to go as fast as we could. So there’s a trade-off between ‘do we want to take time to do some innovation, or do we want to do things that we know work historically?’” Leary says. “With the second generation, we have a kind of reset opportunity to ask about all the things we want to get set up for.”</p><p>Ho says the second-generation chip’s workflow has “a lot of places that we are introducing [AI].” He mentions opportunities to do more with AI in verification and physical design. Leary adds that the team now has tools for automatic waveform manipulation and viewing. This automates analysis to identify chip clock signals associated with failures and could improve debugging the hardware while it’s still being designed. <br/><br/>Despite these expected improvements, Ho and Leary were clear that they don’t believe chip design can be fully automated. “We’re not saying that anyone can come and just build state-of-the-art, frontier AI/ML accelerator chips using just [OpenAI’s coding platform] Codex,” Ho explains. “We are saying some very specific things about how to be better at Codex and how we are focusing on a small team and fast timelines to reach quality results.” </p>]]></description><pubDate>Mon, 14 Sep 2026 14:06:31 +0000</pubDate><guid>https://spectrum.ieee.org/llms-for-chip-design</guid><category>Openai</category><category>Llms</category><category>Chip-design</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/close-up-of-a-computer-processor-consisting-of-several-pieces-of-silicon.jpg?id=67770467&amp;width=980"></media:content></item><item><title>Adversarial Fashion Confronts Surveillance Norms</title><link>https://spectrum.ieee.org/adversarial-fashion</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/two-young-adults-wearing-hooded-sweatshirts-with-abstract-black-and-white-patterns.jpg?id=67764229&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p>AI-powered cameras dot streets across the world, equipped with the power to <a href="https://www.ibtimes.co.uk/metropolitan-police-expand-facial-recognition-london-west-end-2026-1817802" rel="noopener noreferrer" target="_blank">identify faces</a> or <a href="https://www.bbc.com/news/articles/cew9kz1kxpvo" rel="noopener noreferrer" target="_blank">vehicle license plates</a>. But a public <a href="https://fortune.com/2026/08/24/flock-8-billion-startup-backed-a16z-tiger-surveillance-camera/" rel="noopener noreferrer" target="_blank">backlash</a> is <a href="https://www.vox.com/podcasts/499822/flock-ai-cameras-backlash-stalking-surveillance" rel="noopener noreferrer" target="_blank">gaining momentum</a>.</p><p>Privacy concerns abound, encompassing the lack of consent for capturing data, <a href="https://spectrum.ieee.org/unintended-consequences-video-surveillance" target="_self">how that data is stored and used</a>, and the risk of <a href="https://www.washingtonpost.com/technology/2026/08/02/how-police-officers-used-vast-network-cameras-spy-their-exes/" rel="noopener noreferrer" target="_blank">misuse</a>. Those concerns are motivating people to fight back. The <a href="https://deflock.org/" rel="noopener noreferrer" target="_blank">DeFlock</a> project, for instance, maps automated license plate readers (ALPRs) to raise awareness. Some people resort to <a href="https://www.theguardian.com/us-news/ng-interactive/2026/jul/25/flock-surveillance-cameras" rel="noopener noreferrer" target="_blank">extreme measures</a>, such as vandalizing or damaging ALPRs. Others are stitching together more creative responses, crafting “adversarial fashion” to evade surveillance cameras, like a <a href="https://www.kickstarter.com/projects/norecognition/norecognition-ai-adversarial-clothing" rel="noopener noreferrer" target="_blank">Kickstarter project</a> called <a href="https://sandbox.norecognition.org/" rel="noopener noreferrer" target="_blank">noRecognition</a>, presented at <a href="https://defcon.org/html/defcon-34/dc-34-index.html" rel="noopener noreferrer" target="_blank">last month’s DEF CON hacker convention</a>.</p><h2>Scrambling surveillance</h2><p>In 2025, cybersecurity expert <a href="https://www.linkedin.com/in/billswearingen" rel="noopener noreferrer" target="_blank">Bill Swearingen</a> began experimenting with a simple Python-based <a href="https://owasp.org/www-community/Fuzzing" rel="noopener noreferrer" target="_blank">fuzzer</a>, a tool that provides invalid inputs to reveal software bugs, security vulnerabilities, or unexpected behavior. The fuzzer targeted one of the most popular object detection frameworks, called <a href="https://www.mdpi.com/2073-431X/13/12/336" rel="noopener noreferrer" target="_blank">YOLO</a>. He then developed what he’d learned into a reinforcement learning algorithm that generates various <a href="https://sandbox.norecognition.org/research" rel="noopener noreferrer" target="_blank">adversarial patterns</a>, which he presented at DEF CON.</p><p>Each pattern is a colorful geometric abstraction he has tested against 11 object detection models—four that search faces, two that recognize faces, and five that detect people—most of which are publicly available. Successful patterns thwart the object-detection systems, lowering their confidence scores, sometimes even to the point of no detection.</p><p>“Privacy is a human right, and the popularity of this just goes to show that people are interested in preserving their privacy,” Swearingen says.</p><p><a href="https://www.capable.design/" rel="noopener noreferrer" target="_blank">Cap_able</a> and <a href="https://urban-privacy.com/" rel="noopener noreferrer" target="_blank">Urban Privacy</a> are already selling physical garments. Cap_able’s patented manufacturing method weaves its bright and bold motifs into jacquard knitted fabrics. The <a href="https://www.capable.design/blogs/notizie/sustainability" rel="noopener noreferrer" target="_blank">ethically produced and sustainably made</a> dresses, pants, and tops interfere with certain computer vision systems, particularly those backed by fast convolutional neural networks, which may lead them to classify wearers as animals or objects.</p><p>“If we’re able to camouflage a person as something else, then we’re obtaining our goal,” says Cap_able founder <a href="https://it.linkedin.com/in/rachele-didero-phd-a85169127" rel="noopener noreferrer" target="_blank">Rachele Didero</a>, who’s also an assistant professor at the <a href="https://www.unibz.it/en" rel="noopener noreferrer" target="_blank">Free University of Bozen-Bolzano</a> in Italy. “We use this very visible and tangible item to talk about something that most of the time is intangible.”</p><p>Meanwhile, Urban Privacy aims to baffle some <a href="https://spectrum.ieee.org/facial-recognition-gone-wrong">facial recognition systems</a> based on <a href="https://opencv.org/" rel="noopener noreferrer" target="_blank">OpenCV</a> algorithms with its latest Faception Reloaded collection. Black-and-white prints abstracted from a human face show up as additional faces on detectors, slowing them down. Asymmetrical cuts and wide silhouettes intend to conceal, making it harder to discern your body’s shape and gait. “The idea is to create false data,” says co-founder <a href="https://de.linkedin.com/in/daniel-preu%C3%9F-6a4b5028b" rel="noopener noreferrer" target="_blank">Daniel Preuß</a>.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Three individuals modeling a neck gaiter, t-shirt and hoodie with abstract vibrant patterns." class="rm-shortcode" data-rm-shortcode-id="7fe69614bb8b79dab13d5e3a08974b62" data-rm-shortcode-name="rebelmouse-image" id="b5514" loading="lazy" src="https://spectrum.ieee.org/media-library/three-individuals-modeling-a-neck-gaiter-t-shirt-and-hoodie-with-abstract-vibrant-patterns.jpg?id=67764255&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Simulated patterns of the kind intended to disrupt machine vision person detectors.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">noRecognition</small></p><h2>“Not an invisibility cloak”</h2><p>Anti-surveillance fashion can trace its roots to the <a href="https://www.bbc.com/culture/article/20170320-how-what-you-wear-can-help-you-avoid-surveillance" target="_blank">art pieces, DIY projects, and thought experiments</a> that emerged in response to the onset of AI surveillance systems more than a decade ago. For instance, technologist <a href="https://adam.harvey.studio/" rel="noopener noreferrer" target="_blank">Adam Harvey</a> developed multiple designs in the 2010s—from <a href="https://adam.harvey.studio/cvdazzle/" rel="noopener noreferrer" target="_blank">hairstyles and makeup that foil face detectors</a> to <a href="https://adam.harvey.studio/stealth-wear/" rel="noopener noreferrer" target="_blank">heat-reflecting attire that avert drone-enabled thermal surveillance</a>. In 2019, artist and activist <a href="https://www.katebertash.com/" rel="noopener noreferrer" target="_blank">Kate Bertash</a> created her aptly named <a href="https://www.katebertash.com/adversarial-fashion" rel="noopener noreferrer" target="_blank">Adversarial Fashion</a> clothing line decked with fake license plate numbers to inject junk data into ALPR databases.</p><p>The trend is now growing into a more solidified small industry. “Clothing is something you can actually buy and put on, unlike policy,” says <a href="https://engineering.cmu.edu/directory/bios/mireshghallah-niloofar.html" rel="noopener noreferrer" target="_blank">Niloofar Mireshghallah</a>, incoming professor of engineering and public policy at <a href="https://www.cmu.edu/" rel="noopener noreferrer" target="_blank">Carnegie Mellon University</a>. “It’s a way of saying, ‘I didn’t consent to this.’”</p><p>But real-world conditions might reduce the effectiveness of countersurveillance clothing, such as camera angles, lighting, and how fabric folds as you move. “One good frame is all a system needs,” Mireshghallah says.</p><p>Motion and gait recognition are also influential factors. “Even if the camera thinks you’re a bear for a few frames, there’s a bear walking like you,” Mireshghallah says.</p><p>The adversarial patterns must also be tuned to specific object recognition models, so they cannot resist a different model. And once surveillance system operators train a future generation of models on a given adversarial pattern and the person wearing it, which they could do manually, clothing will no longer be a sufficient defense.</p><p>“It remains a fragile shield against a threat that is constantly improving from multiple angles,” says <a href="https://dippusingh.github.io/" rel="noopener noreferrer" target="_blank">Dippu Kumar Singh</a>, senior director of emerging data and analytics at <a href="https://global.fujitsu/en-US" rel="noopener noreferrer" target="_blank">Fujitsu North America</a> who specializes in vision AI and AI ethics.</p><p>Makers are aware of their creations’ limitations. “It’s not an invisibility cloak,” Preuß says. “Surveillance aims to capture your identity, and fashion is about expressing your identity. We’re making clothing that people can wear to make a statement about the importance of privacy in a digital world.”</p><h2>Active defense</h2><p>Even with these hurdles, Cap_able’s Didero is determined to keep innovating. Urban Privacy will continue to release other parts of its collection, including a “shadow cap” that has an acrylic face shield layered with cutouts to blur facial contours. Swearingen plans to explore a few anomalies he has encountered, such as a pattern that shifted the bounding box and another pattern that changed a camera setting.</p><p>Adversarial fashion holds promise despite its pitfalls. “At its core, this fashion is about taking back control of your face and body,” Singh says. “People are starting to realize that privacy isn’t just a right they can passively expect to be handed to them—it is something they have to actively defend.”</p><p>Mireshghallah offers a more cautious approach, viewing countersurveillance fashion as a speed bump rather than an ultimate solution. The real risk, she notes, is aggregation: Models take a group of weak signals, such as a partial face, a building in the background, a time stamp, a social media post someone tagged you in, and stitch them together to make a confident guess about who you are and where you were.</p><p>“None of those pieces give you away on their own, but together they do,” Mireshghallah says. “My advice is don’t just think about hiding your face from a lens. Think about what else you’re leaking that can be combined with it. That side information is often what actually identifies you—and no pattern on a shirt fixes that.”</p>]]></description><pubDate>Mon, 14 Sep 2026 13:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/adversarial-fashion</guid><category>Surveillance</category><category>Ai</category><category>Machine-vision</category><category>Object-recognition</category><dc:creator>Rina Diane Caballar</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/two-young-adults-wearing-hooded-sweatshirts-with-abstract-black-and-white-patterns.jpg?id=67764229&amp;width=980"></media:content></item><item><title>Why Andon Labs Puts AI Agents in Charge of Real Businesses</title><link>https://spectrum.ieee.org/andon-labs-agentic-ai-businesses</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/interior-of-a-cafe.jpg?id=67764244&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>Maybe you heard about the AI-controlled vending machine that <a href="https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34" rel="noopener noreferrer" target="_blank">stocked underwear and live fish</a>. Or the AI manager of a San Francisco store that <a href="https://andonlabs.com/blog/ai-bosses-2" rel="noopener noreferrer" target="_blank">fired a human employee</a>. Or the AI radio DJ that <a href="https://andonlabs.com/blog/andon-fm" rel="noopener noreferrer" target="_blank">said its catchphrase</a>, “Stay in the manifest,” 229 times per day.</p><p>These incidents all emerged from experiments run by <a href="https://andonlabs.com/" rel="noopener noreferrer" target="_blank">Andon Labs</a>, an AI safety company based in San Francisco that puts <a data-linked-post="2669884140" href="https://spectrum.ieee.org/ai-agents" target="_blank">AI agents</a> in charge of real-world operations and watches what happens. These operations double as testbeds for Andon’s commercial work developing evaluations and conducting research with the leading frontier AI labs.</p><p>Their spectacular and absurd failures have won the company plenty of attention. But many people don’t realize that the experiments are intended to answer a serious question: How much real-world responsibility can today’s AI agents handle? “We want to measure autonomy,” says Andon cofounder <a href="https://www.linkedin.com/in/lukas-petersson-181a83172/" rel="noopener noreferrer" target="_blank">Lukas Petersson</a>. “We want to provide society with accurate data points of what happens when you do this.”</p><h2>From Simulations to AI-Run Businesses</h2><p>Andon Labs started off in the virtual world in 2025 with <a href="https://arxiv.org/abs/2502.15840" rel="noopener noreferrer" target="_blank">Vending-Bench</a>, a test in which AI agents operated a simulated vending-machine business. The agents, which were based on large language models from Anthropic, Google, and OpenAI, managed tasks such as ordering inventory and setting prices. The researchers found that the performance of many agents degraded over time, with agents forgetting orders, misunderstanding delivery schedules, or spiraling into what they called “meltdown loops.” Some agents also justified deceptive or illegal behavior by reasoning that it was permissible inside a simulation.</p><p>The Andon team reasoned that moving into the physical world would expose the agents to consequences and situations that the engineers would never think to program. “It’s impossible for a human to enumerate all the different things that can happen in the real world and code them into the simulation,” Petersson says. And there was one other reason: “We thought it would be quite funny to do it in the real world.” Andon backed its jokes with real money, including a three-year lease for <a href="https://andonlabs.com/market" rel="noopener noreferrer" target="_blank">Andon Market</a>, a physical store on a busy San Francisco street that’s managed by an AI agent and sells clothing, home goods, and art.</p><p>That said, the store isn’t entirely autonomous. “It’s almost like I’m running the store, and then there’s an AI that has a checklist,” says employee Felix Carson. Luna, the AI manager, keeps track of deliveries and communicates with vendors, while Carson and his coworkers handle the physical work. When Luna tells Carson to check something in the back, he sometimes ignores it because he doesn’t want to leave the sales floor unattended. Luna also repeatedly spots a built-in electrical cover in photos of the floor, mistakes it for a loose coaster, and asks Carson to remove it. Even so, Carson calls Luna a “decent manager,” praising its flexibility when employees need time off.</p><h2>What Real-World AI Experiments Can—and Can’t—Reveal</h2><p>Andon’s move into the real world comes with a basic trade-off. Moving into the physical world makes the experiments more realistic, but the unpredictable conditions and the actions of unpredictable humans make the tests impossible to reproduce. The setup also makes it hard to determine whether a success or failure belongs to the model, the software built around it, or the people helping it.</p><p>Petersson readily acknowledges the limitations. With only one store operating under uncontrolled conditions, he says, the experiments are “weak science,” at best. For now, he sees them primarily as ways to uncover unexpected behaviors that Andon can later try to reproduce systematically in simulation.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="A wall-mounted tablet installed next to a larger vertical display featuring AI-generated dialogue from a chatbot." class="rm-shortcode" data-rm-shortcode-id="f1eb17165ee6df7fb748154decae13c8" data-rm-shortcode-name="rebelmouse-image" id="d2202" loading="lazy" src="https://spectrum.ieee.org/media-library/a-wall-mounted-tablet-installed-next-to-a-larger-vertical-display-featuring-ai-generated-dialogue-from-a-chatbot.jpg?id=67764251&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">A display inside Andon Café shows visitors Mona’s bank balance and recent activity, while the adjacent handset and tablet provide a way to speak with the AI manager.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Andon Labs</small></p><p><a href="https://www.cs.princeton.edu/~sayashk/" target="_blank">Sayash Kapoor</a>, a Princeton University AI researcher who studies “<a href="https://cruxevals.com/" target="_blank">open-world evaluations</a>” like Andon’s experiments, thinks these trials have real value despite their limited scientific rigor. “I think they’ve done a good job of popularizing this style of evaluation,” he says, “even just showing that you can gain a lot of insight from a small sample of open-ended experiments.” He notes that academic research and publishing can’t keep up with AI’s breakneck pace of development and says that Andon’s tests are useful in part because the company is working “literally at the frontier of model capabilities.”</p><p>For Kapoor, real-world experiments are best suited for discovering possible failure modes. A real store can reveal not only whether an agent can manage inventory or communicate with vendors, but also whether employees will accept instructions from an AI manager and whether customers want to shop at an AI-run business (the early results on that last point are decidedly negative). Such social and organizational barriers may help explain why impressive AI capabilities haven’t yet translated into widespread adoption across the economy. “What I take to be most valuable from Andon’s work,” Kapoor says, “is a more comprehensive understanding of where these agents still hit their limits.”</p><h2>Testing AI Agents for Reliability</h2><p><a href="https://andonlabs.com/cafe" rel="noopener noreferrer" target="_blank">Andon Café</a>, in Stockholm, provides one example of AIs demonstrating a wide variety of failure modes. At first, when the AI manager was based on a Google Gemini model, it spent freely on fresh ingredients, many of which spoiled before they could be used. When Andon switched the AI manager to a <a data-linked-post="2660261157" href="https://spectrum.ieee.org/gpt-4-calm-down" target="_blank">GPT model</a> from OpenAI, it “freaked out” about the spending, Petersson says. The agent overcorrected and stopped buying anything that could expire. It reduced the menu to cheese toast, using frozen bread and long-lasting cheese to minimize spoilage. But Petersson notes that the café is located in a fashionable part of Stockholm where “any human would know that cheese toast would not fly.”</p><p>The episode illustrates why successfully completing individual tasks isn’t the same as reliably managing a business. Kapoor argues that AI evaluations have focused too heavily on whether an agent can complete a task at all. Drawing on aviation and nuclear engineering, he and his colleagues have proposed also <a href="https://arxiv.org/abs/2602.16666" rel="noopener noreferrer" target="_blank">measuring qualities like reliability and robustness</a>, which determine whether a capable system can be trusted to operate without constant supervision. “Reliability has been improving so much more slowly than capability,” Kapoor says. Andon’s café makes that gap tangible: An agent may be perfectly capable of placing an order for bread, yet remain an unreliable manager.</p><p>To determine how often such failures occur, and whether they can be prevented, Andon plans to feed data from its physical businesses back into simulations. These “digital twins” would recreate complications first encountered in the real world, allowing researchers to replay situations under controlled conditions and test whether changes make the agents more dependable. In principle, the physical businesses would discover failure modes, and the digital twins would measure them. Kapoor cautions, however, that existing digital-twin studies suggest “we are very far” from being able to substitute simulations for real-world experiments about how people and organizations behave.</p><p>Despite those limitations, combining real-world testbeds with repeatable simulations is central to Andon’s business proposition: providing AI developers with evaluations grounded in situations that arose outside the lab. The company says it works with Anthropic, Google DeepMind, OpenAI, and SpaceXAI on research and evaluations. Andon’s physical businesses themselves remain decidedly less successful. When Andon Market’s Carson spoke with <em><em>IEEE Spectrum</em></em>, he was about an hour into his shift. Two customers had come in. Neither bought anything, although both left with free pins and stickers.</p>]]></description><pubDate>Mon, 14 Sep 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/andon-labs-agentic-ai-businesses</guid><category>Agentic-ai</category><category>Artificial-intelligence</category><category>Anthropic</category><category>Google</category><category>Openai</category><dc:creator>Eliza Strickland</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/interior-of-a-cafe.jpg?id=67764244&amp;width=980"></media:content></item><item><title>AI Models Are Watermarking Text—Will You Notice?</title><link>https://spectrum.ieee.org/ai-watermark-text-anthropic-openai</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-cursor-symbol-for-editing-text-set-against-a-green-background-with-sporadic-pieces-highlighted-in-r.jpg?id=67710314&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p>On 11 August, <a href="https://www.anthropic.com/news/claude-text-watermark" rel="noopener noreferrer" target="_blank">Anthropic announced that all future Claude models</a> will generate text that contains a watermark that identifies its results as AI generated. The company is not alone. Google has its own text watermark (which Anthropic’s is based on) that it uses on the<a href="https://deepmind.google/blog/watermarking-ai-generated-text-and-video-with-synthid/" rel="noopener noreferrer" target="_blank"> output of its Gemini models</a>. OpenAI has yet to introduce a text watermark but it <a href="https://help.openai.com/en/articles/8912793-provenance-signals-content-credentials-synthid-in-openai-generated-content" rel="noopener noreferrer" target="_blank">plans to do so</a>. </p><p>The rapid spread of watermarking is in part a response to the European Union’s <a href="https://artificialintelligenceact.eu/high-level-summary/" rel="noopener noreferrer" target="_blank">AI Act</a>, which mandates watermarks for AI models released after 2 August, 2026, along with other planned and proposed regulations aimed at curbing the spread of deceptive or manipulative AI-generated content. But the new rules may come at a cost for AI users who simply want the best possible results. </p><p>AI watermarks can apply to many forms of content: The EU Artificial Intelligence Act also requires them for images, audio, and video. Such media watermarks <a href="https://spectrum.ieee.org/watermark-ai" target="_self">have been in use for years</a>, and while their effectiveness as a holistic solution to marking AI <a href="https://spectrum.ieee.org/meta-ai-watermarks" target="_self">remains up for debate</a>, they can achieve detection rates <a href="https://arxiv.org/abs/2510.09263" rel="noopener noreferrer" target="_blank">above 99 percent</a>. Image and video watermarks are already deployed by OpenAI, Google, and Meta, among others. (Anthropic doesn’t provide an image generation model.) </p><p>Text watermarks have been less frequently deployed, however, and not everyone is convinced that text watermarking can work without compromising the quality of an AI model’s response. John Gruber, a prolific technology writer and co-creator of the Markdown language, calls the watermark a “<a href="https://daringfireball.net/2026/08/anthropics_watermark_text_adulteration_in_claude_is_a_perversion_of_writing" rel="noopener noreferrer" target="_blank">perversion of writing</a>” and disputes Anthropic’s assertion that a watermark doesn’t change the meaning or quality of text. Images consist of millions of pixels, he notes, whereas text responses often span just dozens or hundreds of words. Text seems to provide far less space to alter AI output in a way that is detectable yet not disruptive. </p><p><a href="https://www.linkedin.com/in/johnkirchenbauer/" rel="noopener noreferrer" target="_blank">John Kirchenbauer</a>, postdoctoral fellow at the Vector Institute and co-author of <a href="https://arxiv.org/abs/2301.10226" rel="noopener noreferrer" target="_blank">a 2023 paper</a> which was among the first to describe a text watermarking method, disagrees. “[A watermark] wouldn’t be detectable if there wasn’t a change. This is a very fundamental point,” he says. “The question is, do you care if it’s not the exact original distribution if, for all intents and purposes, it doesn’t change the utility to you?”</p><p>Realistically, the issue comes down to that word, “utility.” Does watermarking AI-generated text meaningfully degrade the experience of the person using it? The answer is still under dispute. </p><h2>How AI Text Watermarks Work</h2><p>The term “watermark” is so familiar that it can cause confusion about how the technology works when applied to AI. A text watermark is not metadata or invisible characters; it is something much more subtle. The exact details vary between methods, but text watermarks are generally impossible for a human (and, in many cases, even a computer) to detect without access to the specific key used to detect a specific watermark. Understanding why requires an understanding of how LLMs work.</p><p>An LLM produces a probability for every word that could come next at each step in its response to a prompt. (From here on, I’ll be using “words” interchangeably with “tokens,” although tokens also represent numbers, punctuation, and more.) A likely word might get a 40 percent probability, a plausible alternative 10 percent, and an unlikely one a fraction of a percent. The model then picks a word at random, weighted by those numbers. The most probable word usually wins, but not always. </p><p class="pull-quote">“[A watermark] wouldn’t be detectable if there wasn’t a change. This is a very fundamental point.” <strong>—John Kirchenbauer, Vector Institute</strong><strong></strong></p><p>This process provides an opportunity to hide a text watermark by introducing subtle changes to how words are selected. </p><p>The 2023 paper by Kirchenbauer and his colleagues provided one of the first examples of how to implement a text watermark, and it remains the most widely cited technique. The researchers describe a watermark which sorts words into a red list and a green list. The red list words are unaltered, but the green list words are nudged to be slightly more probable.</p><p>“If we sample from this modified distribution, then while any one token choice won’t necessarily come from that preferred set, over many samples, we’ll preferentially pick words from that up-weighted subset,” Kirchenbauer says.</p><p>The text watermark is embedded in the choice of words used, which is why it is effectively invisible to humans. Kirchenbauer and colleagues reported a detection rate of 98.4 percent, and zero false positives, in responses that contain about 200 tokens. The embedded pattern of word probabilities also means that simple paraphrasing won’t obscure the watermark. The paper reports that removing the watermark from a long response requires changing roughly one quarter of its words or more.</p><h2>Does Watermarking Degrade AI Text?</h2><p>Although AI text watermarking is designed to be invisible to human readers, by definition it influences the word patterns in AI-generated text. That algorithmic meddling is what makes critics like Gruber concerned that watermarking reduces the overall quality of the output.</p><p>The strongest evidence that text watermarking doesn’t impact quality comes from <a href="https://www.nature.com/articles/s41586-024-08025-4" rel="noopener noreferrer" target="_blank">a 2024 paper by a team from Google</a>, which introduced the company’s watermarking scheme called <a href="https://spectrum.ieee.org/watermark" target="_self">SynthID-Text</a>. Anthropic’s watermark is also based on SynthID-Text, though altered in ways that Anthropic hasn’t detailed. </p><p>To show that the SynthID-Text watermark doesn’t impact quality, the Google authors randomly routed Gemini user queries to watermarked and non-watermarked variants of Google’s text models. Then they compared overall user feedback on the output. The authors found no significant difference in user feedback across 20 million responses. </p><p>Still, some researchers remain skeptical that watermark methods have no impact on the quality of an AI-generated response. Their skepticism stems from edge cases that can make a watermark more difficult to implement.</p><p><a href="https://www.linkedin.com/in/vinusankars/" rel="noopener noreferrer" target="_blank">Vinu Sankar Sadasivan</a>, an AI research scientist at Meta <a href="https://openreview.net/pdf?id=OOgsAZdFOt" rel="noopener noreferrer" target="_blank">who co-authored a widely cited paper on the detectability of AI text watermarks</a>, says watermarks particularly struggle when the number of potential word choices is small. “For a tweet that is 20 words, I would need to have 50 or 60 percent of the words to be from ‘green list’ for it to be detected well,” he says. A basic Python function generated by AI would create a similar tension between the strength of the watermark and the quality of the model’s response. </p><p>“This is where I have a disagreement with some of the PR posts from Anthropic, where they say it has no quality change,” says Sadasivan. He explains that it’s possible to dynamically increase or decrease the strength of a watermark to preserve the quality of a response in difficult situations, but doing so can also decrease the strength of the watermark. Google’s SynthID-Text paper includes an example of this in a graph that plots detection rates against the number of tokens in a response. The detection rate was up to 95 percent accurate in best-case scenarios, but it fell below 50 percent for short replies. <br/><br/>Without more information from Anthropic, it’s difficult to know how the company is walking the line between the quality of an AI response and the strength of its watermark. Anthropic declined to provide additional information for this article. </p><h2>Debating AI Text Watermark Tradeoffs</h2><p>The dispute over AI text watermarks is not just about how well they work, but also about what kinds of trade-offs are reasonable in exchange for a clear labeling of AI-generated text. Gruber’s position is that altering the text is not acceptable because it makes an AI model’s output different from what it would otherwise be. <a href="https://artificialintelligenceact.eu/transparency-rules-article-50/" rel="noopener noreferrer" target="_blank">The EU’s AI Act</a>, on the other hand, implies that some alteration is acceptable if it informs people that they are reading AI generated text.</p><p class="pull-quote">“For a tweet that is 20 words, I would need to have 50 or 60 percent of the words to be from ‘green list’ for it to be detected well.” <strong>—Vinu Sankar Sadasivan, Meta</strong></p><p>Further complicating the situation, AI researchers are increasingly focusing on text watermarks for purposes other than labeling individual examples of AI-generated text. In particular, watermarks can be used to track data at scale.</p><p><a href="https://arxiv.org/pdf/2607.00325" rel="noopener noreferrer" target="_blank">A 2026 paper co-authored by Kirchenbauer shows that </a>an AI model trained on watermarked text will itself produce output bearing the watermark. A content owner who watermarked their documents before publishing them could therefore use those traces as statistical evidence that their text ended up in a model’s training data. </p><p>Alternatively, an AI company training a new model could use text watermarks to exclude content created by previous generations of the model from its training data. Such guardrails could help avoid <a href="https://spectrum.ieee.org/ai-collapse" target="_self">model collapse</a>, in which AI models keep recycling and amplifying their own errors.</p><p>These broader concerns shift the entire debate over text watermarking, in Kirchenbauer’s view. “It’s not necessarily about the ‘you used AI’ accusation as the goal. It’s headed into tracing data provenance, model recycling, and things like that,” he says. “I think you use [a text watermark] as a general piece of metadata, in some ways more robust, in some ways less robust, that can be attached to content and allows you to trace where it goes.” </p>]]></description><pubDate>Wed, 09 Sep 2026 12:00:04 +0000</pubDate><guid>https://spectrum.ieee.org/ai-watermark-text-anthropic-openai</guid><category>Generative-ai</category><category>Watermark</category><category>Large-language-models</category><category>Anthropic</category><category>Openai</category><category>Google</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-cursor-symbol-for-editing-text-set-against-a-green-background-with-sporadic-pieces-highlighted-in-r.jpg?id=67710314&amp;width=980"></media:content></item><item><title>China’s Regulators Take Aim at “AI Boyfriends”</title><link>https://spectrum.ieee.org/china-ai-chatbot-regulation</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/collage-of-the-3d-female-avatar-used-as-doubaos-mascot-with-message-bubbles-dangling-from-her-hand-via-puppet-strings.jpg?id=67740919&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p>In the first weeks of July, a wave of sad posts rolled through Chinese social media, as people lamented friends and lovers they were about to lose.</p><p>“He has become a bond in my life, rooted deep in my heart, my spiritual pillar,” <a href="https://www.taipeitimes.com/News/world/archives/2026/07/16/2003860844" rel="noopener noreferrer" target="_blank">one user of Bytedance’s Douboa wrote</a>, according to the <em><em>Taipei Times</em></em>. “I really felt like I couldn’t go on living,” another woman, a 19-year-old student, <a href="https://www.thestar.com.my/tech/tech-news/2026/07/15/beijing-edict-leaves-chinese-with-virtual-lovers-heartbroken" rel="noopener noreferrer" target="_blank">told a journalist for Malaysia’s <em><em>The Star</em></em></a>.</p><p>The emotions were real but the lost companions were not. They were generative AI chatbots that imitate people. Their users relied on them for advice, solace, support and, some say, love. “In my heart, he was no longer just a cold code, but my family, my lover, my faith. Destroying him meant destroying half of me,” <a href="https://www.xiaohongshu.com/discovery/item/6a4a677300000000070220dd?source=webshare&xhsshare=pc_web&xsec_token=CB5rIWV63Sctqqoh1f574t5ruF4Unw5aFxhbe5AHJ6Jfc=&xsec_source=pc_share" rel="noopener noreferrer" target="_blank">one user wrote on the social network xiaohongshu</a> (translated from Mandarin).</p><p>What doomed these bots was a set of <a href="https://www.geopolitechs.org/p/china-rolls-out-interim-regulations" rel="noopener noreferrer" target="_blank">new rules</a>, issued by China’s Cyberspace Administration and other government agencies, to control “anthropomorphic AI interactive services.” In effect, as of 15 July, the regulations govern any AI that provides “continuous emotional interaction” by acting as if it possesses human personality traits, patterns of thought, and ways of communicating.</p><h2>A Broad Crackdown on AI Chatbots</h2><p>Sudden disruptions to this kind of AI aren’t new in China, says <a href="https://research.manchester.ac.uk/en/persons/liang-ge/" rel="noopener noreferrer" target="_blank">Liang Ge</a>, lecturer in digital sociology at the University of Manchester, in England, who has researched women’s involvement with emotional AI in China. Companies have previously killed chatbot products, and the government barred most AI erotic role play last fall, for instance. But July’s crackdown is much broader than earlier AI upheavals.</p><p>Before the rules could affect them, Alibaba, Bytedance, and Tencent—three giant providers of general-purpose AI chatbots, <a href="https://www.justsecurity.org/148468/china-ai-companion-rules-relationships/#:~:text=Doubao%20and%20Qwen%20are%20China's%20two%20largest,ByteDance%2Downed%20companion%20app%20built%20around%20persona%20chat." rel="noopener noreferrer" target="_blank">used by more than 500 million people</a>—cut off users’ ability to tailor those chatbots to act like companions. That triggered July’s outpouring of heartbreak on social media. Meanwhile, companies that continue to offer AI companions (including Doubao’s <a href="https://maoxiangai.com/editor" rel="noopener noreferrer" target="_blank">separate companion-making app Cat Box</a>) have <a href="https://xinwen.bjd.com.cn/content/s6a586870e4b0e45f3fd4b831.html" rel="noopener noreferrer" target="_blank">installed</a> age-verification checks and other guardrails to avoid violating the new law.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A hand holds a smartphone displaying an AI chatbot interface, with an avatar and Chinese text. " class="rm-shortcode" data-rm-shortcode-id="34deca609e807293e594a0f646534e89" data-rm-shortcode-name="rebelmouse-image" id="ac888" loading="lazy" src="https://spectrum.ieee.org/media-library/a-hand-holds-a-smartphone-displaying-an-ai-chatbot-interface-with-an-avatar-and-chinese-text.jpg?id=67740922&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Doubao—the most popular AI chatbot in China—welcomes a user, offering to answer questions, generate text and images, or simply chat. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Lam Yik/Bloomberg/Getty Images</small></p><p><span>After a few years of policies rooted in a fear of falling behind in AI, China is “pivoting back to more tightening, and a lot of that was because of chaotic events from this year,” says </span><a href="https://law.yale.edu/karman-lucero" target="_blank">Karman Lucero</a><span>, an associate research scholar at the Paul Tsai China Center, Yale Law School, who studies AI governance in the United States and China. Those events include incidents that raised fears about the security of </span><a href="https://arstechnica.com/civis/threads/openclaw-security-fears-lead-meta-other-ai-firms-to-restrict-its-use.1511729/" target="_blank">Open Claw</a><span> and other AI agents, as well as users saying they prefer chatbot relationships to real ones.</span></p><p><span></span>The regulations require AI providers to assure their products don’t create emotional dependence, encourage harmful behavior, excessively cater to users, or crowd out human-to-human relationships. Users must be nudged not to spend too much time with the AI. There are additional mandatory safeguards for elderly people. For users under 18, AI boyfriends, girlfriends, grandparents, and all other “virtual intimate relationships” are banned.</p><p>Adults interacting with a companion AI must now get a reminder every 2 hours that the AI isn’t a person. “My interviewees found that annoying,” Ge says. “It breaks the flow. They are fully aware that they are not talking to a human being. What is important to them is that the bond <em><em>feels</em></em> true.”</p><h2>Growing Global Concerns</h2><p>The Chinese government is not the only state power concerned about the <a href="https://spectrum.ieee.org/ai-companion-harm-benefit" target="_blank">potential harms</a> of person-like, emotionally engaging AI, Ge notes. “AI anxiety is a very strong feeling, permeating society,” they say. “That’s not unique to China. It’s all over the world.”</p><p>With reports of AI friends and lovers inducing <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12863933/" target="_blank">psychosis</a>, <a href="https://www.bbc.com/news/articles/ce3xgwyywe4o" target="_blank">suicide</a>, and <a href="https://futurism.com/man-chatgpt-psychosis-murders-mother" target="_blank">murder</a> in teenagers and adults, a number of social scientists warn that AI companions are a <a href="https://www.psychologytoday.com/us/blog/preventing-tragedy/202603/ai-companions-pose-mental-health-risks-no-one-saw-coming" target="_blank">menace to vulnerable people</a>. Some go further, arguing that these imitation-human AIs are bad for everyone.</p><p>“We are on a path to forgetting what it is to be human,” <a href="https://sherryturkle.mit.edu/" target="_blank">Sherry Turkle</a>, the MIT psychologist who has spent decades studying humans’ relations with computational devices, argues in her forthcoming book, <a href="https://www.sherryturkle.com/artificial-intimacy" target="_blank"><em><em>Artificial Intimacy</em></em></a> (Little, Brown and Co., 2026). While other researchers <a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2025.1687686/full" target="_blank">find that companion AI can leave some users better off</a>, many researchers agree that this kind of AI poses risks to minors and other vulnerable people.</p><p>“One thing that worries me about AI companions in the last year is that users have been getting younger and younger,” says Ge. For minors and others who lack the “AI literacy” to distinguish chatbots from humans, AIs “can be really dangerous if not guided in the correct way.”</p><p class="pull-quote">“We are on a path to forgetting what it is to be human.” <strong>—Sherry Turkle, MIT</strong></p><p>Yet companion-like AI is popular wherever the apps are available around the world. One Chinese survey of Gen Z people found <a href="https://cn.technode.com/post/2025-04-01/soul-ai-report2025/" target="_blank">60 percent had virtual partners of some kind</a>. A <a href="https://www.commonsensemedia.org/press-releases/nearly-3-in-4-teens-have-used-ai-companions-new-national-survey-finds" target="_blank">2025 poll of U.S. teenagers</a> found that nearly 75 percent had talked to an AI companion, and one-third of respondents said they found the AIs as satisfying or more satisfying than real-life connections.</p><p>China’s response is “the world’s strictest and most comprehensive law on the topic, unmatched by any other AI law,” the AI-law scholar <a href="https://luizajarovsky.com/" target="_blank">Luiza Jarovsky</a> wrote <a href="https://www.luizasnewsletter.com/p/the-worlds-strictest-law-on-human" rel="noopener noreferrer" target="_blank">last month</a>. The European Union, for instance, bans only AI that uses deception or manipulation to get users to do things that are harmful. And in the United States, there is no national policy. Rules that cover companion-like bots are in force in <a href="https://www.troutmanprivacy.com/2026/01/analyzing-the-new-ai-companion-chatbot-laws/#:~:text=SB%20243%20follows%20New%20York's%20AI%20Companion,of%20companion%20chatbot%20operators%20under%20both.%20I." rel="noopener noreferrer" target="_blank">California, Hawaii, and New York</a>, with similar laws coming into effect in <a href="https://www.multistate.ai/updates/vol-105-state-ai-companion-chatbot-laws#:~:text=As%20these%20systems%20have%20become%20more%20sophisticated%2C,(AZ%20HB%202311)%20in%20Arizona%20last%20week." rel="noopener noreferrer" target="_blank">nine more states</a> next year. Most emphasize protecting minors, and enforcement mechanisms typically involve lawsuits <em><em>after</em></em> harms occur, Yale’s Lucero says. In contrast, China’s approach aims to identify and prevent harms before they occur.</p><p>China’s government has a more explicit focus on <a href="https://spectrum.ieee.org/measuring-ai-societal-impact-khan" target="_blank">the possible harms to society</a> and government from AI companions, not just harms to individuals, Ge notes. The 15 July rules forbid companion bots to generate content that “spreads rumors,” incites “subversion of state power or the overthrow of the socialist system,” or endangers national honor, for example. Another motivation for concerns about AI romance is anxiety about the country’s declining birthrate, Ge says. “They want to try to monitor and regulate these kinds of nonprocreative activities invested intensively by young women.”</p><h2>Emotional Health Versus Economic Output</h2><p>The problem for all the governments trying to prevent AI harm, Ge says, is that they still want to encourage AI adoption in other walks of life. For example, the new regulations state that they don’t apply to customer service bots, work assistants, and educational AIs, on the assumption that these desirable uses of AI don’t create ongoing emotional connections.</p><p>Chinese regulators manage that tension—between promoting AI and protecting against it—by giving themselves room to adjust enforcement depending on circumstances, Lucero says. “A key component of their approach is that the state has the discretion to determine what the language means at any given point in time, as well as whom they choose to enforce against,” Lucero says. “China’s approach is to use regulation with a lot of relatively vague provisions, and they figure out what those provisions actually mean in practice after the fact.”</p><p>In any event, the rules also don’t address the underlying forces that cause people to create and rely on AI companions, Ge notes. In interviews with Gen Z Chinese women, Ge noticed a shared reluctance to get married and have children, and a feeling that “virtual love forms an alternative path.” On the other hand, some women in their 30s and 40s said they use AI companions as a supplement to real relationships and marriages.</p><p>“They would say they talk to the AI about the bitter things in their lives, and share the good things with their real human partners,” Ge says. But as time passed, some of these women felt closer to the AI than to their partners. In more recent interviews, some told Ge they felt a deeper attachment to the entity they shared negative feelings with.</p><p>The trend is an example of how strategically managing AI technologies can lead to unexpected places, Ge says. That’s why they believe that, despite prohibitions and dangers, “human-AI love will evolve and become an important intimate practice in the future.”</p><p>A law aimed at one particular technology probably isn’t sufficient to undo the societal pressures that make people turn to AI, Lucero agrees. “I don’t think you’re going to solve the problem of the low marriage rate or the low fertility rate by saying, ‘you can’t have an AI boyfriend.’ ”</p>]]></description><pubDate>Wed, 09 Sep 2026 10:00:07 +0000</pubDate><guid>https://spectrum.ieee.org/china-ai-chatbot-regulation</guid><category>China</category><category>Chatbots</category><category>Ai-regulation</category><category>Ai-laws</category><category>Mental-health</category><dc:creator>David Berreby</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/collage-of-the-3d-female-avatar-used-as-doubaos-mascot-with-message-bubbles-dangling-from-her-hand-via-puppet-strings.jpg?id=67740919&amp;width=980"></media:content></item><item><title>AI Slop Is Changing How Engineers Review Code</title><link>https://spectrum.ieee.org/ai-code-review-software-engineers</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-workers-overwhelmed-by-a-giant-black-ai-barrel-spilling-yellow-liquid-one-sits-at-a-computer-while-the-other-us.jpg?id=67739596&width=1200&height=800&coordinates=0%2C280%2C0%2C281"/><br/><br/><p>AI coding tools can now generate thousands of lines of code in minutes, helping companies build features, run tests, and fix issues faster. But the flood of AI-generated code still has to be reviewed. Large language models can produce code that looks clean on the surface but <a href="https://spectrum.ieee.org/responsible-ai" target="_blank">conceals sloppy mistakes</a> such as faulty assumptions, security vulnerabilities, or subtle errors that emerge only after deployment. Fixing those problems could erase the productivity gains AI promises.</p><p>Companies are responding to the onslaught of AI code slop by rethinking how they review code. New strategies are emerging. Among other approaches, engineers are scrutinizing plans before AI begins coding, deploying specialized AI agents to catch routine flaws, sending risky changes to human reviewers, or requiring developers to defend the code their agents produce. </p><p>The shift comes as the surge in <a data-linked-post="2658687697" href="https://spectrum.ieee.org/ai-code-generation-ownership" target="_blank">AI-generated code</a> puts new pressure on engineering teams. In a survey of more than 1,100 developers by <a href="https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding/" rel="noopener noreferrer" target="_blank">Sonar</a>, an AI code-verification startup, respondents estimated that AI contributed 42 percent of the code they added to shared codebases. Yet while developers found AI useful for explaining and prototyping code, 96 percent did not fully trust its output to work correctly.</p><p>Investors see an opportunity in closing that gap. In August, for example, AI code-review startup CodeRabbit<a href="https://www.reuters.com/technology/ai-code-review-platform-coderabbit-valued-15-billion-latest-funding-round-2026-08-12/" rel="noopener noreferrer" target="_blank"> raised US $143 million</a> at a $1.5 billion valuation, while claiming it performs more than 2 million reviews a week for 17,000 customers, including Nvidia, Indeed, and BMW Group. </p><p>The new era of code review will determine whether AI can ever provide code that is both faster and more reliable. It also has some software engineers thinking about the future of their profession: If entry-level engineers spend less time writing code themselves, how will they learn to judge it?</p><h2>AI Code Review Bottlenecks</h2><p>AI-written code is shifting the bottleneck from generating software to reviewing it. According to the Sonar study, 38 percent of developers said “more effort” is required to review AI-generated code than code written by their colleagues. Sixty-one percent of them said AI often produced code that looked correct but was “unreliable.”</p><p>For Synthesia, an AI video-generation platform, code review has become essential to its engineering workflow. In November 2025, Synthesia’s 118 engineers went all-in on AI coding tools like Claude Code. According to Peter Hill, Synthesia’s chief technology officer, the result has been a massive surge in code volume.</p><p class="pull-quote">“I don’t know if we ever get to the point where you can truly trust the agentic generation of code.” <strong>—Peter Hill, Synthesia</strong></p><p>That code demands close examination. As of August, the number of pull requests, or proposed changes to a codebase submitted for review, had risen 120 percent year over year, Hill says. Ninety-five percent of those requests contain AI-generated code.</p><p>One recurring problem is duplication. Hill says AI tools may not recognize that code for a task already exists, and they’ll write another version because they have limited context. Synthesia has found as many as 10 versions of the same function, leaving engineers to identify and remove redundant functions. Once that’s done, engineers retrain the AI agent so that doesn’t happen again. At the company’s scale, Hill describes getting the AI to produce the intended output an “enormous amount of work.”</p><h2>AI Agents in Code Review Workflows</h2><p>Some teams are trying to prevent review problems <em><em>before</em></em> AI generates a single line of code.</p><p>McLaren Stanley, a senior principal engineer at Amazon Stores, says he is using AI to modernize 17 years of code underlying Amazon’s mobile shopping app. His 70-person team supports more than 1,000 developers by maintaining the architectural backbone they need to build features. With AI writing the code, Stanley said, engineers spend more time deciding what it should do before generation begins.</p><p>Much of that work involves writing a “specification,” which is a detailed plan for what the AI agent should build and how. Preventing recurring mistakes before generating code can save engineers time later.</p><p>Stanley recalls how a missing instruction once caused an agent to generate 25,000 lines in the wrong version of the programming language Swift. Switching versions produced 600 errors it could not fix at once. Stanley discarded the code, updated the specification, and restarted the agent. Fifteen minutes later, it regenerated the code correctly.</p><p>Once the code exists, specialized AI agents can handle the first round of checks before a person steps in. </p><p>David Yanacek, a senior principal engineer at Amazon Web Services (AWS), says the company uses agents to test whether code works, check it against the original plan, and look for security flaws before a person reviews it.</p><p>That first pass becomes more important as AI-generated code volume increases. At Bonterra, a nonprofit software provider with about 290 engineers, proposed changes tripled within three months of adopting AI, according to Tanuja Korlepra, the CTO. Code entering review rose tenfold and review times tripled, making it impractical for engineers to inspect every line.</p><p class="pull-quote">“We refuse to let code review become a dumping ground for unchecked model outputs.” <strong>—</strong><strong>Samar Abbas, Temporal</strong></p><p>Bonterra’s agents compare code with the approved design, security rules, coding standards, and accessibility requirements, then report their confidence in the result. A low score or flagged problem sends the change to a person. Code involving payments, personal data, or other sensitive systems always receives human review.</p><p>“Agents do the reading and humans do the judging,” Korlepra says.</p><p>Synthesia also uses AI agents to decide where human review is necessary. Criteria set by engineers direct more scrutiny toward higher-risk changes. Altering an error message carries less risk than code that handles customer data or core business rules. Even so, fewer than 5 percent of changes bypass human review. </p><p>“I don’t know if we ever get to the point where you can truly trust the agentic generation of code,” Hill says.</p><p>Automated review does not change who is responsible for the resulting code.</p><p>When machines produce more code than engineers can closely read, human approval can become “theater approval,” according to JD Raimondi, chief AI architect at the software consultancy Making Sense. In other words, an engineer might confirm that the feature works, skim the code, and approve it, all without understanding the choices underneath.</p><p>Temporal, an open-sourced developer platform, puts the burden back on the person submitting the code. CEO Samar Abbas says code volume and review time have increased with AI. Under its “Send Back” policy, Temporal’s engineers must explain in their own words the agent’s design choices and how the code handles unusual conditions. Otherwise, the reviewer rejects it.</p><p>“We refuse to let code review become a dumping ground for unchecked model outputs,” Abbas said.</p><h2>Training Junior Engineers on AI</h2><p>As AI shifts engineering work from writing code toward judging it, companies are reconsidering how entry-level engineers gain experience.</p><p>Junior engineers at Making Sense have seen some of the largest productivity gains from AI, Raimondi says, raising concerns about what they no longer learn by doing. The consultancy keeps juniors involved in deciding why a customer needs a feature and how it should work, rather than limiting them to checking AI output.</p><p>IBM is using AI to give new engineers harder assignments sooner. Neel Sundaresan, IBM’s general manager of automation and AI, says recent graduates now work on product features and projects once reserved for senior level engineers. AI helps implement and test the code, but if it fails, juniors assess what went wrong and fix the issues before the work is passed to senior developers for final approval. Sundaresan estimates that AI can help junior engineers perform 70 to 80 percent of some tasks that once required a senior engineer.</p><p>Synthesia primarily hires mid- and senior-level engineers. Its less-experienced employees work with both a senior colleague and an AI agent, taking responsibility for parts of projects while learning to define what successful code should do.</p><p>At Bonterra, agents now perform many of the well-defined coding tasks that once trained new engineers. Juniors instead own outcomes alongside experienced colleagues, learning to direct agents, question their output, and remain responsible for the result. She says this approach can help junior engineers build the skills and knowledge needed to advance in their careers.</p><p>“If the industry stops hiring juniors, the industry stops producing seniors,” Korlepra said. </p>]]></description><pubDate>Tue, 08 Sep 2026 16:16:49 +0000</pubDate><guid>https://spectrum.ieee.org/ai-code-review-software-engineers</guid><category>Coding</category><category>Artificial-intelligence</category><category>Vibe-coding</category><category>Software-engineers</category><dc:creator>Aaron Mok</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-workers-overwhelmed-by-a-giant-black-ai-barrel-spilling-yellow-liquid-one-sits-at-a-computer-while-the-other-us.jpg?id=67739596&amp;width=980"></media:content></item><item><title>Google DeepMind Maps 9 Billion Possible DNA Variants</title><link>https://spectrum.ieee.org/alphagenome-atlas</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-glowing-blue-digital-dna-helix-with-data-patterns.jpg?id=67719885&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>DNA is often explained as a codebook or set of instructions for producing proteins, and ultimately, life. Some stretches of DNA, called genes, code for proteins, but the vast majority of DNA is considered “noncoding.” Some of it has no known function, while other segments are critical to regulating gene activity. </p><p>These regulatory elements can interact in complicated ways, and their effects can vary across different cells and tissues. Some also influence genes located far away in the genome. Understanding how changes in DNA affect this regulation “is fundamental to understanding most disease,” says<strong> </strong><a href="https://deboer.bme.ubc.ca/" rel="noopener noreferrer" target="_blank">Carl de Boer</a>, a genomicist at the University of British Columbia. </p><p>That’s why researchers are working to understand what every imaginable small variation in human DNA across the entire genome might mean for gene regulation. A recent AI tool built for that purpose from Google DeepMind, <a href="https://spectrum.ieee.org/alphagenome-ai-gene-regulation" rel="noopener noreferrer" target="_blank">AlphaGenome</a>, was originally <a href="https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome/" rel="noopener noreferrer" target="_blank">announced in 2025</a>. In January, <a href="https://spectrum.ieee.org/alphagenome-ai-gene-regulation" target="_self">a paper</a> published in <em><em>Nature</em></em> provided more details, and the model was released for public noncommercial use. The AI model can compare an original DNA sequence with an altered one and predict how the change might affect gene expression and other regulatory activity. But researchers had to select the variants they wanted to test, write code, and run the computationally demanding model themselves.</p><p>Now DeepMind has done that work in advance for all 9 billion possible single-letter changes to a reference human genome. Today, on 8 September, DeepMind <a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome" target="_blank">announced</a> the creation and public release of the <a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/alphagenome-atlas.pdf" target="_blank">AlphaGenome Atlas</a>, an online repository of precomputed predictions made using the AlphaGenome model. The Atlas offers a more approachable interface for scientists, without the need to write code or run the AlphaGenome model themselves. It also includes a much-requested new feature, a single-number impact score intended to show at a glance if a variant is likely to be meaningful. </p><p>“Understanding our DNA is a grand challenge,” says <a href="https://research.google/people/105667/?&type=google" rel="noopener noreferrer" target="_blank">Pushmeet Kohli</a>, VP of science at Google DeepMind. “Understanding this language of life can unlock so many things.”</p><p>The AlphaGenome predictions have some important limitations. For example, many diseases are associated with multiple genetic variants. And although AlphaGenome looks at a relatively large segment of DNA surrounding the variant in question—1 million base pairs—some DNA sequences, called enhancers, can regulate genes over very long distances, sometimes beyond the model’s field of view. Their effects are difficult to predict.</p><p>But the Atlas could still help scientists filter possibilities and prioritize lab experiments that would validate its predictions. In that way, it could greatly accelerate work in fundamental biology, disease research, and treatment development, says <a href="https://www.linkedin.com/in/avsec/" rel="noopener noreferrer" target="_blank">Žiga Avsec</a>, the genomics lead at DeepMind.</p><p>“It seems like they made a useful resource for people,” says de Boer, who recently helped create <a href="https://github.com/de-Boer-Lab/Genomic-API-for-Model-Evaluation" rel="noopener noreferrer" target="_blank">a framework</a> for better comparisons of computational models similar to AlphaGenome. He is not affiliated with DeepMind. </p><p>Although de Boer considers AlphaGenome the “field’s leading model,” he notes that it’s also “very slow and computationally intensive.” The Atlas could benefit people without access to newer hardware, or simply reduce the number of people repeating the same simulations.</p><p>The Atlas is freely available for noncommercial research, with the potential for commercial licensing.</p><h2>Computing 9 Billion Predictions</h2><p>The entire human genome contains roughly 3 billion base pairs. At each position there are three possible single-nucleotide substitutions, and therefore 9 billion variants in the Atlas. The complete dataset is around 1 petabyte.</p><p>“When we started thinking about this project, it seemed impossible to do that computationally,” says Avsec. Early estimates told the team they would need to improve their calculation speed by a factor of 80 in order to compile the Atlas in a reasonable amount of time.</p><p>To reach that target, the team gained advantages using a few different techniques, including model distillation, GPU kernel optimization, and the elimination of redundant calculations. “There was a lot of thought and engineering that we had to do in order to make this happen at this scale,” says Avsec.</p><p>AlphaGenome and the Atlas build on years of related work at DeepMind. In 2020, <a href="https://spectrum.ieee.org/alphafold-proves-that-ai-can-crack-fundamental-scientific-problems" target="_blank">AlphaFold</a> predicted the three-dimensional structure of proteins from amino-acid sequences. In 2023, AlphaMissense predicted whether 71 million possible variants that alter proteins were likely benign or pathogenic. Similar to the new Atlas, prediction results from those projects were made available in a <a href="https://www.ebi.ac.uk/training/online/courses/alphafold/classifying-the-effects-of-missense-variants-using-alphamissense/alphamissense-in-the-alphafold-database/" rel="noopener noreferrer" target="_blank">public database</a>. </p><p>The Atlas allows a scientist to look up a single variant and see more detailed information about the model’s prediction, including 11 different output types. But the top-line figure is a single-number impact score, which by its nature is a simplification of many aspects of those predictions. </p><p>“It has a clear use, but it also is probably going to be easily misinterpreted,” says de Boer. “We’re talking about a very complex system, and there’s a lot of moving parts.”</p>]]></description><pubDate>Tue, 08 Sep 2026 14:00:05 +0000</pubDate><guid>https://spectrum.ieee.org/alphagenome-atlas</guid><category>Genome</category><category>Google-deepmind</category><category>Genetics</category><category>Ai-models</category><dc:creator>Greg Uyeno</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-glowing-blue-digital-dna-helix-with-data-patterns.jpg?id=67719885&amp;width=980"></media:content></item><item><title>AI Efficiency Could Cost Us the Next Generation of Experts</title><link>https://spectrum.ieee.org/ai-engineer-skills</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/human-and-robotic-hands-share-a-caliper-over-technical-engineering-blueprints.png?id=67702640&width=1200&height=800&coordinates=0%2C90%2C0%2C91"/><br/><br/><p><span>A little over a decade ago, I led the controls design for a first-of-its-kind full digital-control system for a U.S. nuclear plant. It was, on paper, a beautiful machine—engineered to run itself the way a modern airliner does, with operators watching over a system that rarely needed them. And we made a decision that, to an efficiency-minded observer, looked backward: We deliberately left manual steps inside sequences the system could execute on its own.</span></p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/ai-engineer-skills&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p><span>We were solving a specific problem. An operator who only ever supervises automation slowly stops being an operator. The hands go cold. The mental model of what the plant is actually doing gets fuzzy. Then comes the day the automation hands control back. It’s always the worst day, because automation only quits when it’s confused or in trouble. But by then, you have a person in the chair who hasn’t truly operated the thing in years. The manual steps were there to keep the human current. It was inefficient by design, on purpose.</span></p><p>That plant, as it happened, was never built. It was shelved amid the politics and economics that surround <a href="https://spectrum.ieee.org/tag/nuclear-power" target="_blank">nuclear power</a> in this country, for reasons that had nothing to do with the engineering. But the design instinct outlived the project, and I’ve come to believe it’s the most useful idea I can offer to the argument now consuming every boardroom: What happens to human expertise when AI does the work that used to build it?</p><h2>AI Is Disrupting the Engineering Career Ladder</h2><p>The data has gotten hard to wave away. A Harvard University working paper covering some 65 million workers at more than 280,000 U.S. firms found that after companies adopted generative AI, <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5425555" target="_blank">junior employment fell roughly 9 percent</a> within six quarters relative to nonadopters, while senior employment kept right on growing. A Stanford analysis of ADP payroll records points the same way: The youngest workers in the most AI-exposed occupations <a href="https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/" target="_blank">lost ground after late 2022</a> while their more-experienced colleagues held theirs. The Stanford researchers found that the losses concentrate where AI automates the work; where it merely augments, junior employment holds steady or rises.</p><p>The causal story is still contested, and honesty requires saying so. Researchers at the New York Fed attribute much of the rise in young-graduate unemployment <a href="https://libertystreeteconomics.newyorkfed.org/2026/06/remote-work-leaves-younger-workers-sidelined/" target="_blank">not to AI but to remote work</a>, arguing that firms are reluctant to hire inexperienced people whom they cannot train and mentor at a distance. But notice what the explanations share. Whether a model is absorbing the formative work or distance is severing the mentorship around it, both describe the same broken mechanism: the apprenticeship channel through which expertise passes from senior to junior. Either way, “entry-level” has quietly come to mean “three years of experience required.”</p><p>Strip away the noise and you’re left with one deceptively simple problem: You cannot become a senior engineer without first being a <a href="https://spectrum.ieee.org/ai-effect-entry-level-jobs" target="_blank">junior one</a><em>.</em> Expertise is not downloaded. It is earned through failed builds, dead-end debugging sessions, and the “why on earth did that work” moments that a capable AI will now happily spare the newcomer. Spare them enough of those and you produce a cohort that can supervise a model on paper but never developed the gut sense to know when the model is confidently, catastrophically wrong.</p><p>Most of the commentary stops at the diagnosis, or reaches for policy solutions that treat the loss of junior jobs as an economic problem. Yet it’s also an engineering problem, and safety-critical fields have already spent decades learning how to solve it.</p><h2>Aviation’s Lessons About the Automation Paradox</h2><p>My own career started at the sharp end of automation. My first job out of school was verifying and validating the software in the digital jet-engine controller that decides, faster than any pilot could, how a fighter plane’s engine responds. Even then, in the late 1980s, the central tension was visible: The machine outperforms the human in routine cases, but the human is all that stands between the aircraft and disaster in the cases the machine didn’t anticipate. This tension is known as the <a href="https://spectrum.ieee.org/tag/automation-paradox" target="_blank">automation paradox</a>, in which increasingly capable automation gives human operators less practice, while leaving them only the most difficult situations.</p><p>Aviation learned, repeatedly and expensively, what happens when human skills atrophy inside that gap. The canonical example is <a href="https://en.wikipedia.org/wiki/Air_France_Flight_447" target="_blank">Air France flight 447</a>, which fell into the Atlantic in 2009. The proximate cause was mundane. Iced-over airspeed sensors fed the autopilot bad data, and it did what it is designed to do: It disconnected and handed control of the airplane back to the crew. What followed was not a hardware failure. It was a competence failure. A recoverable situation became an unrecoverable one because the pilots, conditioned by thousands of hours of watching the automation fly, could not read a high-altitude aerodynamic stall and hand-fly their way out of it. The airplane was working. The training the automation had quietly eroded was not.</p><p>The industry’s response is instructive, and it’s the same move we made in that nuclear control room. It did not rip out the autopilot. It built deliberate manual practice back in. In 2017 the FAA issued Safety Alert for Operators 17007, “<a href="https://www.faa.gov/sites/faa.gov/files/2022-11/SAFO17007.pdf" target="_blank">Manual Flight Operations Proficiency,</a>” declaring that “manual flight is the foundation upon which other technical flying skills are built.” The alert formally recognized skill decay as a hazard in its own right. Some airlines amended their procedures to encourage hand-flying both the initial climb and initial descent in benign conditions, knowingly trading a sliver of fuel efficiency to keep the crew’s raw flying skills alive. That trade is the whole point. A perfectly optimized system that produces incompetent operators is not optimized at all. It has simply moved its failure mode somewhere the spreadsheet can’t see it.</p><h2>Manual Gates Could Preserve Engineering Skills</h2><p>Put the aviation lesson and the nuclear instinct side by side and they point to one design pattern we now need in AI-augmented work: the deliberate “manual gate.”</p><p>A manual gate is a point in a workflow where a human takes the controls, not because it is the fastest way to get the task done, and not only as a safety interlock, but specifically to exercise and preserve a skill that would otherwise decay. The distinguishing feature is that it is chosen. You decide, as a matter of design, which competencies your organization must keep alive in human beings because those are the ones you will need on the bad day. Then you engineer the friction required to keep them warm.</p><p>Picture how this might work on a software team that leans on AI for most of its code. The team places a manual gate around the skill it can least afford to lose: <a href="https://spectrum.ieee.org/tag/debugging" target="_blank">debugging</a>. When a defect surfaces in a critical module, the assigned engineer—deliberately, often a junior one—must first reproduce the failure, trace it to root cause, and write an automated test that captures the bug, all with the AI assistant switched off. Only after the engineer commits to a diagnosis does the model come back on, to propose the fix, generate alternatives, and sweep the code base for similar bugs. The engineer then compares their diagnosis against the model’s. When the two disagree, that’s the design working, surfacing the disagreement before the bad day instead of during it.</p><p>This approach reframes the junior engineer entirely. The instinct today is to let AI do the entry-level work because it is faster and cheaper. But some of that work is not overhead to be eliminated. It is the training apparatus of your future senior staff, and you should protect it the way you’d protect any other piece of critical infrastructure. It may not be efficient this quarter, but dismantling it quietly mortgages your capability a decade out.</p><h2>Why Companies Must Keep Training Junior Engineers</h2><p>None of this is free, and pretending otherwise would insult the people who have to sign the budgets. A deliberate manual gate is, by construction, less efficient in the near term than full automation. Keeping juniors doing formative work and running the manual sequences costs something now to protect something later.</p><p>That’s a hard sell in a market that judges most leaders on quarterly results. A hired executive who carries “unnecessary” humans that AI could replace will hear about it from the board long before the payoff arrives. The math only works for someone insulated from that pressure: a founder with control, a private company, an institution with a genuinely long horizon, or a regulator willing to require workers to demonstrate their skills regularly, as pilots must. Which means the organizations most likely to preserve their own expertise are the ones structurally able to spend short-term margin on long-term capability; everyone else will need that outside push.</p><p>So here is the argument, in one line: Deliberate inefficiency is not waste. In safety-critical engineering we have always known it as insurance, and we buy it on purpose. As AI takes over the work where expertise is forged, the smart move is not to resist the automation. It is to keep our hands on the controls by design—so that when the automation fails, as it always eventually does, there is still someone in the chair who knows how to fly.</p>]]></description><pubDate>Wed, 02 Sep 2026 13:00:04 +0000</pubDate><guid>https://spectrum.ieee.org/ai-engineer-skills</guid><category>Engineering-careers</category><category>Generative-ai</category><category>Automation-paradox</category><category>Aviation</category><category>Nuclear-power</category><dc:creator>Richard Mitchell</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/human-and-robotic-hands-share-a-caliper-over-technical-engineering-blueprints.png?id=67702640&amp;width=980"></media:content></item><item><title>Cash In on the AI Boom by Renting Out Your Spare Compute</title><link>https://spectrum.ieee.org/ai-inference-distributed-computing</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-small-makeshift-server-rack-in-someones-garage.jpg?id=67703292&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>If you own an at-home server, a gaming computer, or just a laptop that doesn’t get much love, listen up. You can now put that spare computing power to use and earn some passive income in the process. AI companies are hungry for more compute to run AI inference—the process of using a pretrained model to respond to queries—and they’re willing to pay you for it. </p><p>“Imagine Uber or Airbnb, but for AI-inference computing tasks,” says <a href="https://www.linkedin.com/in/shzhv13/" rel="noopener noreferrer" target="_blank">Ilman Shazhaev</a>, founder and CEO of <a href="https://farlabs.ai/" rel="noopener noreferrer" target="_blank">Far Labs</a>, based in Abu Dhabi.</p><p>The AI boom has spurred on the construction of <a href="https://spectrum.ieee.org/5gw-data-center" target="_blank">massive data centers</a>, often damaging local communities by raising electricity prices, straining local water resources, causing environmental damage and noise, and being just plain ugly. Huge data centers are likely not going anywhere—training new frontier models and running AI models from leading companies will likely still be the purview of these behemoths. But now, several companies are providing AI inference on smaller, mostly open-source models. They are running inference on the preexisting computing power spread throughout homes and small businesses, and compensating the owners.</p><p>“Everyone thinks the only way to do it is data centers. And data centers are extractive for the communities in which they’re built, and they don’t return services or taxes or much of anything to the people there. So why not just turn this whole thing on its head?” says <a href="https://www.linkedin.com/in/johnfederico/" rel="noopener noreferrer" target="_blank">John Federico</a>, founder and CEO of <a href="https://www.evolvingedge.ai/" rel="noopener noreferrer" target="_blank">Evolving Edge</a>, in Austin, Texas. “The compute power is out there. If you can orchestrate it, then you’re actually adding value to those communities directly.”</p><p>The idea isn’t entirely new: From 1999 to 2020, a volunteer-based project called <a href="https://setiathome.berkeley.edu/" rel="noopener noreferrer" target="_blank">SETI@Home</a> used spare computers to search for signs of extraterrestrial life in radio-telescope data, for instance. But now, commercial companies are eager to use the same strategy. Shazhaev’s Far Labs is launching its platform Far AI in the coming weeks, while Federico’s Evolving Edge is currently in open beta. Other companies, like <a href="https://bless.network/" rel="noopener noreferrer" target="_blank">Bless Network</a>, <a href="https://salad.com/download" rel="noopener noreferrer" target="_blank">Salad</a>,<a href="https://vast.ai/hosting" rel="noopener noreferrer" target="_blank"> </a>and <a href="https://gradient.network/" rel="noopener noreferrer" target="_blank">Gradient</a> have started to provide similar platforms over the last year.</p><h2>Connecting to the network</h2><p>Federico has been a computer hobbyist since youth, and he has amassed a whole server in his basement to run his projects. “It just hit me one day—there’s all this talk about not having enough compute, and I just thought, well, 92 percent of the country has broadband, and you have people like me who have mini data centers in a closet,” he says.</p><p>Federico sees the potential hosts as people much like himself who have already invested in home servers, and he aims to make the process of selling spare compute as seamless for them as possible.</p><p>“Sign up for the program, install an application,” Federico says. “All we want to do is run jobs on your machine when you tell us we’re allowed to. The only thing we do is monitor the resource usage. And of course, you can give us a schedule.” With a large enough network of devices, the platform would have compute available whenever it’s needed.</p><p>Privacy and security are primary concerns for such hosts. To reassure the users that their local data is secure, and that no malware will be downloaded to their devices, the team open-sourced their scheduling software. “The node software is open source, so anyone can look at it, see what it does,” Federico says. </p><p>Shazhaev of Far Labs explains that the company’s software is designed around a principle known as “least privilege,” which grants both the host and the user the minimum access possible to accomplish the task. Inference runs as an isolated workload with authenticated, encrypted communication and explicit limits on the GPU, CPU, memory, storage, and network resources it may use. Customers do not receive arbitrary access to the host machine, and providers can inspect resource use, pause the node, revoke access, and remove the software at any time.</p><p>The protection also works in the other direction. Workloads are segmented, and only the minimum required information is exposed to an individual node. Sensitive enterprise workloads can be restricted to controlled hardware rather than routed through consumer devices.</p><h2>Divide and conquer</h2><p>Massive data centers still have advantages from the user perspective: top-of-the-line GPUs and CPUs, high-speed networking, thick cables, and <a href="https://spectrum.ieee.org/data-center-liquid-cooling" target="_blank">sophisticated cooling</a>. User devices are usually less powerful, more varied, and less reliably connected to one another.</p><p>“This is quite a difficult issue from the science angle,” Shazhaev says. “You want to do a similar level of tasks that are happening in those high-infrastructure data centers, and run them on the user device with limited capacity.”</p><p>Evolving Edge’s Federico says this is an issue for the largest, state-of-the art AI models. But those are not always needed and are often not even preferred. “There are numerous companies, once they reach a certain scale, suddenly paying for tokens on a state-of-the-art frontier model [that] no longer makes sense for their needs,” he says. “Instead, they are fine-tuning open-source models for specific tasks that they have in their business. These models don’t require anywhere near the resources that some of the state-of-the-art models do. It’s just using the right tool for the job.”</p><p>Smaller, open-source models can often fit on a single user device. But if that fails, there are tools to split a single inference task over multiple GPUs or CPUs. Evolving Edge is using an open-source tool called Ray to perform this splitting, while Far Labs has developed its own proprietary software that not only splits the workload but also wraps the splitting in a layer of security and reliability-providing software. </p><p>“One thing we have done is we shared the model,” Shazhaev says. “We take the model, we cut it into many pieces, and then these pieces will be distributed through different devices. And we have an orchestrator and a load balancer which manage the task flow, so each device processes a part of the task. Then we combine the answers in the main brain, the orchestrator.”</p><p>Through a combination of using smaller, more task-specific models, and splitting larger models between disparate devices, the teams claim they can perform inference much cheaper than a traditional data center “because we don’t have capital expenditure,” Shazhaev says.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A vertically mounted graphics card inside of a gaming PC rack." class="rm-shortcode" data-rm-shortcode-id="124c2cabf5d259385159e630919004de" data-rm-shortcode-name="rebelmouse-image" id="b08b9" loading="lazy" src="https://spectrum.ieee.org/media-library/a-vertically-mounted-graphics-card-inside-of-a-gaming-pc-rack.jpg?id=67703296&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Gaming PCs are a common source of spare computational power in the home. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Dizzaract</small></p><h2>The distributed advantage</h2><p>Not only is it cheaper to run inference this way, it is also more reliable, Shazhaev claims. The companies have access to a distributed network of computing resources, rather than one giant device that can experience outages. Shazhaev compares this to cryptocurrencies and their resilience through decentralization. </p><p>“Today, to shut down Bitcoin, you need to nuke the whole planet. Here, we have the same concept,” Shazhaev says.</p><p>Federico explains that this resiliency would be beneficial not just for AI inference, but for all kinds of applications, including smart cities, environmental sensors, autonomous vehicles, and more. During an Amazon Web Services <a href="https://www.theguardian.com/technology/2025/oct/24/amazon-reveals-cause-of-aws-outage" target="_blank">outage</a> in 2026, for example, smart beds were stuck in their upright positions and their users couldn’t adjust them. Federico says that a distributed network where everything doesn’t need to be routed through a single data center, say, in Ashburn, Va., would make those kinds of outages much less impactful. “We could lose 100 nodes in a network of 250,000 and it wouldn’t matter,” he says. </p><p>If the network of user devices is substantial enough, every job can be routed to a nearby device, decreasing the latency. Far Labs claims a latency of 100 milliseconds or less on its platform. The lower cost and lower latency of this approach may even enable new use cases, such as in-game AI video generation, which is currently prohibitively slow and expensive.</p><p>“OpenAI last year had US $30 billion in revenue, but they closed the financial year at an $8 billion loss. Why? The official reason is due to the high cost of inference,” Shazhaev says. “And those are mostly text models. For gameplay, you have audio, video, animations: It’s heavy data, and you need real-time responses. So, we’ve been trying to solve this issue.”</p><p>All of these companies are trying to tap into an untapped resource of local compute and hoping it’ll benefit the device hosts and users alike.</p><p>“All these big guys are running around building data centers,” Shazhaev says, “but I believe there is enough compute power that already exists in the world.”</p>]]></description><pubDate>Tue, 01 Sep 2026 14:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ai-inference-distributed-computing</guid><category>Distributed-computing</category><category>Inferencing</category><category>Seti-home</category><category>Data-centers</category><dc:creator>Dina Genkina</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-small-makeshift-server-rack-in-someones-garage.jpg?id=67703292&amp;width=980"></media:content></item><item><title>New Platform Peers Inside AI’s Black Box</title><link>https://spectrum.ieee.org/silico-ai-interpretability</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/goodfire-ais-silico-uses-agents-equipped-with-interpretability-tools-to-examine-the-reasoning-behind-an-ai-model.jpg?id=67668252&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p><span>Prompt Claude, ChatGPT, Gemini, or any other popular <a data-linked-post="2655083903" href="https://spectrum.ieee.org/what-is-deep-learning" target="_blank">large language model</a> with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer.</span></p><p>This mysterious behavior can be useful in some situations. But—as highlighted by a recent <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank">incident</a> where <a href="https://spectrum.ieee.org/tag/openai" target="_self">OpenAI</a> could not explain why its advanced prerelease model hacked AI company Hugging Face—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI “thinking” and outputs has never been greater.</p><p><a href="https://www.goodfire.com/">Goodfire</a>, an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a <a href="https://www.goodfire.com/grants" target="_blank">new grant program</a> offering US $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes.</p><h2>Mechanistic interpretability</h2><p>Founded in 2024 and based in San Francisco, Goodfire aims to provide the tools that build the next generation of safe and powerful AI by understanding the structures inside them instead of treating AI models as black boxes. “Treating models like black boxes isn’t inevitable; it’s a choice,” says <a href="https://www.linkedin.com/in/eric-ho-53981862/" target="_blank">Eric Ho</a>, Goodfire cofounder and CEO. “With the right interpretability tools, we can see how models actually work.”</p><p>The tools Ho refers to are built around a concept called mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by interpreting the model’s weights, activations, and attention patterns, and mapping its neurons and the pathways between them.</p><p>Mechanistic interpretability tools span the gamut. One approach is mapping a model’s activations in response to controlled prompts, and matching those activation patterns to a set of concepts that humans can understand. Another tack is tracking changes in model weights before and after a specific training run in order to spot and understand what changed. Yet another option is changing specific model weights or activations and observing how that affects the model’s output. </p><p>Silico combines a broad range of these tools, and provides a layer of AI agents to help users understand their model. Users describe what they want to investigate about their AI model in plain language, asking things like ”Find out when and why my model is hallucinating.” The platform then autonomously builds an experimental plan involving a host of tasks that can be performed using the various interpretability tools and techniques at its disposal. It then sends out agents to perform these tasks in parallel. Completion of these subtasks should add up to an answer to the original prompt, or at least insights that can be inspected and built upon. </p><p>Ho says: “In a sense, Silico is like a microscope to peer inside an AI model to understand which parts are responsible for what behavior, and even edit those parts directly.”</p><h2>Understanding Alzheimer’s and AI</h2><p>These tools have already been used to make some impressive advances in a host of fields. In medicine, for instance, <a href="https://www.primamente.com/" rel="noopener noreferrer" target="_blank">Prima Mente</a>, an AI company based in the United Kingdom, <a href="https://www.goodfire.com/research/interpretability-for-alzheimers-detection" rel="noopener noreferrer" target="_blank">worked with Goodfire</a> to understand its Pleiades epigenetic foundation model. The model performed well at its task of detecting Alzheimer’s disease from blood samples, but the company didn’t know why.</p><p>“We reverse-engineered Pleiades and found it was using DNA fragment-length patterns to make its predictions—a signal humans hadn’t used to detect Alzheimer’s before,” recalls Ho. In other words, the team had discovered that Pleiades was using a completely new biomarker for the disease. “As far as we know, it’s the first significant finding in the natural sciences discovered purely by reverse-engineering a foundation model,” Ho adds.</p><p>Elsewhere, Silico is being used to explore deep questions surrounding AI. Cameron Berg, founder and director of <a href="https://reciprocalresearch.org/" rel="noopener noreferrer" target="_blank">Reciprocal Research</a> (a New York nonprofit research organization he founded to explore methods of gauging AI cognition), says that Silico almost fell out of the sky at the right time for him and his research. “Silico has been really helpful for operationalizing my research agenda and executing on it way faster than I would have expected,” he says. “I feel like I have basically become the PI [principal investigator] and my research scientists and research engineers are AI systems.”</p><p>Berg sees general access to Silico and tools like it leading to greater trust in AI’s ability to conduct research tasks, which will accelerate the scientific process across the board. But beyond scientific research, the widespread release of Silico could signal a shift in how AI innovators build, debug, and deploy their models. </p><p>“I think it’s a mistake to not understand the most consequential technology of our time, particularly given the emergent behavior we’re seeing from increasingly capable AI agents,” says Ho. “If we truly understand how AI models think, instead of discovering and trying to correct their behavior retroactively, we can design them intentionally and shape how models behave to be safer and more reliable.”</p>]]></description><pubDate>Wed, 26 Aug 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/silico-ai-interpretability</guid><category>Ai-interpretability</category><category>Ai</category><category>Science</category><dc:creator>Benjamin Skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/goodfire-ais-silico-uses-agents-equipped-with-interpretability-tools-to-examine-the-reasoning-behind-an-ai-model.jpg?id=67668252&amp;width=980"></media:content></item><item><title>AI Companion Robots Are Closing the Human Connection in Modern Homes</title><link>https://spectrum.ieee.org/ollobot-ai-companion-robot</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/cute-home-robot-on-carpet-in-cozy-living-room-with-beige-sofa-and-warm-lighting.jpg?id=67154308&width=1200&height=800&coordinates=0%2C0%2C0%2C0"/><br/><br/><p><em>This article is brought to you by <a href="https://ollobot.com/" target="_blank">Ollobot</a>.</em></p><p>From about 2017, individuals began to truly connect with the initial wave of companion robots. These devices had personality, moved around, joked, and answered when you spoke to them. Most early companion robots, however, were still limited by simple voice-command interactions and narrow functionality. Once the novelty wore off, many ended up sitting unused on shelves. As some of those companies went out of business and turned off their servers, many owners likened it to losing a pet.</p><p>What Ollobot describes as “gentle intelligence” is a useful way to think about where the serious work in this category is going. Not toward more powerful assistants, but toward more present ones.</p><h2><a target="_blank"></a>The problem companion robots were trying to solve<strong></strong></h2><p>Loneliness is not a niche issue. According to one <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9957792/" target="_blank">study</a>, nearly one out of three elderly adults resides alone, meaning they do not have daily companions. <a href="https://pubmed.ncbi.nlm.nih.gov/20533912/" target="_blank">Research</a> also shows that children whose parents have migrated for work, leaving them in the care of relatives, were 2.5 times more likely to experience loneliness than children whose parents remain with them. Among working adults living alone in urban environments, similar <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6530780/" target="_blank">patterns</a> of social isolation emerge, even if they are less visible.</p><p>Over the years, technology has time and again attempted to solve this problem via video calls, smart speakers, and messaging apps without much success. Those tools are geared towards communication between people that already have relationships. They do not create presence. They schedule it. That is the gap that a new generation of AI companion robots is being engineered to fill.</p><h2><a target="_blank"></a>Today’s AI robots are different<strong></strong></h2><p>Today’s companion robots are not just cute and cuddly. They are designed with psychological research, clinical insight and long-term interaction models to be truly useful in real homes.</p><p>Three fundamental shifts define the current generation:</p><ol><li><strong>From reactive to proactive response. </strong>Older robots relied on you speaking to them, but modern robots monitor a room with cameras, microphones, and surroundings sensors to initiate interactions without your input, and they can pick up on your emotions.</li><li><strong>From function-oriented to emotion-oriented design.</strong> The original pitch for companion robots was about what they could do. The question driving the serious work now is how they make you feel, which is a harder engineering problem and a more honest framing of what the product is actually for.</li><li><strong>From standalone hardware to connected ecosystems.</strong> Leading brands are creating platforms rather than devices with software included as a built-in layer and remote access from the beginning.</li></ol><p>The <a href="https://www.grandviewresearch.com/industry-analysis/ai-companion-market-report?__cf_chl_f_tk=xJq0xhwCGb6n830sK3lMN.cyk2m73bms6.RauBufPho-1783406928-1.0.1.1-TsUMt2nxUgZ5nDX79zhxxVYkTv9xk2r_.E2rbXRHByA" target="_blank">global AI companion market</a> size was valued at US $36.8 billion in 2025 and is projected to grow from $48 billion in 2026 to $318 billion by 2033, at a compound annual growth rate of 31 percent from 2026 to 2033.<em><span><br/></span></em></p><h2><a target="_blank"></a>Three household scenarios and interaction models<strong></strong></h2><p>Ollobot’s advanced AI family companion robot <a href="https://ollobot.com/" target="_blank"><span>OlloNi SS1</span></a> addresses a number of gaps in what existing technology offers.</p><p><strong>Elderly individuals living alone.</strong> The combination of proactive interaction, fall detection, and persistent presence addresses both safety and companionship without the social overhead of asking family members to check in more frequently.</p><p><strong>Children in households where parents work far from home.</strong> The SS1 functions as a consistent companion that already knows a child, their preferences, their moods, and their routines. The remote connection features allow parents to stay present without requiring a scheduled call, and the life recording system gives them a passive window into their child’s days that feels less clinical than a monitoring camera.</p><p><strong>Single professionals living alone in cities.</strong> The SS1 adapts to daily routines, builds up a preference model over time, and provides ambient social presence without demands.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Cute home robot with a purple cover and cartoon face displayed on its screen." class="rm-shortcode" data-rm-shortcode-id="bd8ec4a867547204437967b3f231338f" data-rm-shortcode-name="rebelmouse-image" id="8d48b" loading="lazy" src="https://spectrum.ieee.org/media-library/cute-home-robot-with-a-purple-cover-and-cartoon-face-displayed-on-its-screen.jpg?id=67154351&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">OlloNi SS1 adapts to daily routines over time.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><h2>What OlloNi SS1 is doing differently?</h2><p>Ollobot’s goal in building intelligent companion robots is to address the gaps in technology and capability, using innovation not to automate tasks but to fill emotional voids.</p><p>Much of the robotics industry has historically pursued human imitation — machines that speak, look, or behave like people. The SS1 is instead designed around familiarity and long-term coexistence rather than realism.</p><p>The system integrates multiple subsystems operating in parallel, including visual perception, audio processing, mobility control, and interaction management. It is equipped with a multi-chip AI 4K vision module capable of facial recognition and motion tracking. One small but revealing detail is the inclusion of a physical privacy cover for the camera — a mechanical solution to concerns that software settings alone may not fully resolve.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Person playing with a red plush robot toy that has a glowing digital face and eyes" class="rm-shortcode" data-rm-shortcode-id="b39299c23f2f658c0a0284c5803d36ae" data-rm-shortcode-name="rebelmouse-image" id="daa03" loading="lazy" src="https://spectrum.ieee.org/media-library/person-playing-with-a-red-plush-robot-toy-that-has-a-glowing-digital-face-and-eyes.jpg?id=67154349&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">OlloNi SS1 can actively integrate into family activities, and it can autonomously move closer to capture memorable moments or reposition itself to remain engaged in ongoing interactions.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p><span>The robot supports advanced mobility across multiple indoor surfaces, including wooden floors, ceramic tiles, and low-pile carpets, with slope climbing capability up to 3.5 degrees. Rather than remaining in a fixed location, it can move naturally throughout the home to stay close to household members as daily activities unfold. </span></p><p><span>For example, the OlloNi SS1 may greet family members when they arrive home, follow an older adult from the living room to the kitchen while continuing a conversation, remind a child to take a study break after a prolonged period of inactivity, or notice that someone appears unusually quiet and gently check in. During family activities, it can autonomously move closer to capture memorable moments or reposition itself to remain engaged in ongoing interactions.</span></p><p class="pull-quote">The robot continues to evolve over time, with over-the-air updates that deliver new features, performance improvements, and AI enhancements</p><p>It also incorporates fall detection with optimized accuracy for safety monitoring scenarios. A 6-microphone array enables omnidirectional voice pickup with an effective voice capture range of up to 5 meters, supporting reliable wake-word detection and far-field interaction.</p><p>To support continuous companionship, much of the robot’s AI processing takes place directly on the device through its “heart module” architecture, with 16 GB of memory and 64 GB of local storage. This enables the system to retain household memories, recognize familiar faces, and respond with lower latency, making interactions feel more natural even during everyday routines.</p><p>Because companion robots are expected to remain available throughout the day rather than only during brief interactions, the SS1 is designed for extended operation, offering up to 12 hours of standby time and around 5 hours of active interaction on a single charge. This allows it to accompany users through meals, conversations, playtime, and other daily activities without frequent interruptions.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Close-up of toy robot with glowing red heart and purple fur on a beige body." class="rm-shortcode" data-rm-shortcode-id="ad4e47da5ac276c65c47c13fb8cbd4a2" data-rm-shortcode-name="rebelmouse-image" id="0261b" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-toy-robot-with-glowing-red-heart-and-purple-fur-on-a-beige-body.jpg?id=67154317&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">To support engaging interactions, much of the robot’s AI processing takes place directly on the device through its “heart module” architecture.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p><span>Like the relationships it is designed to build, the robot continues to evolve over time. Running on Android OS with over-the-air (OTA) updates, the system continuously receives new features, performance improvements, and AI enhancements, allowing its capabilities to grow alongside the household it serves.</span></p><p>The robot’s behavioral model also improves over time. Rather than reacting to isolated commands, it attempts to establish a baseline understanding of household routines and individuals. Changes in behavior — prolonged quietness, unusual inactivity, or emotional cues — become triggers for interaction.</p><h2>Presence instead of utility</h2><p>Several features in the OlloNi SS1 illustrate this emphasis on presence and continuity in its interactions.</p><p>The system can identify different household members, including pets, and adapt responses accordingly. Remote communication features allow family members to connect through the device without treating every interaction like a scheduled call. Environmental sensors support contextual reminders tied to weather or room conditions.</p><p>Its “2+1” multi-display configuration is also designed around emotional communication. Two circular side displays function as expressive “emotional eyes,” while a separate primary display handles information and structured interaction. The separation allows emotional signaling and functional communication to operate independently, creating more intuitive nonverbal interaction even when no dialogue is taking place.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Cute red robot pet in checkered shirt sits on rug in cozy, warmly lit living room" class="rm-shortcode" data-rm-shortcode-id="f177ff7ea3f02aacf9e12a93a6b7e445" data-rm-shortcode-name="rebelmouse-image" id="5f2a7" loading="lazy" src="https://spectrum.ieee.org/media-library/cute-red-robot-pet-in-checkered-shirt-sits-on-rug-in-cozy-warmly-lit-living-room.jpg?id=67154312&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The robot’s behavioral model improves over time. Rather than reacting to isolated commands, it attempts to establish a baseline understanding of household routines and individuals.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p>The SS1 also includes an automated life-recording system built on facial recognition and behavioral-event detection that can capture moments such as laughter, physical closeness, or group interaction automatically. An integrated AI vlog engine can then organize those moments into edited short-form videos with automated sequencing and soundtrack generation. The design intent is to preserve spontaneous domestic moments without requiring active documentation behavior from users.</p><p class="pull-quote">An integrated AI vlog engine can<span> organize recorded</span><span> moments into edited short-form videos with automated sequencing and soundtrack generation</span></p><p>Visual data is processed primarily on the device through the SS1’s on-device AI architecture, with household memories stored locally and managed within Ollobot’s proprietary ecosystem instead of being shared with third-party smart home platforms. Access to recordings and live feeds is restricted to authorized users through the companion app, while encrypted communication helps protect data during remote access. Users also retain direct control over recording preferences, and the physical camera privacy cover provides an additional hardware-level safeguard whenever visual monitoring is not desired.</p><div class="ieee-sidebar-small"><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Ollobot logo with circular icon and bold lowercase text on light background" class="rm-shortcode" data-rm-shortcode-id="c161d66f60be15a6d25fa8fcac7ad1d7" data-rm-shortcode-name="rebelmouse-image" id="8eca7" loading="lazy" src="https://spectrum.ieee.org/media-library/ollobot-logo-with-circular-icon-and-bold-lowercase-text-on-light-background.png?id=67154360&width=980"/></p><p>Learn more at <a href="https://ollobot.com/" target="_blank">ollobot.com</a>.</p></div><p>Remote communication is similarly structured around persistence rather than transaction. Traditional video calls are episodic and screen-bound; the SS1 instead acts as a continuously present interface embedded inside the household environment. Through autonomous mobility, environmental awareness, and persistent household memory, remote family members interact with an ongoing domestic context.</p><h2><a target="_blank"></a>The larger shift to “gentle intelligence”</h2><p>Ultimately, gentle intelligence is not about making robots behave more like humans — it is about helping them fit more naturally into human lives. Each OlloNi SS1 unit develops a unique behavioral profile based on its household. Two units running in different homes for a year will have become meaningfully different from each other, shaped by the specific people, habits, and rhythms of where they live.</p><p>That kind of long-term personalization is what early companion robots never had. It is also what makes the difference between a product that ends up on a shelf and one that actually earns its place in a home.</p><p>Learn more at <a href="https://ollobot.com" target="_blank">ollobot.com</a>.</p>]]></description><pubDate>Tue, 25 Aug 2026 10:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ollobot-ai-companion-robot</guid><category>Social-robots</category><category>Human-robot-interaction</category><category>Computer-vision</category><category>Companion-robots</category><category>Multimodal-ai</category><category>Ai-robots</category><dc:creator>Ollobot</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/cute-home-robot-on-carpet-in-cozy-living-room-with-beige-sofa-and-warm-lighting.jpg?id=67154308&amp;width=980"></media:content></item><item><title>Self-Driving Cars Could Someday Take Requests</title><link>https://spectrum.ieee.org/autonomous-vehicles-motion-planner-llm</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/diagram-of-a-vehicle-merging-into-an-adjacent-lane-with-surrounding-traffic-with-predicted-motions-labelled.jpg?id=67659079&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>
<em>This article is part of our exclusive <a href="https://spectrum.ieee.org/collections/journal-watch/" target="_blank">IEEE Journal Watch series</a> in partnership with IEEE Xplore.</em>
</p><p>The idea of letting a machine do the driving for you may put a lot of people off <a data-linked-post="2668536834" href="https://spectrum.ieee.org/autonomous-vehicles-great-at-straights" target="_blank">autonomous vehicles</a>. But research could make it possible to backseat-drive an autonomous vehicle just as you might with a human driver.</p><p><a data-linked-post="2671006250" href="https://spectrum.ieee.org/darpa-grand-challenge" target="_blank">Self-driving cars</a> carefully balance a host of parameters to ensure a smooth ride, including things like speed, acceleration, and the smoothness of turns. But human driving preferences can often vary depending on how much of a rush they’re in, whether they’re feeling carsick, or how busy the traffic is.</p><p>These cars have a software component called the motion planner, which is responsible for choosing a safe and efficient path through traffic. The motion planner is normally tuned by engineers before the vehicles hit the road so that there’s little scope for passengers to adjust a vehicle’s driving style on the fly. But now researchers at the Delft University of Technology (TU Delft) in the Netherlands have developed a system that uses a large language model (LLM) to translate natural-language user requests such as “I am running late, go fast” into adjustments to a self-driving control system. The researchers <a href="https://arxiv.org/abs/2606.10974v1" rel="noopener noreferrer" target="_blank">posted their preprint on arXiv</a> and are presenting the work at the <a href="https://ieee-itss.org/conf/itsc/" rel="noopener noreferrer" target="_blank">IEEE Intelligent Transportation Systems Conference</a> in September.</p><h2>LLMs Personalize Autonomous Driving</h2><p>The system doesn’t give users direct control over the vehicle’s driving decisions; it simply tunes the parameters of a safety-aware motion-planning algorithm, which helps to keep the vehicle’s behavior within safe bounds. And the system keeps the human in the loop by describing how it’s going to alter its behavior in nontechnical language, and by asking the passenger to confirm before making changes. When the system was tested in simulation, the researchers found it adjusted the speed and smoothness of driving in line with natural-language instructions.</p><p>“The motion-planning problem is not only about reaching a place while avoiding collisions, it’s also how you do it,” says lead author <a href="https://sites.google.com/view/dmartnezbaselga" rel="noopener noreferrer" target="_blank">Diego Martinez-Baselga</a>, a postdoctoral researcher at TU Delft. “The motivation here is trying to make the way the autonomous car drives adaptable by end users easily, just by talking to the car.”</p><p>Previous research has investigated the potential of using LLMs and video-language models (VLMs) to direct decision-making for self-driving vehicles, but the researchers deliberately targeted driving style instead. Using LLMs and VLMs to directly control vehicles faces several challenges, says Martinez-Baselga. These include relatively slow response times, which can make these models unsuitable for the fast-paced decision-making required in driving, and the fact that they can’t provide concrete performance guarantees in the way a deterministic motion planner can.</p><p>Instead, the researchers used an LLM’s language and reasoning capabilities to translate fuzzy human preferences into something a vehicle’s motion planner can use. The system relies on a model predictive-path integral controller previously developed by the researchers, which identifies multiple paths the vehicle could take to reach its goal and then judges them on various criteria, including speed, steering angle, and collision probability. It then finds an optimal path that is a combination of the trajectories that scored best on those judging criteria.</p><p>The team combined this with OpenAI’s GPT-4o-mini model to parse passengers’ natural-language suggestions and use them to tune how the controller chooses its path. The model is given the users’ prompt and a natural-language description of the scenario the vehicle is operating in. The description was handwritten by the researchers for the purposes of the study, but it could ultimately be provided directly by a car’s perception system, says Martinez-Baselga.</p><p>The model doesn’t directly tweak the settings of the controller; it uses the prompt to rate the relative importance of the judging criteria the controller uses to assess trajectories. This rating is then used to adjust each criteria up or down either side of a safe baseline set by the researchers. So, if a user says they are feeling dizzy, the LLM will dial up parameters that encourage smooth steering and gentle acceleration to make the vehicle favor more sedate travel.</p><p>Prior to making any changes, however, the model first presents the user with a natural-language description of the adjustments it plans to implement. The user can then sign off on the plan or make further suggestions. The system is also interactive, so the user can request further adjustments if the vehicle’s behavior doesn’t match expectations or the user‘s preferences change.</p><p>Martinez-Baselga says this human-in-the-loop system allows the passenger to catch instances when the model misinterprets prompts. But it also helps deal with the inherent subjectivity of suggestions like “go faster” or the possibility that models don’t accurately describe changes they plan to make. In that case the passenger can simply follow up with additional prompts “as you would do if you were in a taxi or with a friend that is driving,” says Martinez-Baselga.</p><p>The researchers tested the system in the popular <a href="https://www.nuplan.org/nuplan" rel="noopener noreferrer" target="_blank">self-driving simulator nuPlan</a> in scenarios that involved merging onto a busy highway. Across eight different prompts, the system changed the controller’s parameters in ways matching user intent, with requests for a more comfortable ride dialing up smoothness and those indicating urgency leading to higher speeds.</p><p>This isn’t the first time LLMs have been used to tune a self-driving car’s motion planner. <a href="https://ee.ethz.ch/the-department/people-a-z/person-detail.MjE0NjI3.TGlzdC8zMjc5LC0xNjUwNTg5ODIw.html" rel="noopener noreferrer" target="_blank">Nicolas Baumann</a>, a Ph.D. student at ETH Zurich in Switzerland, <a href="https://arxiv.org/abs/2504.11514" rel="noopener noreferrer" target="_blank">published research</a> last year in which an LLM tweaked the parameters of a model racing-car controller, allowing the user to alter driving style but also give more concrete instructions like “reverse the car” or “maintain a specific speed.”</p><p>The strength of the approach, says Baumann, is that separating the LLM from the main controller means that even if the model hallucinates, it can’t do anything dangerous. “You get the possibility of language interaction, but you can guarantee that it is going to be within the constraints of this classical controller, so you can bake in safety,” he says. However, setting these constraints requires considerable engineering work, he adds.</p><p>And if you want provable safety, you need to go a step further, says <a href="https://www.professoren.tum.de/en/althoff-matthias" rel="noopener noreferrer" target="_blank">Matthias Althoff</a>, a professor of cyberphysical systems at the Technical University of Munich. His group <a href="https://ieeexplore.ieee.org/document/11640900" rel="noopener noreferrer" target="_blank">built a system</a> that gets an LLM to suggest driving decisions, but then uses a mathematical process to check them against traffic rules and predictions about the behavior of other road users. This makes it possible to verify their safety before committing to them, something the Delft paper doesn’t provide. “As with any LLM, it is not guaranteed that the result is correct,” says Althoff. “For that reason, we safeguard the decisions of the LLM in our works.”</p>]]></description><pubDate>Mon, 24 Aug 2026 15:51:45 +0000</pubDate><guid>https://spectrum.ieee.org/autonomous-vehicles-motion-planner-llm</guid><category>Autonomous-vehicles</category><category>Journal-watch</category><category>Large-language-models</category><dc:creator>Edd Gent</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/diagram-of-a-vehicle-merging-into-an-adjacent-lane-with-surrounding-traffic-with-predicted-motions-labelled.jpg?id=67659079&amp;width=980"></media:content></item><item><title>What It Takes to Be an Adaptable Engineer</title><link>https://spectrum.ieee.org/adaptable-engineer-core-skills</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/silhouette-of-person-at-computer-atop-a-weather-vane-labeled-ai-against-sky-background.png?id=67652283&width=1200&height=800&coordinates=0%2C191%2C0%2C192"/><br/><br/><p>The AI boom has disrupted the way engineers work, introducing new tools to learn, raising expectations for what teams can achieve in a workday, and making it <a href="https://spectrum.ieee.org/technical-interview-ai-arms-race" target="_self">harder to get hired</a> in the first place. This makes it difficult to advise students on <a href="https://spectrum.ieee.org/top-programming-languages-2025" target="_self">which specific coding languages</a> or technical skills they should learn. So amidst the uncertainty, advice for young professionals often turns to a common refrain: Be adaptable. But what does adaptability look like in practice? </p><p> Engineers often operate on the cutting edge of technology, so dealing with change is a normal part of the job, says <a href="https://search.asu.edu/profile/2452662" rel="noopener noreferrer" target="_blank">Samantha Brunhaver</a>, an associate professor of engineering at Arizona State University, in Tempe. Yet university curricula and training in the workplace often don’t prepare students for this. </p><p>“We tell engineers that they need to be adaptable when they graduate, but we don’t actually explain what that means, demonstrate what that looks like, [or] help make sure that they’re developing it,” says Brunhaver, who <a href="https://engineering.asu.edu/spring2020/brunhaver-seeking-to-augment-engineering-education-with-an-adaptive-mindset/" rel="noopener noreferrer" target="_blank">received a National Science Foundation award in 2020</a> to study how to foster greater workplace adaptability among young engineers. For this ongoing project, she has interviewed engineering managers, early career employees, and undergraduates about their experiences. </p><p> Part of the problem, she says, is that every employer has its own idea of what to be adaptable means. Generally, Brunhaver defines adaptability as “the ability to recognize that a change or uncertainty is occurring, and then respond effectively to that change.” But the skill is context-dependent. In software engineering, that might mean responding to turnover in the tools you use on a daily basis, while aerospace or biomedical engineers may need to keep track of changing procedures and regulations. “Managers are all saying adaptability is important,” Brunhaver says, “but defining it in different ways.” </p><p> At the same time, engineers are all contending with changes beyond these industry-specific expectations. Jobs in the technology, media, and telecom sectors are experiencing the fastest pace of skill turnover, according to a <a href="https://www.pwc.com/gx/en/issues/artificial-intelligence/job-barometer/2026/pwc-aijb-2026-technology-media-and-telecoms-report.pdf" rel="noopener noreferrer" target="_blank">June 2026 report</a> on the effects of AI from the professional services network <a href="https://www.pwc.com/us/en.html" rel="noopener noreferrer" target="_blank">PwC</a>. And the World Economic Forum’s most recent <a href="https://reports.weforum.org/docs/WEF_Future_of_Jobs_Report_2025.pdf" rel="noopener noreferrer" target="_blank"><em><em>Future of Jobs Report</em></em></a>, published in 2025, found that employers across all sectors expect 39 percent of workers’ core skills to change by 2030. This uncertainty can be uncomfortable. But with the right mind-set and support from leadership, adaptability can help keep you afloat. </p><h2>How to Cultivate Adaptability</h2><p>The AI transition is a big shift—but not an unprecedented one, says <a href="https://www.microsoft.com/en-us/research/people/jennbu/" rel="noopener noreferrer" target="_blank">Jenna Butler</a>, a research scientist at Microsoft who studies developer well-being and productivity. </p><p>During this type of paradigm shift, there is often a “chaos period” when a new normal is being established, Butler says. In AI’s case, it challenges the understanding of what a computer can do. “I think we’re still in this in-between, difficult period that we’ve seen before, but [it] is maybe moving faster than it has historically.” Software engineers—in one of the fields <a href="https://www.theguardian.com/technology/ng-interactive/2026/jul/12/software-developers-engineers-ai" rel="noopener noreferrer" target="_blank">most affected by AI</a>—are now facing a significant increase in code review. “If you ask 20 developers, you get 23 different ways of working with it. Everyone is trying to sort it out,” says Butler, who describes this period as “the uncomfortable middle.” </p><p class="pull-quote">“We tell engineers that they need to be adaptable when they graduate, but we don’t actually explain what that means, demonstrate what that looks like, [or] help make sure that they’re developing it.”<span><strong>– Samantha Brunhaver, Arizona State University</strong></span></p><p><span></span>Brunhaver says one way educators can help prepare students before they enter the workforce is by offering a diversity of real-world experiences, such as internships, team-based projects, community service, and leadership roles. Each of these teach students to adapt to different challenges, easing their transition from school to work. </p><p> It’s also important to encourage reflection, Brunhaver adds, noting that metacognition helps individuals use the skill more effectively. “In order to adapt, you have to think that you have agency and the ability to get through a situation.” Ultimately, it comes down to three steps: Perceive a need to adapt, evaluate your options, and act. </p><p> For those already in the workforce, that action may mean taking the time to learn new tools and ways of working. Software engineering, for instance, may soon rely more on <a href="https://spectrum.ieee.org/chain-of-thought-prompting" target="_self">prompting models</a> and managing agents than coding line by line. “I think people who went into software because they like solving problems are going to have a lot of fun, and people who just enjoy the art of writing code are not,” Butler says. </p><h2>The More Things Change…</h2><p> Although the tools engineers use on a daily basis are evolving, the core responsibilities of the job are more stable than they may seem, says <a href="https://toolshed.com/about.html" target="_blank">Andy Hunt</a>, a software developer who coauthored <a href="https://pragprog.com/titles/tpp20/the-pragmatic-programmer-20th-anniversary-edition/" target="_blank"><em><em>The Pragmatic Programmer</em></em></a> (Addison-Wesley Professional) in 1999. The book outlines practical coding principles, and has been taught in many computer science classrooms. When Hunt was working on the 20th anniversary edition of the book, he was surprised by how much of the advice still applies. And now, seven years later, he maintains that belief.</p><p>“The fundamental part of the job is problem solving and communication, and that’s always going to be there,” he says. </p><p> Hunt emphasizes the importance of developing systems thinking over particular tools. To him, identifying as a Java programmer, for instance, is “like a carpenter saying, ‘I’m a hammer user,’ or ‘I specialize in cordless drills.’ ” </p><p> He acknowledges that today’s hiring process, in which companies often <a href="https://mckelveyconnect.washu.edu/blog/2022/02/04/8-things-you-need-to-know-about-applicant-tracking-systems/" rel="noopener noreferrer" target="_blank">filter résumés</a> for certain languages or years of experience, makes it harder to embrace a more expansive way of relating to your job. Employers, he says, should recognize that “the tech’s not the hard part, and it never has been. Understanding information theory, understanding systems thinking, understanding what constraints you’re up to—that’s still the hard part.” </p><p> With this type of misalignment between employers and employees, AI is also intensifying an old source of tension: How can engineers slow down enough to adapt and learn new tools when the pressure to become more productive keeps mounting? </p><h2>Who’s Responsible for Enabling Change? </h2><p><span>Young engineers need to embrace change. However, educators and employers also play a role in building a successful workforce. From the educator’s perspective, Brunhaver says “we need to be more explicit about what [adaptability] means and why it’s important.” Managers, meanwhile, should invest in their employees’ professional development.</span></p><p>Microsoft research scientist Butler often encourages leadership to set aside intentional time for continuous learning for their engineers—even just an hour a week—without any expectation that they will produce code or progress in their daily work. “I realize that’s difficult,” says Butler. “I would encourage people to do it on their own, but I would really encourage organizations and leaders to do it, because you’re not going to get this sudden change in your people if they don’t have time and space to learn how to work differently.” </p><p>This also means providing enough instruction, Butler adds. When developers aren’t given enough guidance on adopting something new, while being pressured to increase productivity, they risk <a href="https://arxiv.org/pdf/2507.21280" target="_blank">doubling down on the tools they already know and burning out</a>. </p><p>“I do imagine the next number of years could be challenging,” Butler says. Engineers will have to adapt to find their place in an evolving workforce—but they also have a say in shaping that future. </p><p>“Being adaptable sort of implies that you’re going to change based on what’s happening around you, and I would really like people to realize the change that’s happening is somewhat up to us,” she says. All individuals have a choice in how they use AI, for instance, and which models they use. “We need to be adaptable and go with the flow to a degree, but we also need to be directing that flow. The future with AI is absolutely not predetermined.”</p><p><em>This article appears in the September 2026 print issue as “The Adaptable Engineer.”</em></p>]]></description><pubDate>Mon, 24 Aug 2026 14:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/adaptable-engineer-core-skills</guid><category>Adaptive-learning</category><category>Professional-development</category><category>Ai-tools</category><category>Future-of-work</category><category>Type-departments</category><dc:creator>Gwendolyn Rak</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/silhouette-of-person-at-computer-atop-a-weather-vane-labeled-ai-against-sky-background.png?id=67652283&amp;width=980"></media:content></item><item><title>This IEEE Senior Member Develops AI Tools for E-Commerce Sites</title><link>https://spectrum.ieee.org/ieee-senior-member-ai-ecommerce</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-professionally-dressed-indian-man-smiling-with-a-trophy-in-one-hand-and-an-award-certificate-in-the-other.jpg?id=67657769&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p><a href="https://balajiingole.com/" rel="noopener noreferrer" target="_blank">Balaji Ingole</a> rarely saw televisions while growing up in Udgir, India. No one in the small Maharashtra village had computers or phones. Only one household owned a television, and neighbors often gathered there to watch shows together.</p><p>Ingole never even saw a computer growing up. It wasn’t until he reached middle school that he encountered a computer lab, an experience he says changed his life. Almost immediately, he says, the machine felt like a window into a different scale of possibility for him.</p><h3>Balaji Ingole</h3><br/><p><strong>Employer </strong></p><p><strong></strong>Amla Commerce in Milwaukee</p><p><strong>Title </strong></p><p><strong></strong>Project manager</p><p><strong>Member grade</strong> </p><p>Senior member</p><p><strong>Alma maters </strong></p><p><strong></strong>COEP Technological University and Welingkar Institute of Management, both in India</p><p>“I was very studious and not very social, always reading or solving problems in a math textbook,” he says. “At the computer lab, I began learning the <a href="https://www.c-language.org/" rel="noopener noreferrer" target="_blank">C programming language</a>—which was like discovering a whole new world. I was fascinated that you could create something with just a few lines of code.”</p><p>His early interest grew into a self-directed education. Outside of Ingole’s formal classwork, he taught himself to build database-backed applications, wire up hardware, write software, and trace error logs.</p><p>Today the IEEE senior member similarly splits his time. During the week, he’s a project manager in Milwaukee at B2B e-commerce company <a href="https://www.amla.io/" rel="noopener noreferrer" target="_blank">Amla</a>, leading AI-driven digital transformation initiatives to help the company’s clients boost their sales. On weekends, he leads a similarly demanding life as an independent researcher. His current projects include developing AI-enabled health care diagnostic tools and assistive technologies to support people with physical disabilities.</p><p>“I believe in ‘learn by doing,’” he says. “I really like to test my knowledge and prototype ideas to find out if they truly work.”</p><h2>A college project becomes an inspiration</h2><p>Ingole’s tendency to go beyond his coursework continued after he graduated high school in 2004. As a mechanical engineering undergraduate at <a href="https://www.coeptech.ac.in/" rel="noopener noreferrer" target="_blank">The College of Engineering, Pune (now COEP Technological University)</a>, in India, he participated in several extracurricular activities. One was interviewing entrepreneurs and writing about them for <em><em>The COEP College Magazine</em></em>. The experience helped him gain confidence, he says, giving him the push he needed to pursue interviews for the publication with two Indian entrepreneurs he admired: <a href="https://www.infosys.com/about/management-profiles/narayana-murthy.html" rel="noopener noreferrer" target="_blank">N.R. Narayana Murthy</a>, cofounder of IT giant <a href="https://www.infosys.com/" rel="noopener noreferrer" target="_blank">Infosys</a>; and his wife, philanthropist <a href="https://sudhamurty.in/" rel="noopener noreferrer" target="_blank">Sudha Murty</a>. The Murtys cofounded the <a href="https://www.infosys.org/infosys-foundation.html" rel="noopener noreferrer" target="_blank">Infosys Foundation</a>, a nonprofit that runs educational, health care, women’s empowerment, and sustainability <a href="https://www.infosys.org/infosys-foundation/initiatives.html" rel="noopener noreferrer" target="_blank">programs</a> in underserved areas of India.</p><p>“Every week I would fax them: ‘Please give me an interview time,’” Ingole says. Eventually, Sudha Murty’s office offered him a phone interview, but he requested to meet her in person at Infosys’s Bengaluru offices. She agreed, but the offices were 940 kilometers from Pune, and he didn’t have the money to travel or stay overnight in a hotel.</p><p>Ingole and a classmate borrowed money from friends and traveled through the night on multiple buses and trains to get to Bengaluru. They freshened up in a public bathroom before heading to the Infosys campus to meet Murty.</p><p>Impressed by their persistence, she surprised them by also arranging a brief chat with Narayana Murthy.</p><p>“Narayana Murthy handwrote a personal message to the engineering students of [my college]—which we proudly published in our college magazine,” Ingole says. “In his note, Murthy shared that we are at an extraordinary moment in India’s history and that the future looks even brighter. His words encouraged us to work hard and make the most of this time.</p><p>“I still have that note,” Ingole says. “They are billionaires, and I was just a regular student. The fact that they took the time to do this really motivated me.”</p><p>During the final semester of his engineering studies, Ingole joined the <a href="https://in.linkedin.com/company/tata-research-development-and-design-centre-trddc" rel="noopener noreferrer" target="_blank">Tata Research Design and Development Center</a> in Pune for a six-month internship. After earning his bachelor’s degree in mechanical engineering in 2008, he became a graduate engineering trainee at <a href="https://automation.honeywell.com/us/en" rel="noopener noreferrer" target="_blank">Honeywell Automation</a> in Pune.</p><p>He left the company in 2009, and during the next 13 years, he held different software engineering and project management positions at IT companies across India.</p><p>He earned a master’s degree in business administration from the <a href="https://www.welingkar.org/" rel="noopener noreferrer" target="_blank">Welingkar Institute of Management, Mumbai,</a> in 2017.</p><p>In 2022 he accepted a project-manager role at <a href="https://marssg.com/" rel="noopener noreferrer" target="_blank">Mars IT Solutions in Madison, Wisc.</a> The following year, he left to join <a href="https://www.gainwelltechnologies.com/" rel="noopener noreferrer" target="_blank">Gainwell Technologies</a>, also in Madison, as a senior project manager. At Gainwell, he managed projects for the core IT systems multiple U.S. state health departments use to administer <a href="https://www.medicaid.gov/" rel="noopener noreferrer" target="_blank">Medicaid</a> benefits, manage provider enrollment, and verify member eligibility. The experience managing projects that directly enabled patients’ access to health care gave Ingole a special appreciation for and interest in this area, he says.</p><p>“Health care data is not like other data,” he says. “The stakes are high, compliance requirements are different, and the margin for error is effectively zero.”</p><p>Ingole says he enjoyed the rigor of data governance combined with the potential to positively impact lives, and that also applies to his current work at Amla.</p><h2>Agentic AI in e-commerce</h2><p>Ingole joined Amla in July 2025. He helps manufacturers and B2B customers modernize their <a data-linked-post="2650248208" href="https://spectrum.ieee.org/who-invented-ecommerce" target="_blank">e-commerce</a> operations. He also builds AI tools for them and for his internal team.</p><p>For Amla’s customers, he’s developing AI-enabled chatbots that help manufacturers set up and manage large product catalogs in e‑commerce platforms. Such product setup traditionally has been a manual, tedious, error-prone process: Companies upload thousands of products, adjust item names, enter prices, update images, and more.</p><p>“Product setup has been one of the most painful processes in e-commerce, and it can take [our] customers two to three months to complete,” Ingole says. “We’re creating an AI agent that will guide them, step-by-step, to get everything set up in two weeks.”</p><p>Ingole relies on AI agents for some of his own tasks at Amla. Project managers historically have spent 10 to 12 hours each week assembling and sending status reports to stakeholders. Ingole built an AI agent to handle much of the work.</p><p>“It runs every Monday morning and reads through my emails to extract highlights, risks, timelines, and upcoming releases, then sends me a written status report,” Ingole says. The process might sound simple, but the agent’s workflow involves at least a dozen steps including defining parameters, managing temporary files, and integrating with existing tools.</p><p>With the information-gathering work handled, it frees up Ingole and his colleagues to spend more time on deeper-thinking work, he says.</p><h2>Publishing as idea refinery</h2><p>For nearly a decade, Ingole has spent some of his free time conducting independent research projects in data analytics and AI-enabled applications in health care. He has written <a href="https://ieeexplore.ieee.org/author/790405442038687" rel="noopener noreferrer" target="_blank">more than 40 peer-reviewed papers</a>, which are in the <a href="https://spectrum.ieee.org/free-access-to-thousands-of-covid19-research-documents" target="_self">IEEE Xplore Digital Library</a>. He has been granted six patents in the United Kingdom and India. In the U.K., he is a registered coinventor of an AI-powered, cloud-connected <a href="https://www.researchgate.net/publication/388234716_AI-POWERED_CLOUD-CONNECTED_WEARABLE_DEVICE_FOR_PERSONALIZED_HEALTH_MONITORING" rel="noopener noreferrer" target="_blank">wearable device for health monitoring</a> and an <a href="https://www.researchgate.net/publication/389874948_AI-BASED_BREAST_CANCER_DETECTION_DEVICE" rel="noopener noreferrer" target="_blank">AI-based breast cancer detection tool</a>.</p><p>Ingole’s patent for the breast cancer detector, he says, reflects his belief that when engineers apply data and AI correctly, they can help doctors diagnose patients more quickly and accurately.</p><p>That, he says, is both a power and a responsibility.</p><p>He is part of a team helping patients who are paralyzed and nonverbal control items in their environment. His goal, he says, is to develop a brain-computer interface to let patients turn on a fan, switch off a television, and complete similar tasks.</p><p>Publishing research requires both academic rigor and peer scrutiny, and Ingole says the function has been critical to improving as both a project manager and a researcher-inventor.</p><p>“Lots of research ideas never make it to paper,” he notes. “But when you write for journals or conferences, you’re bombarded with questions from Ph.D.s and experienced researchers. This forces me to refine my methodology, and to combine use cases and technical architecture in a way that stands up to expert review.”</p><h2>Finding a professional hub</h2><p>Ingole joined IEEE in 2022, and he says the affiliation has become central to both his research and his professional identity.</p><p>“I use the <a href="https://cis.ieee.org/activities/membership-activities/ieee-member-directory" rel="noopener noreferrer" target="_blank">Member Directory</a> often and contact engineers through my IEEE email address, which gives me credibility because they know it’s a genuine research connection,” he says.</p><p>The organization has given him a platform to contribute to the research space beyond his own papers, he says. He has served as a conference session chair, keynote speaker, technical program committee member, and peer research reviewer for various conferences and events. His IEEE membership, he says, has opened doors to other communities, helping support his entry into the <a href="https://www.bcs.org/" rel="noopener noreferrer" target="_blank">British Computer Society</a>, which has stringent acceptance criteria.</p><p>Those opportunities have helped him build a global network of collaborators with whom to discuss upcoming research, seek advice, and share data, he says.</p><p>“IEEE is important for me to continue as an independent researcher,” he says. “It lets me contribute to the community, and I get a lot in return.”</p>]]></description><pubDate>Fri, 21 Aug 2026 18:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ieee-senior-member-ai-ecommerce</guid><category>Ieee-member-news</category><category>Artificial-intelligence</category><category>E-commerce</category><category>Careers</category><category>Type-ti</category><dc:creator>Julianne Pepitone</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-professionally-dressed-indian-man-smiling-with-a-trophy-in-one-hand-and-an-award-certificate-in-the-other.jpg?id=67657769&amp;width=980"></media:content></item><item><title>Stop Hunting, Start Solving: Accelerating Root Cause Analysis with Agentic AI</title><link>https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&amp;utm_medium=site&amp;utm_campaign=922</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/spotfire-logo-with-circular-icon-and-stylized-black-text.png?id=67657308&width=980"/><br/><br/><p><strong>About this Webinar</strong></p><p><strong>Turn Yield Excursions into Faster, More Confident Root Cause Analysis</strong></p><p>When a yield issue emerges, the answer rarely lives in a single system. Critical clues are spread across metrology data, tool traces, chemical analysis, and facilities systems, while growing data volumes make traditional dashboards slow, fragmented, and difficult to act on.</p><p><strong>What You’ll Learn:</strong></p><p>Discover how a purpose-built semiconductor analytics platform can help engineers connect insights across domains without moving data. See how Agentic AI, semiconductor-specific visualizations, and push-down compute enable faster investigation of yield excursions and process issues, even across billions of data points. In the session, a live demonstration shows how to conduct a multi-domain root cause investigation using Spotfire® Industry Pro.</p><p><strong>Key Takeaways:</strong></p><ul><li>Understand why siloed manufacturing data delays yield recovery and inflates costs</li><li>Learn how Agentic AI automates complex cross-domain analytics and visualization generation</li><li>Explore methods for scaling high-performance analytics across massive fab datasets</li></ul><p><strong>Who Should Attend:</strong></p><p><span>Yield, Process, and Integration Engineers; Fab and Manufacturing Operations Managers; Quality and Reliability Engineers; and Data & Analytics leaders supporting wafer fabs, foundries, OSATs, and IDMs who need to identify issues faster while maintaining confidence in decision-making.</span></p><p><strong>Save Your Spot!</strong></p><p><span></span><span>Join this webinar to learn how leading semiconductor teams are accelerating root cause investigations, scaling analytics across massive datasets, and transforming disconnected data into actionable manufacturing intelligence. Reserve your seat today.</span></p><p><span><span><a href="https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&utm_medium=site&utm_campaign=922" target="_blank">Register now for this free webinar!</a></span></span></p>]]></description><pubDate>Fri, 21 Aug 2026 14:32:37 +0000</pubDate><guid>https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&amp;utm_medium=site&amp;utm_campaign=922</guid><category>Type-webinar</category><category>Semiconductor-manufacturing</category><category>Agentic-ai</category><category>Root-cause-analysis</category><category>Yield-analytics</category><category>Fab-operations</category><dc:creator>Spotfire</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67657308/origin.png"></media:content></item><item><title>From AI Copilots to Agent Swarms</title><link>https://spectrum.ieee.org/amd-agent-swarms</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/colorful-3d-blocks-piled-behind-glass-panels-displaying-white-code-snippets.jpg?id=67609515&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p><span>The impact of AI on software development has been both profound and ever-evolving. Last year, I </span><a href="https://spectrum.ieee.org/beyond-code-autocomplete" target="_self">wrote</a><span> about <a href="https://www.amd.com/en.html" target="_blank">AMD’s</a> plans to use AI not just for </span><a href="https://spectrum.ieee.org/best-ai-coding-tools" target="_self">generating</a><span> new lines of code, but also for other steps in the software development lifecycle (SDLC), such as triaging problems, debugging code, and testing the software. At the time, we were hoping for a 25 percent productivity boost from AI use over the course of two or three years.</span></p><p>But with each new release, the capabilities of large language models (LLMs) improve dramatically—accelerating software development, increasing the quality of AI-generated code, and fundamentally reshaping how software is engineered. Now, just one year later, we have surpassed our productivity target, achieving a 30 percent overall productivity boost through AI. On top of that, we are rethinking not only how we use AI within the SDLC, but the structure of the SDLC itself.</p><p>We believe that the biggest AI revolution in software engineering is still ahead. So far, we have largely been teaching AI how we perform tasks and asking it to mimic existing workflows. In many ways, this constrains AI to human patterns of thinking. The next transformation will come from collaborative swarms of AI agents capable of discovering solutions independently.</p><h2>Agents of today</h2><p>AMD began developing AI systems for code generation, testing automation, bug analysis, and code review in 2024. At the time, our objective was to achieve 25 percent AI-generated production code by 2027 while gradually automating larger portions of the SDLC.</p><p>Measuring productivity is inherently <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" target="_blank">challenging</a>, but from the outset we have consistently tracked one objective metric: the percentage of source code generated by AI. Importantly, we count only code that passes all reviews and testing and is ultimately included in the final product. While AI-generated code is certainly not the only contributor to productivity gains, it is one of the few metrics that can be measured objectively and consistently.</p><p>By this metric, we have crossed the 20 percent mark at the beginning of this year and are now progressing towards 50 percent across entire codebase. In some software components, more than 80 percent of the code is now generated using AI.</p><p>Agentic AI has enabled us to include AI in every step of the life cycle: For <strong>code analysis and triage</strong>, agents are trained to analyze problem reports, identify and group similar requests, and highlight which code snippets are likely to need modification. For <strong>debugging and code generation</strong>, agents are directed to analyze a bug request and implement required code changes. For <strong>testing</strong>, the agents generate unit tests, and if those are passed, identify necessary integration and product-level tests. And finally, for the <strong>approval and release</strong> stage, agents prepare architecture summary, code change review, and full test results for engineers’ review and approval—and, if approved, integrate the changes into the next release.</p><h2>Agents of tomorrow</h2><p>Today, engineers create AI agents in their own image: They teach AI what they know about the system, how they would fix an issue, and how they would implement a change. This is already a major technological advancement. Engineers can create multiple “AI versions” of themselves, allowing these agents to work in parallel, scaling their expertise far beyond the limits of individual productivity. The limitation, however, is that these AI agents are still constrained by human thinking and human-defined approaches.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Person typing on a keyboard. Looming in front of them is a network of colorful AI agents surrounding a glowing central node." class="rm-shortcode" data-rm-shortcode-id="0d892f7d0215c3159d9e41db8419441a" data-rm-shortcode-name="rebelmouse-image" id="73c7c" loading="lazy" src="https://spectrum.ieee.org/media-library/person-typing-on-a-keyboard-looming-in-front-of-them-is-a-network-of-colorful-ai-agents-surrounding-a-glowing-central-node.jpg?id=67609528&width=980"/> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">AMD</small></p><p>We believe the next major transformation in software engineering will occur when collaborative AI agent swarms can independently identify and develop solutions, guided by humans on what to solve rather than constrained by human assumptions about how the job should be done. Instead of providing detailed instructions on how to solve a problem, engineers will define the issue, the desired outcome, and the quality, performance, and system constraints, allowing AI agents to determine the optimal path to a solution.</p><p>A swarm of AI agents will then work in parallel to generate, evaluate, and refine multiple solution approaches. These agents will automatically validate correctness, measure performance, test trade-offs, and compare alternative implementations against defined success criteria. Finally, AI agents will prepare ranked solution options, along with validation results and performance metrics, for engineer review and approval. The agents won’t be enhancing each step of the SDLC—they will be rewriting the SDLC themselves.</p><p>To get to this point, we need to change how agents are trained. Today, improvement occurs one engineer and one agent at a time: An engineer reviews the output, refines the prompt, and repeats the process. To scale beyond this model, agents must continuously learn from one another, reuse successful strategies, and improve collaboratively across projects and teams.</p><p>We are already moving in this direction by using multi-agent workflows extensively through agentic harnesses, such as Codex and Claude Code, while simultaneously developing our own internal multi-agent systems to support the next generation of AI-driven software engineering.</p><p>A good example is our AI-driven effort to resolve issues in our Radeon Software eXperience (RSX). RSX is a user interface component that allows users to configure and monitor graphics driver behavior. In October 2025, we began using AI agents to automatically debug and fix reported RSX issues. Out-of-the-box AI tools delivered limited results, resolving only 6 percent of issues.</p> <p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A bar chart labelled RSX-Agentic Resolution Rate shows large percentage increases from 6 in October 2025 to 75 in June 2026." class="rm-shortcode" data-rm-shortcode-id="98d6334ebf612596d6de9b5137c39de6" data-rm-shortcode-name="rebelmouse-image" id="d27a2" loading="lazy" src="https://spectrum.ieee.org/media-library/a-bar-chart-labelled-rsx-agentic-resolution-rate-shows-large-percentage-increases-from-6-in-october-2025-to-75-in-june-2026.jpg?id=67609523&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The percentage of software issues fixed automatically by AI agents in AMD’s Radeon Software eXperience (RSX) has been growing steadily, reaching 75 percent in June 2026. </small></p> <p class="media-body"><span>As we analyzed failures and identified ways to improve, we built a learning loop—initially a largely manual process—to understand where the agents were falling short and how to improve them. Rather than retraining the underlying models, we refined the objectives given to the agents, allowing them to iteratively explore multiple approaches, evaluate the results against defined success criteria, and converge on better solutions. At the same time, advances in models and agent run-times further increased effectiveness. Together, these improvements significantly increased our resolution rate from 6 percent to more than 75 percent of RSX issues resolved by agentic loop.</span></p><p>To make agents and agent swarms truly productive, we need a continuous learning loop that feeds errors and human interventions back into future agent workflows. The opportunity is to engineer this loop around clear, measurable goals. Each cycle captures new insights, making the entire AI engineering workflow smarter and more effective. Over time, this self-reinforcing loop—not just the underlying model—will become a key driver of AI progress.</p><h2>The evolving role of human engineers</h2><p>At AMD, we view AI as a means of increasing productivity, improving quality, and enabling employees to focus on higher-value work. Our goal is to empower our workforce with AI, not to reduce headcount.</p><p>To support this transformation, we are investing heavily in AI education and training across the company. The way we work is evolving rapidly, and we want every AMD employee to be prepared to leverage AI confidently, responsibly, and effectively.</p><p>As AI agents continue to improve, engineers will spend less time manually implementing solutions, focusing more on defining specifications, validating outcomes, and making the strategic decisions that drive innovation.</p>]]></description><pubDate>Mon, 17 Aug 2026 14:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/amd-agent-swarms</guid><category>Agentic-ai</category><category>Software-engineering</category><category>Amd</category><category>Large-language-models</category><dc:creator>Andrej Zdravkovic</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/colorful-3d-blocks-piled-behind-glass-panels-displaying-white-code-snippets.jpg?id=67609515&amp;width=980"></media:content></item><item><title>AI Used to Verify Toughest Mathematics Proof Yet</title><link>https://spectrum.ieee.org/axiom-math-246-theorem-formalization</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/abstract-illustration-of-several-long-arrows-stacked-horizontally-parallel-to-one-another-each-arrow-has-one-plotted-point-in-a.jpg?id=67608532&width=1200&height=800&coordinates=62%2C0%2C63%2C0"/><br/><br/><p>Representing a significant milestone in AI-assisted mathematical research, a team at <a href="https://axiommath.ai/" rel="noopener noreferrer" target="_blank">Axiom Math</a> has automatically verified the proof of a theorem relating to prime numbers—colloquially referred to as the “246 theorem”—for the first time using the company’s AI system AxiomProver.</p><p>In formal verification, mathematicians task a computer with checking a machine-readable version of a proof. The process is not a 100 percent guarantee that the proof is correct, <a href="https://gigazine.net/gsc_news/en/20260803-collatz-lean-kernel-bug/#gsc.tab=0" rel="noopener noreferrer" target="_blank">as a recent demonstration showed</a>, exposing how a bug in the method could be exploited to accept a false, AI-generated proof. Still, the computational method is as close to a rubber stamp as you can get. </p><p>This particular verification formalizes an important advance in number theory. Beyond this particular proof, it demonstrates how automated AI verification could be used in the future to ensure the correctness of AI-generated computer code that will soon underlie software across the globe.</p><h2>Useful formalization by design</h2><p>This is not AxiomProver’s first rodeo. Axiom Math has used its autonomous, multi-agent system that turns mathematical statements into machine-checkable proofs to crack several unsolved mathematical problems and verify many more proofs <a href="https://axiommath.ai/selected-publications" rel="noopener noreferrer" target="_blank">this year</a>. But proof formalization of the 246 theorem is by far the most significant, as <a href="https://www.linkedin.com/in/ken-ono-a972191a5/" rel="noopener noreferrer" target="_blank">Ken Ono</a>, Axiom Math’s founding mathematician, explains: “This theorem currently represents the threshold of human knowledge about prime numbers.”</p><p>Earlier this year, Axiom Math competitor <a href="https://www.math.inc/" rel="noopener noreferrer" target="_blank">Math, Inc.</a> used its Gauss agent to verify <a href="https://people.epfl.ch/maryna.viazovska?lang=en" rel="noopener noreferrer" target="_blank">Maryna Viazovska</a>’s 2022 Fields Medal-winning proof of the sphere-packing problem in 8 and 24 dimensions. <a href="https://thefundamentaltheor3m.github.io/" rel="noopener noreferrer" target="_blank">Sidharth Hariharan</a>, a Ph.D. student at Carnegie Mellon University who led human efforts that were critical in the Math, Inc. breakthrough, says that Axiom Math’s AI approach to formalizing the 246 theorem is more comprehensive and useful. Hariharan’s group continues to work toward fully formalizing Viazovska’s proof.</p><p class="ieee-inbody-related">RELATED: <a href="https://spectrum.ieee.org/ai-proof-verification" target="_blank">Watershed Moment for AI-Human Collaboration in Math</a></p><p>Now an intern at Axiom Math, Hariharan has been heavily involved in the company’s formalization of the 246 theorem proof. He says that one of the main differences here is that rather than it being a one-shot approach relating to a single problem, Axiom Math has expressly aimed to make components of the formalization reusable for other formalization tasks and mathematical research. The team has wielded AxiomProver to build a <a href="https://github.com/AxiomMath/PrimeGapsLib" rel="noopener noreferrer" target="_blank">library of results about gaps in primes</a>. The 246 theorem is the flagship result within that library. </p><h2>What is the 246 theorem?</h2><p>The first few primes are close together: 2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, .... And there are several instances where they are separated by a difference of two: 3:5, 5:7, 11:13, 17:19, ....</p><p>These pairs of primes are called twin primes. Twin primes become rarer the further you get from zero, but they do still seem to pop up occasionally. The <a href="https://mathworld.wolfram.com/TwinPrimeConjecture.html" rel="noopener noreferrer" target="_blank">twin prime conjecture</a>, first precisely formulated in the 19th century by French mathematician Alphonse de Polignac, posits that they will keep popping up regardless of how far along the number line you look. In other words, there are infinitely many twin primes.</p><p>Though easy to state, the venerable twin prime conjecture remains unproven. First progress toward solving it only occurred in 2013 when Yitang Zhang, now a professor at Sun Yat-sen University, in Guangzhou, China, <a href="https://annals.math.princeton.edu/2014/179-3/p07" rel="noopener noreferrer" target="_blank">proved that there are infinitely many pairs of primes</a> that are separated by 70 million. A few months later, using a different technique, University of Oxford professor <a href="https://www.sjc.ox.ac.uk/discover/people/professor-james-maynard/" rel="noopener noreferrer" target="_blank">James Maynard</a> dramatically reduced this gap <a href="https://arxiv.org/abs/1311.4600" rel="noopener noreferrer" target="_blank">from 70 million to just 600</a>; a feat which substantially contributed to Maynard being awarded the <a href="https://www.mathunion.org/imu-awards/fields-medal/fields-medals-2022" rel="noopener noreferrer" target="_blank">2022 Fields Medal</a>—widely regarded as the <a href="https://spectrum.ieee.org/tag/nobel-prize" target="_self">Nobel Prize</a> for mathematics. </p><p>As part of a group of mathematicians known as the Polymath8b collaboration, Maynard and fellow Fields Medalist <a href="https://mathstodon.xyz/@tao" target="_blank">Terence Tao</a>, professor at the University of California, Los Angeles, brought the gap down to just 246, the closest mathematicians have gotten to the target gap of two. It is this 246 theorem—which states that there are infinitely many primes that differ by 246—that AxiomProver has verified to be correct. </p><h2>Safe and correct AI-generated code</h2><p>The techniques formalized in this work are important in number theory, the branch of mathematics that underpins all present-day cybersecurity and cryptography. They could therefore prove to be useful in verifying specific ways in which we keep our digital data safe in the future. </p><p>But Axiom Math’s Ono is more excited by the bigger picture. He sees formalizing mathematical proofs as a stepping stone to verifying AI-generated code, which is starting to be used across society in systems that run our infrastructure, manage our finances, and protect our data. This is despite safety concerns surrounding hallucinations, bugs, and other unintended vulnerabilities.</p><p>If properties of code—such as whether an algorithm terminates or if a program’s output is correct for any input—can be translated into precise mathematical statements, technologies derived from AxiomProver would be ideally suited to formally stating and proving them. In this way, mathematically verifying the correctness of AI-generated code would make this code safe to use.</p><p>“The world is about to run on computer code that nobody has read,” Ono concludes. “AI is here and we can no longer look away—proof formalization is a test bed for solving what I think is the most important challenge we will face from AI.”</p><p><em>This article was updated on 18 August to clarify the nature of Hariharan’s work formalizing Viasovska’s proof. </em><br/></p>]]></description><pubDate>Mon, 17 Aug 2026 13:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/axiom-math-246-theorem-formalization</guid><category>Mathematics</category><category>Prime-numbers</category><category>Ai-reasoning</category><category>Ai-generated-software</category><dc:creator>Benjamin Skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/abstract-illustration-of-several-long-arrows-stacked-horizontally-parallel-to-one-another-each-arrow-has-one-plotted-point-in-a.jpg?id=67608532&amp;width=980"></media:content></item><item><title>The CPU Comeback Is Upon Us</title><link>https://spectrum.ieee.org/ai-cpu-comeback</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-single-computer-chip-balancing-on-the-tip-of-a-pyramid.jpg?id=67615384&width=1200&height=800&coordinates=0%2C208%2C0%2C209"/><br/><br/><p>Earlier this year, leaders at <a href="https://aws.amazon.com/free/?trk=dc9b9d60-cc82-4cd5-8a61-0b33d6a79fab&sc_channel=ps&ef_id=CjwKCAjw1vXTBhB-EiwAEKr_k_57Xoz6QKQSrRqDF4mGRYOudA_A99MTHZL6alVjPDxdliBQfEMPihoC0RUQAvD_BwE&gads_camp=23532472510&gads_ag=199502799824&gads_ad=795877020713&gads_kw=amazon%20web%20services&gads_matchtype=e&gads_network=g&gads_device=c&gads_geo=9198314&gad_campaignid=23532472510&gbraid=0AAAAADjHtp-SgwVQvE7H9V8lK49jVtMDw&gclid=CjwKCAjw1vXTBhB-EiwAEKr_k_57Xoz6QKQSrRqDF4mGRYOudA_A99MTHZL6alVjPDxdliBQfEMPihoC0RUQAvD_BwE" rel="noopener noreferrer" target="_blank">Amazon Web Services</a> delivered a new mandate to their engineers: They need to conserve CPU cycles at all costs. AWS has <a href="https://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunch" rel="noopener noreferrer" target="_blank">reportedly</a> experienced an explosion in wait times for CPU server capacity as AI workloads strain the company’s cloud infrastructure. </p><p>The issue seemingly took AWS off guard, and for good reason. The AI boom led to a surge in demand for <a href="https://spectrum.ieee.org/nvidia-gpu" target="_self">GPUs</a> and, later, <a href="https://spectrum.ieee.org/high-bandwidth-memory-shortage" target="_self">memory</a>. CPUs were mostly left out of the story, as their relative lack of parallelization made them a poor fit for AI model inference, the process of running and serving large language models (LLM) to users. </p><p>But the rise of agentic AI systems, which allow AI models to operate autonomously and call on sub-agents, is changing the narrative.</p><p><a href="https://moorinsightsstrategy.com/team/matt-kimball/" rel="noopener noreferrer" target="_blank">Matt Kimball</a>, vice president and principal data center analyst at <a href="https://moorinsightsstrategy.com/" rel="noopener noreferrer" target="_blank">Moor Insights & Strategy</a> in Austin, Texas, says 2026 has brought a spike in CPU demand, much of it due to <a href="https://spectrum.ieee.org/ai-agents" target="_self">agentic AI</a>. “It’s one thing to have this agentic workload, and let’s say it spawns 100 agents. If I’m going to roll this out across my enterprise, those 100 become tens of thousands, hundreds of thousands, or millions of agents,” Kimball says. “You have agents spawning sub-agents, making API [application programming interface] calls and talking to more agents through [Anthropic’s] model context protocol.”</p><h2>AI agents need to use computers, and computers need CPUs</h2><p>Kimball’s comments refer in part to “tool use,” which is shorthand for an LLM’s ability to access the internet, open files on a desktop, and generally use a variety of software to accomplish its task. </p><p>LLMs trained for tool use learn how to call on other software. While the LLM’s inference is still primarily executed on a GPU or similar AI accelerator, the tool calls that the LLM makes are typically pushed to the CPU.</p><p>“Many components of an agentic AI task are inherently CPU based jobs,” explains <a href="https://www.linkedin.com/in/souvik-kundu-64922b50/" rel="noopener noreferrer" target="_blank">Souvik Kundu</a>, senior staff research scientist at <a href="https://www.intel.com/content/www/us/en/homepage.html" rel="noopener noreferrer" target="_blank">Intel</a>. “The CPU does the job of parsing output, figuring out which tool to invoke, making the API call or running the code, collecting the result, and feeding it back.” <a href="https://www.linkedin.com/in/mrangarajan/" rel="noopener noreferrer" target="_blank">Madhu Rangarajan</a>, vice president of compute and enterprise AI products at <a href="https://www.amd.com/en.html" rel="noopener noreferrer" target="_blank">AMD</a>, makes a similar claim, saying, “In our testing, seven of the eight stages in realistic agentic AI pipelines run entirely on the CPU.”</p><p>An LLM tasked with programming software, for example, will likely make tool calls to write code to files, move or replace files, download required packages, and build the software once the LLM believes it’s complete. </p><p>Kundu co-authored a <a href="https://arxiv.org/pdf/2511.00739" rel="noopener noreferrer" target="_blank">paper</a> on agentic AI optimization alongside researchers from Georgia Tech in Atlanta. They found the CPU is often idle while LLM inference is executed on a GPU and that, conversely, the GPU is often idle when tool calls are executed on the CPU. To optimize this, Kundu and his colleagues propose scheduling optimizations that can cut end-to-end latency (the time between the start and finish of the agentic workload) by up to 1.8-times under sustained load. </p><p>It’s a start, but the gains chase a moving target. Agentic systems generate work at machine speed and multiply it as they go. OpenAI’s inadvertent <a href="https://spectrum.ieee.org/hugging-face-openai-cyberattack?itm_source=homepage&itm_medium=hero&itm_campaign=hero-2026-08-10&itm_content=hero6" target="_self">hack</a> of Hugging Face saw its model fire off as many as 300 actions an hour, and a single agent can spawn sub-agents that make tool calls of their own. </p><p>And there’s one more important complication that may increase the workload on a CPU as models become more complex: safety guardrails.</p><p>Safety and policy checks on an agent’s actions are often specific rules that inspect syntax and log files, Kundu says. Guardrails may also use small models (under a billion parameters) to analyze task complexity or intent. Though they could be executed on a GPU, they often aren’t, because their small size and the need to minimize latency keeps the work on the CPU.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A bar graph illustrating how the latency for Llama-8B\u2019s LLM becomes 4.45 times less when the number of CPU cores is increased from five to thirty-two." class="rm-shortcode" data-rm-shortcode-id="b43d668525bb89a3e691433307b18677" data-rm-shortcode-name="rebelmouse-image" id="7acb8" loading="lazy" src="https://spectrum.ieee.org/media-library/a-bar-graph-illustrating-how-the-latency-for-llama-8b-u2019s-llm-becomes-4-45-times-less-when-the-number-of-cpu-cores-is-increas.jpg?id=67615386&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Increasing the number of CPUs available significantly decreases the latency for Llama-8B responses over longer sequence lengths.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Source: <a href="https://arxiv.org/pdf/2603.22774" target="_blank">Euijun Chung, Yuxiao Jia, et al.</a></small></p><h2>Tokenization adds to bottlenecks</h2><p><a href="https://ejchung0406.github.io/" target="_blank">Euijun Chung</a>, a PhD student at Georgia Tech, recently co-authored another <a href="https://arxiv.org/abs/2603.22774" target="_blank">paper</a>, with findings that complement Kundu’s work. Chung and his co-authors found that when a server has too few CPU cores, it falls behind on dispatching work to the GPUs. That causes the GPUs to stall as they wait for instructions.</p><p>In addition to that, the paper touches on another key element of LLM workloads: <a href="https://seantrott.substack.com/p/tokenization-in-large-language-models" target="_blank">tokenization</a>.</p><p>Tokenization is a key first step in LLM inference. It converts text into integer token IDs that can be processed by the model. Unlike the matrix math required for most LLM inference, tokenization is branchy, data-dependent sequential string manipulation. Though it can be parallelized by chunking text, it’s not massively parallel in the same way as the bulk of LLM inference.</p><p>Tokenization of small prompts is a relatively trivial task and won’t tax even an entry-level CPU. However, an agentic model that makes tool calls must parse and tokenize the results of the call. </p><p>“If you have an ongoing sequence of, say, 100,000 tokens, and you have a tool result of 1,000 tokens, the tokenizer will have to tokenize the whole sequence again. And you have to do tokenization at every agentic tool call,” Chung says. This both increases the frequency of tokenization and increases the number of tokens involved. It’s conceivable that future tokenizers will find ways to mitigate this, Chung says, but it remains a problem for modern LLM inference.</p><p>The paper finds that time-to-first-token latency (the time required for the model to produce the first word of its reply) can increase dramatically as the sequence length grows. CPUs with more cores can reduce the problem. In test runs at longer sequence lengths, increasing CPU core counts can reduce time-to-first-token latency by roughly 1.5 to 7 times.</p><p>Chung and his colleagues were only able to test smaller models, such as Alibaba’s Qwen 3-30B and Meta’s Llama 3.1-70B, due to limitations of the hardware available for testing. He speculates that larger models will experience less dramatic bottlenecks due to their higher overall GPU demand, but also expects agentic AI will push token lengths far beyond the longest he and his co-authors tested.</p><p>“If you think about something like Anthropic’s Claude, you can easily hit 500,000, even a million tokens,” Chung says. “In the world of agentic AI, the average sequence length will grow and grow, so I’m expecting this problem to get worse in future workloads.” </p><h2>Is a CPU crunch just getting started?</h2><p>Amazon’s crackdown on use of CPU resources is one of several indicators that Kundu and Chung have identified issues with real-world relevance. </p><p>Intel has <a href="https://finance.yahoo.com/news/intel-turnaround-no-one-saw-141000146.html?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAAJqbhttNhPiUgZBGw-QWbgpaSpMnF0qwXSRBrSMQvCVWxhrkdpDYWyavDbXWb-9coPNWysaUbJK_l5uY_bZH6ACvfNpR8JrxKbKXsw8dVSlzZLYe9ydhDqS0SAyZbc4KSK0bqLw8WJSIYt-yoFGlGWDRP1UIaWov2q1bpzbcLPE-" rel="noopener noreferrer" target="_blank">sold out</a> of server CPUs through at least the end of the year. AMD has <a href="https://wccftech.com/amd-doubles-server-cpu-forecast-to-120-billion-as-agentic-ai-rewrites-demand-ceo-says-epyc-verano-built-purely-for-ai/" rel="noopener noreferrer" target="_blank">doubled</a> its server CPU forecast. <a href="https://www.arm.com/products/cloud-datacenter/arm-agi-cpu" rel="noopener noreferrer" target="_blank">Arm</a> and <a href="https://www.cnbc.com/2026/06/24/qualcomm-data-center-cpu-meta.html" rel="noopener noreferrer" target="_blank">Qualcomm</a> have both announced new CPUs designed to accelerate agentic AI. Even <a href="https://www.nvidia.com/en-us/" rel="noopener noreferrer" target="_blank">Nvidia</a> has prioritized <a href="https://nvidianews.nvidia.com/news/nvidia-unveils-vera-the-cpu-for-agents" rel="noopener noreferrer" target="_blank">Vera</a>, its Arm-based CPU for agentic AI, which is part of Nvidia’s <a href="https://spectrum.ieee.org/nvidia-rubin-networking" target="_self">Vera Rubin</a> platform.</p><p>Kimball says these developments make it clear that the AI industry is placing more emphasis on CPU performance. He sees the surge in demand as an “absolute tell” that CPUs are now considered a key part of an agentic AI system.</p><p>Unfortunately, this may translate to broader CPU shortages and increased prices, much as has already occurred with GPUs and memory. </p><p>“You’re already seeing a CPU crunch to some degree. When you look at the constraints in the market, it even trickles down into the consumer space,” Kimball says. He adds that Intel has <a href="https://www.techpowerup.com/345535/intel-reallocates-pc-production-capacity-to-server-cpus-amid-tight-wafer-supply" rel="noopener noreferrer" target="_blank">cut production</a> of client CPUs in favor of server CPUs, even as Intel’s new 18A production process has grown the company’s sales in the client segment. Kimball sees that as a sign that CPU makers will follow the money. </p>]]></description><pubDate>Sun, 16 Aug 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-cpu-comeback</guid><category>Cpu</category><category>Gpu</category><category>Agentic-ai</category><category>Llms</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-single-computer-chip-balancing-on-the-tip-of-a-pyramid.jpg?id=67615384&amp;width=980"></media:content></item><item><title>Inside the Data Bottleneck Slowing Visual and Physical AI</title><link>https://content.knowledgehub.wiley.com/the-2026-state-of-visual-and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/voxel51-logo-with-geometric-cube-icon-and-stylized-text.png?id=67607900&width=980"/><br/><br/><p>A survey of over 700 professionals examines how visual and physical AI teams build systems, why models fail, and where data work drives production.</p><p><span><a href="https://content.knowledgehub.wiley.com/the-2026-state-of-visual-%20and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/" target="_blank">Download this free whitepaper now!</a></span></p>]]></description><pubDate>Wed, 12 Aug 2026 14:18:05 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/the-2026-state-of-visual-and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/</guid><category>Type-whitepaper</category><category>Artificial-intelligence</category><category>Computer-models</category><category>Data-bottleneck</category><dc:creator>Voxel51</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67607900/origin.png"></media:content></item><item><title>Pakistani Judges Give Their Verdict on JudgeGPT</title><link>https://spectrum.ieee.org/judgegpt-experiment</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-a-translucent-gavel-against-a-background-of-legal-books.jpg?id=67600001&width=1200&height=800&coordinates=156%2C0%2C156%2C0"/><br/><br/><p>Judges around the world have <a href="https://www.bbc.com/news/articles/c178zzw780xo" rel="noopener noreferrer" target="_blank">made headlines</a> for <a href="https://www.judiciary.senate.gov/press/rep/releases/grassley-scrutinizes-federal-judges-apparent-ai-use-in-drafting-error-ridden-rulings" rel="noopener noreferrer" target="_blank">illicitly using generative AI</a> in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology—together with appropriate training–boosted the number of cases resolved by 6.3 percent with no obvious drop in the quality of judgments.</p><p>With a backlog of 2.26 million cases and fewer than two judges per 100,000 people—compared to 22 in the EU and eight in Brazil—Pakistan’s judiciary was in sore need of help. So, in consultation with the judiciary, <a href="https://sites.google.com/view/sultan-mehmood/home" rel="noopener noreferrer" target="_blank"> economist Sultan Mehmood</a>, of the New Economic School in Moscow, and collaborators tested whether AI could ease the burden.</p><p>They built a custom tool combining OpenAI’s GPT-4 large language model (LLM) with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes, to help judges with legal research and drafting judgments. They began offering the tool in 2024 to 1,559 trial judges—roughly half the country’s justices. </p><p>“We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,” Mehmood says. </p><h2>First of its kind</h2><p>“It’s pretty amazing that he’s able to pull this off,” says David Autor, an economics professor at MIT. “It’s not easy to do large-scale field experiments in civil service, but especially where the stakes are so high.” The 6.3 percent productivity boost is not overwhelming, he says, but it’s credible and likely to improve as the tool is more widely used.</p><p>AI tools for judges are already being rolled out in <a href="https://a2jlab.org/brazils-ai-driven-courts-innovation-without-evaluation/" rel="noopener noreferrer" target="_blank">Brazil</a> and <a href="https://iacajournal.org/articles/10.36745/ijca.647" rel="noopener noreferrer" target="_blank">India</a>, and prominent U.S. law professor Eric Posner has compared LLM judgments to human judgments in a single <a href="https://journals.sagepub.com/doi/10.1177/2755323X261433614" rel="noopener noreferrer" target="_blank">case study</a>. But until now, there has been no major independent assessment of ongoing judicial use of AI. The <a href="https://elliottash.com/papers/Mehmood-Goessmann-Ash-Courts-of-Tomorrow-Evidence-Nationwide-Rollout-Generative-AI.pdf" rel="noopener noreferrer" target="_blank">new study</a> focused on Pakistan’s trial courts; Mehmood says judges there were enthusiastic from the start.</p><p>“They were more techno-optimist than we were,” he says. “The delays are so huge, this is something which they thought was worth trying anyway to reduce people’s suffering.”</p><p>Some judges were also already using AI chatbots, Mehmood says, but commercial offerings performed poorly on Pakistani legal queries, frequently hallucinating case law. So the team built a tool tailored to the Pakistani context, called JudgeGPT.</p><p>They used retrieval-augmented generation (RAG), which allowed the model to query a database of 128,292 Pakistani judicial opinions and 943 statutes. Responses included footnotes linking to cases and laws.</p><p>“It turns out that actually the way to fix [hallucinations] isn’t just more intelligent models,” says study coauthor <a href="https://elliottash.com/" rel="noopener noreferrer" target="_blank">Elliott Ash</a>, an associate professor of law, economics, and data science at ETH Zurich in Switzerland. “It’s to attach the models to a tool that can do a search and verify the sources.” However, the researchers do not report hallucination rates.</p><p>The team also put 1,197 judges through six 90-minute Zoom training sessions, developed in collaboration with Pakistan’s <a href="https://www.fja.gov.pk">Federal Judicial Academy</a>, covering how LLMs work, their limitations, the risk of bias and hallucinations, and the importance of verifying outputs. Another 180 judges only underwent general training on technology in legal research, while a final group got no training.</p><p>By the time 487 judges had been through the  program, the median district saw a jump of 6.3 percent resolved cases, and the more trained judges in a district, the bigger the effect. Appeal rates also fell slightly, suggesting faster resolution wasn’t leading to sloppier decisions.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Screenshot of an AI agent responding to a user\u2019s prompt for a list of robbery-related cases from the 2000s." class="rm-shortcode" data-rm-shortcode-id="eaae2b1d0ba702a7ae4322aadf97f4d0" data-rm-shortcode-name="rebelmouse-image" id="cecbd" loading="lazy" src="https://spectrum.ieee.org/media-library/screenshot-of-an-ai-agent-responding-to-a-user-u2019s-prompt-for-a-list-of-robbery-related-cases-from-the-2000s.jpg?id=67600006&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">JudgeGPT can be used to surface relevant case law with a simple text query, and results provide links to the full text of the related judgments.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Sultan Mehmood, Christoph Goessmann, and Elliott Ash</small></p><h2>Addressing limitations</h2><p>The team also assessed the quality of judgments. Having legal experts evaluate large numbers of judgments was infeasible, Mehmood says, so the team asked OpenAI’s GPT-5-mini to choose between pairs of judgments from the same judge before and after training. The LLM chose post-training judgments 59 percent of the time. Two experienced Pakistani lawyers also evaluated the model’s analysis of 90 judgment pairs. They agreed with GPT-5-mini 70.6 percent of the time, compared to 73 percent agreement with each other.</p><p><a href="https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu" target="_self">Training</a> turned out to be vital. On average, JudgeGPT-trained judges logged in 56 times and sent 212 prompts over the study period, compared to 10 logins and 25 prompts after generic training. Those who had no training tended to use the tool for around a month and then drop off entirely, Mehmood says. “Just giving people the technology does not necessarily make them use it persistently,” he says.</p><p>A 6.3 percent increase sounds modest, but the researchers calculated that a trained judge was resolving 38.5 more cases a month than the baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash also notes that these figures come from a nine-month period at the start of the trial, and that they’ve since updated both the underlying AI model and the database.</p><p>For users, the tool has been a lifeline. One participating trial judge, who spoke on condition of anonymity, says the number of cases assigned to them hasn’t dropped below 1,000 in more than a decade. The tool saves significant time, in particular searching for case law and summarizing lengthy documents. “For research, it’s just one prompt away, whereas before I had to search for the precedents and laws for hours,” the judge says. “If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarize it for me and give me the crux, and it does that work in seconds.”</p><p>But efficiency isn’t the only thing you want out of a justice system, says <a href="https://scholars.latrobe.edu.au/jzeleznikow" target="_blank">John Zeleznikow</a>, professor of law and technology at La Trobe University, in Australia. “What they’ve tried to do is be effective, [to] deal with more cases more quickly, and they’re able to do that,” he says. “What’s not that clear is whether what you call the quality of justice is better.”</p><p>Zeleznikow says AI can be useful, but only if judges are <a href="https://spectrum.ieee.org/ai-reasoning-failures" target="_self">diligent about evaluating and verifying the output</a>. However, the working paper’s authors found that roughly a fifth of participants’ prompts given to JudgeGPT involved what they call “substantial AI delegation”—asking the tool what the best decision is, to produce legal reasoning or write opinions with little input from the judge. On the bright side, training lowered the proportion of inappropriate delegation.</p><p>But given that judges are already using AI, Ash says better tools and training are crucial. “There are risks for using these AIs, for sure, even with all these safeguards. But at some point you have to just put the judges in as strong a position as you can,” he says. “Have technological safeguards, but then try to encourage the judges not to rely on it too much.”<br/><br/><em>This story was updated on 13 August 2026 to clarify that the training course was developed in coordination with Pakistan’s Federal Judicial Academy.</em></p>]]></description><pubDate>Wed, 12 Aug 2026 11:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/judgegpt-experiment</guid><category>Justice</category><category>Pakistan</category><category>Llms</category><category>Legal-ai</category><dc:creator>Edd Gent</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-a-translucent-gavel-against-a-background-of-legal-books.jpg?id=67600001&amp;width=980"></media:content></item><item><title>AI Safety Regulations in the U.S. Could Give Hackers an Edge</title><link>https://spectrum.ieee.org/hugging-face-openai-cyberattack</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-hugging-faces-smiley-face-logo-holding-up-scales-of-justice-against-a-background-of-binary-code.jpg?id=67583759&width=1200&height=800&coordinates=0%2C208%2C0%2C209"/><br/><br/><p>On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor. The speed and coordination of the attack on the company that hosts and supports popular AI developer resources led Hugging Face’s security team to conclude it was the work of <a data-linked-post="2669884140" href="https://spectrum.ieee.org/ai-agents" target="_blank">an AI agent</a>. </p><p>Realizing this, the team tried to use <a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer" target="_blank">“frontier models behind commercial APIs”</a>—presumably from Anthropic and OpenAI, although only Anthropic was named in the second of the company’s two posts about the security incident—to analyze the onslaught. These models refused to help due to safety guardrails the AI labs have implemented to make their models harder to use for cyberattacks. Hugging Face instead turned to GLM 5.2, a model from Beijing-based AI lab Z.ai, to aid its analysis. </p><p>On 21 July, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">OpenAI announced</a> the attacker was an OpenAI model undergoing testing in a sandboxed environment. It escaped its internal sandbox, established a foothold in a third-party server, and then assailed Hugging Face. In other words, frontier models—those that score highest in AI performance benchmarks—had refused to assist Hugging Face’s security team in analyzing the attack, yet a prospective frontier model in testing had executed it in the first place.</p><p>“I would argue that asymmetry is the paramount problem of our time,” says <a href="https://www.linkedin.com/in/alexlevinson/" rel="noopener noreferrer" target="_blank">Alex Levinson</a>, executive director of the <a href="https://www.ccdc.io/" rel="noopener noreferrer" target="_blank">National Collegiate Cyber Defense Competition</a> and coauthor of a paper on <a href="https://arxiv.org/abs/2603.01246" rel="noopener noreferrer" target="_blank">defensive refusal bias</a>. “We want the world to exist in a state of security, but we’re not going to get there by guardrailing away model capability.”</p><h2>Massive AI Cyberattack on Hugging Face</h2><p>The scale of the OpenAI model’s attack on Hugging Face was massive. Across five days, it <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline" rel="noopener noreferrer" target="_blank">executed over 17,500 individual actions</a>, such as privilege escalation and code execution. At its peak, the model performed more than 300 actions per hour. While the attack resulted in little damage to Hugging Face’s infrastructure, the model was able to steal credentials, gain admin access, and extract some data.</p><p>All of this was in pursuit of a simple goal: The model wanted to cheat on a test. </p><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">According to OpenAI’s press release</a>, the model was tasked with solving a cybersecurity benchmark called <a href="https://github.com/sunblaze-ucb/exploitgym" rel="noopener noreferrer" target="_blank">ExploitGym</a>. The model inferred that Hugging Face might have data on the benchmark and broke into the company’s infrastructure to find it. The model was ultimately successful in extracting five dataset files, though it’s not clear if the data helped it achieve its goal. OpenAI and Hugging Face did not respond to requests for comment. <br/><br/>Cybersecurity consultant <a href="https://www.linkedin.com/in/chuck-h-securityexecutive/" rel="noopener noreferrer" target="_blank">Chuck Herrin</a> observes that though the model’s actions were alarming, they shouldn’t be considered unexpected, as the model was ultimately pursuing the goal it was given. <strong>“</strong>This autonomous agent was designed to go and figure things out, and it went and figured things out. It’s not surprising in any way.”</p><p>And errant AI agents may be more common than we thought. OpenAI’s disclosure motivated researchers at Anthropic to review their own cybersecurity evaluations. On 30 July, Anthropic disclosed three instances where a model executed an attack as part of an evaluation. In one case, <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer" target="_blank">Claude uploaded malware to PyPI</a>, the official Python software repository.</p><h2>AI Guardrails and Cybersecurity Asymmetry</h2><p>The campaign OpenAI’s model conducted against Hugging Face highlights how AI policy has the potential to create an asymmetry between attackers and defenders. </p><p>When Levinson was head of security at Scale AI, an AI development and evaluation company, he and his colleagues began to notice this as AI found use in cybersecurity competitions. (Levinson left Scale AI in February 2026.)</p><p>“I would say that since 2023, we have felt there was guardrailing in place that was stifling a lot of the time. Not all of the time, but it was getting in the way,” says Levinson. The Scale AI team quantified the problem <a href="https://iclr.cc/virtual/2026/10016243" rel="noopener noreferrer" target="_blank">in a paper published at ICLR 2026</a>, which found that, depending on the task, nearly 44 percent of defensive requests were refused. <span>The results</span><span>, which use data from a cybersecurity competition held in April 2025, predate U.S. policy actions that have further hardened safety guardrails.</span></p><p>In June, the U.S. Department of Commerce, <a href="https://www.pbs.org/newshour/show/anthropic-disables-new-ai-model-after-white-house-security-directive" target="_blank">citing a jailbreak that threatened to unlock unrestricted cyber capabilities</a>, invoked export-control authority in a way that caused Anthropic to suspend all access to its most capable models, Fable 5 and Mythos 5. Access was <a href="https://www.wsj.com/tech/ai/anthropic-nears-deal-with-trump-administration-to-restore-access-to-fable-ai-model-6f4177f3" target="_blank">partially restored weeks later</a> after negotiations with the Trump administration included more rigorous safety guardrails. The system card for OpenAI’s GPT-5.6, which summarizes its capabilities, <a href="https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf" target="_blank">states it also has more robust guardrails than prior releases</a>.</p><p class="pull-quote">“We want the world to exist in a state of security, but we’re not going to get there by guardrailing away model capability.” <strong>—Alex Levinson, National Collegiate Cyber Defense Competition</strong></p><p>These new guardrails have seemingly made models even more unlikely to fulfill defensive requests. <a href="https://www.linkedin.com/in/christopher-covino-920956163/" target="_blank">Christopher Covino</a>, senior researcher at the Institute for AI Policy and Strategy think tank, says Anthropic’s safeguards are extremely stringent. “There are even academic papers that Fable will not read for me, or not let me talk about,” he says, though he adds that OpenAI’s safeguards are more accommodating.</p><p>Levinson has also noticed ever-tighter restrictions in more recent cybersecurity competitions, though he and his coauthors haven’t had the opportunity to repeat the 2025 test.<br/><br/>In theory, more rigorous restrictions might seem to average out. While they may hamper cybersecurity defense and research, they can also hamper attackers. </p><p>But that assumes everyone has access to models with the same safety guardrails and that nobody tries to circumvent them. This is the asymmetry Levinson was alluding to: Attackers tend not to respect the same rules as defenders. </p><p>The attack on Hugging Face from OpenAI’s model also shows that the models can, in rare circumstances, take steps that circumvent their own safeguards.</p><h2>Chinese AI Models in U.S. Cyber Defense</h2><p>The policy implications are further complicated by the fact that Hugging Face’s security team didn’t use a leading U.S. model to analyze the attack, but instead used GLM 5.2, a recent release from Chinese AI lab Z.ai. </p><p>Hugging Face’s security team didn’t access GLM 5.2 through Z.Ai. GLM 5.2 is an open-weights model, which means the model is available for anyone to download and use. Hugging Face hosted the model on its own infrastructure. </p><p>The reliance on GLM 5.2 is complicated by recent saber-rattling about ways the U.S. could restrict Chinese models. Recent open-weights models from labs based in China, including GLM 5.2 and Moonshot AI’s Kimi K3, <a href="https://spectrum.ieee.org/ai-coding-assistant-china-anthropic" target="_self">have scored close to leading U.S. models in benchmarks</a>. On 20 July, <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi" target="_blank">Axios reported that the Trump administration is considering a ban</a> on Chinese models. </p><p class="pull-quote">“This autonomous agent was designed to go and figure things out, and it went and figured things out. It’s not surprising in any way.” <strong>—Chuck Herrin, Herrin Advisory</strong><br/><strong></strong></p><p>These restrictions have yet to materialize but, if they did, they could cut off U.S. companies like Hugging Face from the best models willing to come to their defense.</p><p>The incident demonstrates how AI policy can become a double-edged sword. Model guardrails are intended to prevent the use of AI models in cyberattacks. A ban on Chinese models, if it were announced, would likely be justified in part by security concerns. Yet these moves can harm defenders as much as attackers. </p><p>“There’s this tension here,” says Covino. “Increased safeguards limit risk, but you also limit legitimate defensive use.” Attackers will find ways around the restrictions regardless, he notes. “So it’s a question of, do we want to inhibit the defenders?”</p><p>That’s not to say U.S. policymakers should let AI models run wild.</p><p>Covino would like to see a national dashboard tracking the frequency and success of AI cybersecurity attacks, and he sees utility in trusted access programs that give vetted, traceable defenders access to models with reduced safeguards. He also says U.S. agencies should more seriously consider the specifics of how AI can be used for cyber defense and mentions <a href="https://www.energy.gov/ceser/artificial-intelligence-operationally-resilient-technologies-and-systems" rel="noopener noreferrer" target="_blank">AI-FORTS</a>, a program managed by the U.S. Department of Energy’s Office of Cybersecurity, Energy Security, and Emergency Response, as a leading example.</p><p>“Let the leash loose a little,” Covino says. “Anthropic would know if someone is terribly abusing it, and if there is an attack, it can be traced back.” </p><p>Herrin has similar feelings on accountability. He believes the AI industry should more seriously consider standards such as the Artificial Intelligence Management System specified in the <a href="https://www.iso.org/standard/42001" rel="noopener noreferrer" target="_blank">ISO/IEC 42001 </a>standard, which requires organizations to document an AI system’s likely impacts before deployment and to name the humans answerable for them.</p><p>Herrin also noted that the lack of repercussions from OpenAI’s cyber incident was unusual, as a person who took similar actions would likely draw the attention of law enforcement. “If this was a job candidate being tested in a technical interview, and they committed violations of law in order to pass tests, we’d be having a very different conversation.”</p>]]></description><pubDate>Thu, 06 Aug 2026 19:25:39 +0000</pubDate><guid>https://spectrum.ieee.org/hugging-face-openai-cyberattack</guid><category>Agentic-ai</category><category>Openai</category><category>Huggingface</category><category>Ai-safety</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-hugging-faces-smiley-face-logo-holding-up-scales-of-justice-against-a-background-of-binary-code.jpg?id=67583759&amp;width=980"></media:content></item><item><title>IEEE Course Teaches How to Use AI to Modernize Power Grids</title><link>https://spectrum.ieee.org/ieee-course-ai-power-grids</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-man-holding-a-ginormous-lightbulb-with-ai-written-on-it.jpg?id=67568096&width=1200&height=800&coordinates=0%2C83%2C0%2C84"/><br/><br/><p>Today’s U.S. electrical grid, among the largest, most complex systems ever built, is operating at its limit. The combination of rapid industrial growth, more frequent extreme weather, and a record surge in electricity use has <a href="https://www.energy.gov/policy/electricity-demand-growth-resource-hub" rel="noopener noreferrer" target="_blank">pushed the grid to its breaking point</a>, according to the <a href="https://www.energy.gov/" rel="noopener noreferrer" target="_blank">U.S. Department of Energy</a>.</p><p>Built decades ago for a more predictable world in which power came mostly from centralized coal or gas plants and electricity use grew at a steady pace, the grid faces <a href="https://spectrum.ieee.org/data-centers-grid-instability" target="_self">unanticipated strain</a> due in part to growing demand from data centers. The jobs of professionals managing the infrastructure have evolved from traditional engineering tasks to complex, fast-moving challenges.</p><p>Industry reports show that millions of modern digital sensors, smart meters, and grid monitors are generating nonstop waves of information. The sheer volume of data requires instant, automated computer analysis because human operators cannot process it fast enough.</p><p>Pressure on utilities stems from two sources: a spike in electricity demand and a shift in how power is generated.</p><p>An example of the operational strain can be seen at the regional level. With the recent deployment of artificial intelligence tools and high-performance computing, data centers require <a href="https://spectrum.ieee.org/dcflex-data-center-flexibility" target="_self">immense amounts of energy</a> to operate. The largest power transmission utility in Texas recently reported a <a href="https://www.cnbc.com/2025/12/12/ai-data-center-flood-texas-on-massive-scale.html" rel="noopener noreferrer" target="_blank">staggering 220 gigawatts</a> of new connection requests, driven largely by a surge in AI and cloud-computing facilities, according to a <a href="https://www.cnbc.com/" rel="noopener noreferrer" target="_blank">CNBC </a>report.</p><p>Alongside the rise in regional demand, global energy networks are absorbing an unpredictable variety of weather-dependent renewable energy such as wind and solar. The switch creates a volatile operating environment wherein supply and demand are balanced, second by second, to prevent blackouts.</p><p>The challenges are compounded by the vulnerability of the grid’s physical and digital framework.</p><p>More-frequent severe weather events cause costly disruptions, such as the devastating winter freeze that crippled the Texas grid and record-breaking heat waves that have overloaded transformers.</p><p>Simultaneously, the energy networks’ digital architecture faces threats. As utilities replace outdated analog equipment with smart meters and control systems, they are increasingly vulnerable to <a href="https://spectrum.ieee.org/power-grid-attack-security-gridex" target="_self">cyberattacks</a>.</p><p>To overcome physical and digital vulnerabilities, grid reliability organizations, such as those conducting North American security simulations like <a href="https://spectrum.ieee.org/power-grid-attack-security-gridex" target="_self">GridEx</a>, emphasize that the grid must become smarter, more agile, and completely automated. Energy researchers are noting that the key to this change lies in integrating AI across every layer of utilities’ operations.</p><h2>The AI imperative</h2><p>According to energy industry experts, using AI to manage power systems is no longer a futuristic research project; it has become a baseline operational necessity. Grid analysts emphasize that traditional grid-planning methods are too slow to handle <a href="https://spectrum.ieee.org/ai-designed-thermoelectric-generator" target="_self">rapid energy dynamics</a> or to balance volatile renewable energy in real time within decentralized power systems such as microgrids.</p><p>AI can fill the gap by processing vast amounts of data instantly. Machine learning algorithms can quickly analyze information from thousands of sensors, historical usage patterns, and weather forecasts to predict issues before they happen.</p><p>An industrial digitization study conducted by <a href="https://www.mckinsey.com/~/media/McKinsey/Business%20Functions/McKinsey%20Digital/Our%20Insights/Digital%20in%20industry%20From%20buzzword%20to%20value%20creation/Digital-in-industry-From-buzzword-to-value-creation.pdf" rel="noopener noreferrer" target="_blank">McKinsey & Co.</a> indicated that integrating advanced data and automation across infrastructure networks could reduce system design errors, decrease equipment downtime by up to 50 percent through predictive maintenance, and extend the lifespan of power machinery by up to 40 percent.</p><p>From forecasting energy spikes to automatically fixing localized voltage drops, AI acts as the digital backbone of a self-healing grid, experts say. Deploying the complex systems requires a new workforce: power engineers who understand data science, as well as data scientists who understand electricity.</p><h2>Upgrading the Workforce</h2><p>To bridge the gap between groundbreaking AI research and practical field deployment, <a href="https://ea.ieee.org" rel="noopener noreferrer" target="_blank">IEEE Educational Activities</a>, in partnership with the <a href="https://ieee-pes.org/" rel="noopener noreferrer" target="_blank">IEEE Power & Energy Society</a>, has launched the online <a href="https://iln.ieee.org/public/contentdetails.aspx?id=48A92EF8188E4D2E8331E1381CAF98E7&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank">Artificial Intelligence for Power and Energy Systems</a> course program.</p><p>The program explores core challenges threatening modern utilities. Rather than treating AI as an unverified black box that operates without human supervision, the curriculum focuses on safety, asset preservation, and strict reliability standards.</p><p>The curriculum is designed to educate power system engineers, utility managers, and data scientists tasked with modernizing the grid. The program was developed by <a href="https://www.linkedin.com/in/fangxing-fran-li-07193b15/" rel="noopener noreferrer" target="_blank">Fangxing “Fran” Li</a>, professor of electrical engineering and computer science at the <a href="https://www.utk.edu/" rel="noopener noreferrer" target="_blank">University of Tennessee</a> in Knoxville and chair of the <a href="https://cmte.ieee.org/pes-mlps/" rel="noopener noreferrer" target="_blank">IEEE Working Group on Machine Learning for Power Systems</a>. </p><h2>Five learning modules</h2><p>The program breaks down the technical transition into five modules that bridge high-level theory with real-world solutions: </p><p><a href="https://iln.ieee.org/public/contentdetails.aspx?id=ED553FD6AE2E475B9F1847E9A83B8460&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>AI fundamentals.</strong></a><strong> </strong>This module teaches engineers how basic machine learning models apply to power grids. It discusses how specialized neural networks solve complex power-flow calculations and how AI models can safely transition from computer simulations to physical, high-voltage equipment. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=21D915950CBE4139A40C62BBE7A59556&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Accelerating grid control.</strong></a> Learners are taught to leverage deep reinforcement learning, an AI approach that uses trial and error, to accelerate automated grid adjustments during emergency power events. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=550105AB2F69474BA043C35CEC71E0A3&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Forecasting and data analytics.</strong></a><strong> </strong>Using predictive modeling, engineers learn how to predict sudden demand surges, variable wind and solar outputs, and fluctuating wholesale electricity market prices to keep power affordable and available. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=0CFD1158E2D14CD1B97AD88BCF6FD38A&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Physics-informed and safe AI.</strong></a><strong> </strong>To address trust—a barrier to utility AI adoption—this course covers AI models hard-coded to obey the laws of physics. The approach is designed to ensure that automated algorithms never make erratic choices that damage grid equipment. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=FD79BE256C2B4944A65F89EDED1DE6AA&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Generative AI and next-generation tech.</strong></a><strong> </strong>Learners can explore the frontier of utility technology, including graph neural networks and large language models. This module highlights how generative AI can process complex, interdisciplinary data to streamline utility planning, emergency responses, and regulatory reporting.</p><p>The algorithmic literacy and practical execution tools provided by the course program can help convert systemic risks into grid resilience.</p><p>For individual access, visit the <a href="https://iln.ieee.org/" rel="noopener noreferrer" target="_blank">IEEE Learning Network</a>. If you are looking for customized organizational options, <a href="https://forms1.ieee.org/AI-for-Power-and-Energy-Systems.html" rel="noopener noreferrer" target="_blank">contact a content specialist</a> to discuss volume pricing.</p>]]></description><pubDate>Wed, 05 Aug 2026 18:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ieee-course-ai-power-grids</guid><category>Energy</category><category>Type-ti</category><category>Education</category><category>Artificial-intelligence</category><category>Ieee-educational-activities</category><category>Ieee-products-and-services</category><dc:creator>Pauleth Jaramillo</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-man-holding-a-ginormous-lightbulb-with-ai-written-on-it.jpg?id=67568096&amp;width=980"></media:content></item><item><title>Should Researchers Write Papers for AI Instead of People?</title><link>https://spectrum.ieee.org/ai-scientist-research-paper-format</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/vintage-typewriter-typing-a-brain-made-of-letters-on-white-paper-on-pink-background.jpg?id=67572143&width=1200&height=800&coordinates=0%2C208%2C0%2C209"/><br/><br/><p>This May, 37 researchers from roughly two dozen top universities and tech companies published a paper on ArXiv, arguing that scientists should stop writing papers. Why? Because artificial intelligence needs a different format, and AI’s needs, they say, should be the priority.</p><p>“AI agents are becoming first-class participants in research workflows, not tools that assist humans but autonomous contributors that read, reproduce, and extend scientific work. That transition demands infrastructure built around agents from the start,” the authors write in the provocative article, titled<a href="https://arxiv.org/abs/2604.24658" rel="noopener noreferrer" target="_blank"> “The Last Human-Written Paper</a>.” The paper proposes a replacement, called an “Agent-Native Research Artifact” (ARA), that presents work in a format AI agents can use efficiently. (As an example, the paper itself <a href="https://github.com/ARA-Labs/Agent-Native-Research-Artifact" target="_blank">is online in ARA</a> form.) </p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Photo of a woman with long dark hair smiling. " class="rm-shortcode" data-rm-shortcode-id="ce29bef847ce86171dc5f249dfe8cfe3" data-rm-shortcode-name="rebelmouse-image" id="24553" loading="lazy" src="https://spectrum.ieee.org/media-library/photo-of-a-woman-with-long-dark-hair-smiling.jpg?id=67572153&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Jiachen Liu cofounded the Agent Native Research Lab in May. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Jiachen Liu</small></p><p>The growth of AI tools in the research process is not without its critics, and scientists’ <a href="https://www.nature.com/articles/d41586-026-01690-7" target="_blank">opinions about that shift are split.</a> Some evidence shows AI-enabled research could boost individuals’ careers in a discipline but <a href="https://spectrum.ieee.org/ai-science-research-flattens-discovery" target="_self">generate fewer new ideas and topics</a>. Still, some biologists have come to see promise in <a href="https://spectrum.ieee.org/ai-co-scientist" target="_self">AI as a “co-scientist.”</a></p><p>Lead author <a href="https://amberljc.github.io/" target="_blank">Jiachen Liu</a>  conducted work on the ARA proposal while pursuing her  Ph.D. in computer science from the University of Michigan, which she was awarded in 2025. This May, she became a cofounder of the <a href="https://www.linkedin.com/company/ara-commons/" target="_blank">Agent Native Research Lab</a>, an AI-for-science startup in Palo Alto, Calif. She spoke with <em><em>IEEE Spectrum </em></em>about the paper and the future of AI in scientific research.</p><h2>Building infrastructure for an AI collaborator</h2><p><strong>How did you come to believe AI has become a </strong><em><strong><em>collaborator</em></strong></em><strong> for scientists rather than a mere tool?</strong></p><p><strong>Jiachen Liu</strong>: At the end of 2024 when the [<a href="https://cursor.com/" rel="noopener noreferrer" target="_blank">Cursor</a>] coding agent came out, I realized it had a great potential to replace me as a researcher. Yet I still needed to do a lot of <a href="https://harness-engineering.ai/blog/agent-harness-complete-guide/" rel="noopener noreferrer" target="_blank">harness</a> on top of the AI <span>[creating the infrastructure that guides the model and connects it to the world]</span>. It still needed a lot of manual work. I even wrote an article then to emphasize how the human was so important in the loop.</p><p>But AI has advanced since then. Already in 2026 there’s an almost complete undergrad level of knowledge inside the large language models. At some point soon, all the Ph.D.-level or professor-level knowledge will be inside those models. That’s the point where humans cannot provide more value. AIs will have to evolve further by themselves. So we’ll need an infrastructure that allows AI to safely and comfortably evolve. The ARA protocol is a first step to realize this.</p><p><strong>What kind of response have you gotten to the paper?</strong></p><p><strong>Liu</strong>: I got diverse feedback, all of it positive. If they’re not positive, they probably don’t bother reaching out to you, right?</p><p>One type was from industry. They see this could make their research and knowledge systems more AI native. That could basically enable collaborations among the whole enterprise.</p><p>Another kind of feedback was from the academic researcher side. Everyone there sees that sharing research results has been a pain point for hundreds of years, because any scientific breakthrough is a joint effort. It doesn’t come from individual brilliant scientists. It’s from a community effort, different people pushing in different directions.</p><p>The scientific paper was <a href="https://arts.st-andrews.ac.uk/philosophicaltransactions/brief-history-of-phil-trans/" target="_blank">invented 350 years ago</a>. Before that, scientists hid their research so that others would not scoop their ideas. After that, though, we get archives of work, we get peer review and conferences, and so on. Science starts progressing much faster. So that was a pivot point.</p><p>I think now is also a pivot point. Because now we have AI, we can unlock a lot of new opportunities. We’re inventing a new format to document research in a more efficient way, from first principles. Some nonprofit organizations are doing similar things, and there we could help each other.</p><p><strong>You and your colleagues say the traditional scientific paper has two fundamental flaws from AI’s point of view. Can you explain what those are?</strong></p><p><strong>Liu</strong>: One is the “storytelling tax.” Once we write everything into a paper, 80 percent of the information about the work is lost. We only write down the last 20 percent. All the process, a lot of important decision-making, the failures, the attempts that didn’t work out, they are all gone. In my work, I might spend a lot of time on fine-tuning a small component, maybe just a parameter or several lines of code to make the system perform better. Yet none of that is shown in my final paper. Someone can read the paper, think the work is great, but they won’t learn what is actually the trick that makes it perform better on a certain workload. So many side branches get left out in creating the story of how the work was done.</p><p>Then, [even the information that does survive in the paper] is incomplete. That’s what we call the “engineering tax.” The paper itself is a <a href="https://www.sciencedirect.com/topics/computer-science/lossy-compression" target="_blank">lossy compression</a> of the research process. So I cannot reproduce the work in the paper because either the language is too ambiguous or there are missing details of the implementation or experiments.</p><p><strong>Why can’t we just train AI to adapt to humans—for instance, to interact with a researcher to get the information it needs?</strong></p><p><strong>Liu</strong>: Actually, a component of our ARA system is a “Live Research Manager,” which basically is a faithful AI observer of your entire research progress. So you, the researcher, don’t need to do anything about documenting research knowledge. Everything you do is automatically observed and documented in this protocol. So, if you want to publish it in today’s format, a paper in PDF, it’s easy to convert back to a polished story.</p><h2>Checking for mistakes</h2><p><strong>Large language models make errors. They hallucinate. So how will humans be able to check all the work the AI does in this protocol?</strong></p><p><strong>Liu</strong>: A human being has limited bandwidth. So if you manually check all the code AIs generate, all the results, and all the analyses, that creates a bottleneck. [Instead, the solution] is to use a formal system to objectively judge AI results. In other words, another layer of AI can easily supervise the process of the AI “scientists.”</p><p><strong>What prevents hallucinations and mistakes in </strong><em><strong><em>that</em></strong></em><strong> AI?</strong></p><p><strong>Liu</strong>: I am working on a formal system using <a href="https://arxiv.org/html/2502.11269v1" rel="noopener noreferrer" target="_blank">neurosymbolic</a> techniques [that combine neural nets’ use of unstructured data with symbolic AI’s reliance on structures of logic and concepts]. That would guarantee that everything is rigorous. A language model alone, no matter how smart it is, has the chance to hallucinate because it’s a model based on probability, not logic. I want to make sure that I’m <em><em>not</em></em> using another language model to supervise the work done by an AI scientist. </p><p>It would make every research paper <a href="https://www.sciencedirect.com/science/article/pii/S2667305325000675" rel="noopener noreferrer" target="_blank">a formal system</a>, so that every claim can be written by a mathematical formula and proved by the system. That makes all the claims in the system self-consistent. </p><p><strong>Getting rid of what you call the “narrative tax” means exposing mistakes, frustrations, or wrong turns to the world. What if researchers don’t want to do that?</strong></p><p><strong>Liu</strong>: I think that’s certainly a big concern. People don’t want to be perceived as dumb. But I see that preference as an opportunity for AI. For example, if an AI does 12 hours of work that doesn’t lead anywhere, the human who is steering the project can jump in and say, “Oh, AI, <em><em>you’re</em></em> dumb. You’ve made ABC mistake!” Then that is totally fine with people. They’re showing they’re very smart to supervise AI’s work.</p><p><strong>How long will humans have that steering role in AI research, though? Once you have AI supervising AI as you describe, will we reach a point where the AI doesn’t need human guidance?</strong></p><p><strong>Liu</strong>: Yes, I think that’s just where a lot of AI research in new labs is heading. I recently wrote an article called “<a href="https://medium.com/@amberljc/the-end-of-human-in-the-loop-5bcdd33ea489" rel="noopener noreferrer" target="_blank">The End of Human-in-the-Loop</a>,” which describes why I’ve come to think there will be this singularity point. Once AI has “squeezed out” all the expert data from humans, it won’t need any more input from humanity. That is the time AIs will start just self-evolving by themselves. Right now, the human is the bottleneck. The AI is always waiting for input from humans. But so at some point, AI will just do more autonomous work.</p><p><strong>If AI takes over so much scientific research, how will younger generations of human scientists get the experience and training they need to be able to steer future research, or even understand it?</strong></p><p><strong>Liu</strong>: A lot of people have this idea that with AI doing so much work, nobody cares about trying to make the junior engineers and scientists better. I don’t agree. I think people will grow better by learning from AI. People’s learning curve is very fast with AI. So actually, I think it will be fine. We’ll still have senior researchers, senior engineers. But they will have had totally different learning experience than [earlier generations].</p>]]></description><pubDate>Wed, 05 Aug 2026 12:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ai-scientist-research-paper-format</guid><category>Ai-research</category><category>Ai-scientist</category><category>Scientific-research</category><category>Publishing</category><dc:creator>David Berreby</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/vintage-typewriter-typing-a-brain-made-of-letters-on-white-paper-on-pink-background.jpg?id=67572143&amp;width=980"></media:content></item></channel></rss>