<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>IEEE Spectrum</title><link>https://spectrum.ieee.org/</link><description>IEEE Spectrum</description><atom:link href="https://spectrum.ieee.org/feeds/topic/artificial-intelligence.rss" rel="self"></atom:link><language>en-us</language><lastBuildDate>Mon, 07 Sep 2026 12:10:56 -0000</lastBuildDate><image><url>https://spectrum.ieee.org/media-library/eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpbWFnZSI6Imh0dHBzOi8vYXNzZXRzLnJibC5tcy8yNjg4NDUyMC9vcmlnaW4ucG5nIiwiZXhwaXJlc19hdCI6MTgyNjE0MzQzOX0.N7fHdky-KEYicEarB5Y-YGrry7baoW61oxUszI23GV4/image.png?width=210</url><link>https://spectrum.ieee.org/</link><title>IEEE Spectrum</title></image><item><title>AI Efficiency Could Cost Us the Next Generation of Experts</title><link>https://spectrum.ieee.org/ai-engineer-skills</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/human-and-robotic-hands-share-a-caliper-over-technical-engineering-blueprints.png?id=67702640&width=1245&height=700&coordinates=0%2C194%2C0%2C195"/><br/><br/><p><span>A little over a decade ago, I led the controls design for a first-of-its-kind full digital-control system for a U.S. nuclear plant. It was, on paper, a beautiful machine—engineered to run itself the way a modern airliner does, with operators watching over a system that rarely needed them. And we made a decision that, to an efficiency-minded observer, looked backward: We deliberately left manual steps inside sequences the system could execute on its own.</span></p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/ai-engineer-skills&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p><span>We were solving a specific problem. An operator who only ever supervises automation slowly stops being an operator. The hands go cold. The mental model of what the plant is actually doing gets fuzzy. Then comes the day the automation hands control back. It’s always the worst day, because automation only quits when it’s confused or in trouble. But by then, you have a person in the chair who hasn’t truly operated the thing in years. The manual steps were there to keep the human current. It was inefficient by design, on purpose.</span></p><p>That plant, as it happened, was never built. It was shelved amid the politics and economics that surround <a href="https://spectrum.ieee.org/tag/nuclear-power" target="_blank">nuclear power</a> in this country, for reasons that had nothing to do with the engineering. But the design instinct outlived the project, and I’ve come to believe it’s the most useful idea I can offer to the argument now consuming every boardroom: What happens to human expertise when AI does the work that used to build it?</p><h2>AI Is Disrupting the Engineering Career Ladder</h2><p>The data has gotten hard to wave away. A Harvard University working paper covering some 65 million workers at more than 280,000 U.S. firms found that after companies adopted generative AI, <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5425555" target="_blank">junior employment fell roughly 9 percent</a> within six quarters relative to nonadopters, while senior employment kept right on growing. A Stanford analysis of ADP payroll records points the same way: The youngest workers in the most AI-exposed occupations <a href="https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/" target="_blank">lost ground after late 2022</a> while their more-experienced colleagues held theirs. The Stanford researchers found that the losses concentrate where AI automates the work; where it merely augments, junior employment holds steady or rises.</p><p>The causal story is still contested, and honesty requires saying so. Researchers at the New York Fed attribute much of the rise in young-graduate unemployment <a href="https://libertystreeteconomics.newyorkfed.org/2026/06/remote-work-leaves-younger-workers-sidelined/" target="_blank">not to AI but to remote work</a>, arguing that firms are reluctant to hire inexperienced people whom they cannot train and mentor at a distance. But notice what the explanations share. Whether a model is absorbing the formative work or distance is severing the mentorship around it, both describe the same broken mechanism: the apprenticeship channel through which expertise passes from senior to junior. Either way, “entry-level” has quietly come to mean “three years of experience required.”</p><p>Strip away the noise and you’re left with one deceptively simple problem: You cannot become a senior engineer without first being a <a href="https://spectrum.ieee.org/ai-effect-entry-level-jobs" target="_blank">junior one</a><em>.</em> Expertise is not downloaded. It is earned through failed builds, dead-end debugging sessions, and the “why on earth did that work” moments that a capable AI will now happily spare the newcomer. Spare them enough of those and you produce a cohort that can supervise a model on paper but never developed the gut sense to know when the model is confidently, catastrophically wrong.</p><p>Most of the commentary stops at the diagnosis, or reaches for policy solutions that treat the loss of junior jobs as an economic problem. Yet it’s also an engineering problem, and safety-critical fields have already spent decades learning how to solve it.</p><h2>Aviation’s Lessons About the Automation Paradox</h2><p>My own career started at the sharp end of automation. My first job out of school was verifying and validating the software in the digital jet-engine controller that decides, faster than any pilot could, how a fighter plane’s engine responds. Even then, in the late 1980s, the central tension was visible: The machine outperforms the human in routine cases, but the human is all that stands between the aircraft and disaster in the cases the machine didn’t anticipate. This tension is known as the <a href="https://spectrum.ieee.org/tag/automation-paradox" target="_blank">automation paradox</a>, in which increasingly capable automation gives human operators less practice, while leaving them only the most difficult situations.</p><p>Aviation learned, repeatedly and expensively, what happens when human skills atrophy inside that gap. The canonical example is <a href="https://en.wikipedia.org/wiki/Air_France_Flight_447" target="_blank">Air France flight 447</a>, which fell into the Atlantic in 2009. The proximate cause was mundane. Iced-over airspeed sensors fed the autopilot bad data, and it did what it is designed to do: It disconnected and handed control of the airplane back to the crew. What followed was not a hardware failure. It was a competence failure. A recoverable situation became an unrecoverable one because the pilots, conditioned by thousands of hours of watching the automation fly, could not read a high-altitude aerodynamic stall and hand-fly their way out of it. The airplane was working. The training the automation had quietly eroded was not.</p><p>The industry’s response is instructive, and it’s the same move we made in that nuclear control room. It did not rip out the autopilot. It built deliberate manual practice back in. In 2017 the FAA issued Safety Alert for Operators 17007, “<a href="https://www.faa.gov/sites/faa.gov/files/2022-11/SAFO17007.pdf" target="_blank">Manual Flight Operations Proficiency,</a>” declaring that “manual flight is the foundation upon which other technical flying skills are built.” The alert formally recognized skill decay as a hazard in its own right. Some airlines amended their procedures to encourage hand-flying both the initial climb and initial descent in benign conditions, knowingly trading a sliver of fuel efficiency to keep the crew’s raw flying skills alive. That trade is the whole point. A perfectly optimized system that produces incompetent operators is not optimized at all. It has simply moved its failure mode somewhere the spreadsheet can’t see it.</p><h2>Manual Gates Could Preserve Engineering Skills</h2><p>Put the aviation lesson and the nuclear instinct side by side and they point to one design pattern we now need in AI-augmented work: the deliberate “manual gate.”</p><p>A manual gate is a point in a workflow where a human takes the controls, not because it is the fastest way to get the task done, and not only as a safety interlock, but specifically to exercise and preserve a skill that would otherwise decay. The distinguishing feature is that it is chosen. You decide, as a matter of design, which competencies your organization must keep alive in human beings because those are the ones you will need on the bad day. Then you engineer the friction required to keep them warm.</p><p>Picture how this might work on a software team that leans on AI for most of its code. The team places a manual gate around the skill it can least afford to lose: <a href="https://spectrum.ieee.org/tag/debugging" target="_blank">debugging</a>. When a defect surfaces in a critical module, the assigned engineer—deliberately, often a junior one—must first reproduce the failure, trace it to root cause, and write an automated test that captures the bug, all with the AI assistant switched off. Only after the engineer commits to a diagnosis does the model come back on, to propose the fix, generate alternatives, and sweep the code base for similar bugs. The engineer then compares their diagnosis against the model’s. When the two disagree, that’s the design working, surfacing the disagreement before the bad day instead of during it.</p><p>This approach reframes the junior engineer entirely. The instinct today is to let AI do the entry-level work because it is faster and cheaper. But some of that work is not overhead to be eliminated. It is the training apparatus of your future senior staff, and you should protect it the way you’d protect any other piece of critical infrastructure. It may not be efficient this quarter, but dismantling it quietly mortgages your capability a decade out.</p><h2>Why Companies Must Keep Training Junior Engineers</h2><p>None of this is free, and pretending otherwise would insult the people who have to sign the budgets. A deliberate manual gate is, by construction, less efficient in the near term than full automation. Keeping juniors doing formative work and running the manual sequences costs something now to protect something later.</p><p>That’s a hard sell in a market that judges most leaders on quarterly results. A hired executive who carries “unnecessary” humans that AI could replace will hear about it from the board long before the payoff arrives. The math only works for someone insulated from that pressure: a founder with control, a private company, an institution with a genuinely long horizon, or a regulator willing to require workers to demonstrate their skills regularly, as pilots must. Which means the organizations most likely to preserve their own expertise are the ones structurally able to spend short-term margin on long-term capability; everyone else will need that outside push.</p><p>So here is the argument, in one line: Deliberate inefficiency is not waste. In safety-critical engineering we have always known it as insurance, and we buy it on purpose. As AI takes over the work where expertise is forged, the smart move is not to resist the automation. It is to keep our hands on the controls by design—so that when the automation fails, as it always eventually does, there is still someone in the chair who knows how to fly.</p>]]></description><pubDate>Wed, 02 Sep 2026 13:00:04 +0000</pubDate><guid>https://spectrum.ieee.org/ai-engineer-skills</guid><category>Engineering-careers</category><category>Generative-ai</category><category>Automation-paradox</category><category>Aviation</category><category>Nuclear-power</category><dc:creator>Richard Mitchell</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/human-and-robotic-hands-share-a-caliper-over-technical-engineering-blueprints.png?id=67702640&amp;width=980"></media:content></item><item><title>Cash In on the AI Boom by Renting Out Your Spare Compute</title><link>https://spectrum.ieee.org/ai-inference-distributed-computing</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-small-makeshift-server-rack-in-someones-garage.jpg?id=67703292&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>If you own an at-home server, a gaming computer, or just a laptop that doesn’t get much love, listen up. You can now put that spare computing power to use and earn some passive income in the process. AI companies are hungry for more compute to run AI inference—the process of using a pretrained model to respond to queries—and they’re willing to pay you for it. </p><p>“Imagine Uber or Airbnb, but for AI-inference computing tasks,” says <a href="https://www.linkedin.com/in/shzhv13/" rel="noopener noreferrer" target="_blank">Ilman Shazhaev</a>, founder and CEO of <a href="https://farlabs.ai/" rel="noopener noreferrer" target="_blank">Far Labs</a>, based in Abu Dhabi.</p><p>The AI boom has spurred on the construction of <a href="https://spectrum.ieee.org/5gw-data-center" target="_blank">massive data centers</a>, often damaging local communities by raising electricity prices, straining local water resources, causing environmental damage and noise, and being just plain ugly. Huge data centers are likely not going anywhere—training new frontier models and running AI models from leading companies will likely still be the purview of these behemoths. But now, several companies are providing AI inference on smaller, mostly open-source models. They are running inference on the preexisting computing power spread throughout homes and small businesses, and compensating the owners.</p><p>“Everyone thinks the only way to do it is data centers. And data centers are extractive for the communities in which they’re built, and they don’t return services or taxes or much of anything to the people there. So why not just turn this whole thing on its head?” says <a href="https://www.linkedin.com/in/johnfederico/" rel="noopener noreferrer" target="_blank">John Federico</a>, founder and CEO of <a href="https://www.evolvingedge.ai/" rel="noopener noreferrer" target="_blank">Evolving Edge</a>, in Austin, Texas. “The compute power is out there. If you can orchestrate it, then you’re actually adding value to those communities directly.”</p><p>The idea isn’t entirely new: From 1999 to 2020, a volunteer-based project called <a href="https://setiathome.berkeley.edu/" rel="noopener noreferrer" target="_blank">SETI@Home</a> used spare computers to search for signs of extraterrestrial life in radio-telescope data, for instance. But now, commercial companies are eager to use the same strategy. Shazhaev’s Far Labs is launching its platform Far AI in the coming weeks, while Federico’s Evolving Edge is currently in open beta. Other companies, like <a href="https://bless.network/" rel="noopener noreferrer" target="_blank">Bless Network</a>, <a href="https://salad.com/download" rel="noopener noreferrer" target="_blank">Salad</a>,<a href="https://vast.ai/hosting" rel="noopener noreferrer" target="_blank"> </a>and <a href="https://gradient.network/" rel="noopener noreferrer" target="_blank">Gradient</a> have started to provide similar platforms over the last year.</p><h2>Connecting to the network</h2><p>Federico has been a computer hobbyist since youth, and he has amassed a whole server in his basement to run his projects. “It just hit me one day—there’s all this talk about not having enough compute, and I just thought, well, 92 percent of the country has broadband, and you have people like me who have mini data centers in a closet,” he says.</p><p>Federico sees the potential hosts as people much like himself who have already invested in home servers, and he aims to make the process of selling spare compute as seamless for them as possible.</p><p>“Sign up for the program, install an application,” Federico says. “All we want to do is run jobs on your machine when you tell us we’re allowed to. The only thing we do is monitor the resource usage. And of course, you can give us a schedule.” With a large enough network of devices, the platform would have compute available whenever it’s needed.</p><p>Privacy and security are primary concerns for such hosts. To reassure the users that their local data is secure, and that no malware will be downloaded to their devices, the team open-sourced their scheduling software. “The node software is open source, so anyone can look at it, see what it does,” Federico says. </p><p>Shazhaev of Far Labs explains that the company’s software is designed around a principle known as “least privilege,” which grants both the host and the user the minimum access possible to accomplish the task. Inference runs as an isolated workload with authenticated, encrypted communication and explicit limits on the GPU, CPU, memory, storage, and network resources it may use. Customers do not receive arbitrary access to the host machine, and providers can inspect resource use, pause the node, revoke access, and remove the software at any time.</p><p>The protection also works in the other direction. Workloads are segmented, and only the minimum required information is exposed to an individual node. Sensitive enterprise workloads can be restricted to controlled hardware rather than routed through consumer devices.</p><h2>Divide and conquer</h2><p>Massive data centers still have advantages from the user perspective: top-of-the-line GPUs and CPUs, high-speed networking, thick cables, and <a href="https://spectrum.ieee.org/data-center-liquid-cooling" target="_blank">sophisticated cooling</a>. User devices are usually less powerful, more varied, and less reliably connected to one another.</p><p>“This is quite a difficult issue from the science angle,” Shazhaev says. “You want to do a similar level of tasks that are happening in those high-infrastructure data centers, and run them on the user device with limited capacity.”</p><p>Evolving Edge’s Federico says this is an issue for the largest, state-of-the art AI models. But those are not always needed and are often not even preferred. “There are numerous companies, once they reach a certain scale, suddenly paying for tokens on a state-of-the-art frontier model [that] no longer makes sense for their needs,” he says. “Instead, they are fine-tuning open-source models for specific tasks that they have in their business. These models don’t require anywhere near the resources that some of the state-of-the-art models do. It’s just using the right tool for the job.”</p><p>Smaller, open-source models can often fit on a single user device. But if that fails, there are tools to split a single inference task over multiple GPUs or CPUs. Evolving Edge is using an open-source tool called Ray to perform this splitting, while Far Labs has developed its own proprietary software that not only splits the workload but also wraps the splitting in a layer of security and reliability-providing software. </p><p>“One thing we have done is we shared the model,” Shazhaev says. “We take the model, we cut it into many pieces, and then these pieces will be distributed through different devices. And we have an orchestrator and a load balancer which manage the task flow, so each device processes a part of the task. Then we combine the answers in the main brain, the orchestrator.”</p><p>Through a combination of using smaller, more task-specific models, and splitting larger models between disparate devices, the teams claim they can perform inference much cheaper than a traditional data center “because we don’t have capital expenditure,” Shazhaev says.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A vertically mounted graphics card inside of a gaming PC rack." class="rm-shortcode" data-rm-shortcode-id="124c2cabf5d259385159e630919004de" data-rm-shortcode-name="rebelmouse-image" id="b08b9" loading="lazy" src="https://spectrum.ieee.org/media-library/a-vertically-mounted-graphics-card-inside-of-a-gaming-pc-rack.jpg?id=67703296&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Gaming PCs are a common source of spare computational power in the home. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Dizzaract</small></p><h2>The distributed advantage</h2><p>Not only is it cheaper to run inference this way, it is also more reliable, Shazhaev claims. The companies have access to a distributed network of computing resources, rather than one giant device that can experience outages. Shazhaev compares this to cryptocurrencies and their resilience through decentralization. </p><p>“Today, to shut down Bitcoin, you need to nuke the whole planet. Here, we have the same concept,” Shazhaev says.</p><p>Federico explains that this resiliency would be beneficial not just for AI inference, but for all kinds of applications, including smart cities, environmental sensors, autonomous vehicles, and more. During an Amazon Web Services <a href="https://www.theguardian.com/technology/2025/oct/24/amazon-reveals-cause-of-aws-outage" target="_blank">outage</a> in 2026, for example, smart beds were stuck in their upright positions and their users couldn’t adjust them. Federico says that a distributed network where everything doesn’t need to be routed through a single data center, say, in Ashburn, Va., would make those kinds of outages much less impactful. “We could lose 100 nodes in a network of 250,000 and it wouldn’t matter,” he says. </p><p>If the network of user devices is substantial enough, every job can be routed to a nearby device, decreasing the latency. Far Labs claims a latency of 100 milliseconds or less on its platform. The lower cost and lower latency of this approach may even enable new use cases, such as in-game AI video generation, which is currently prohibitively slow and expensive.</p><p>“OpenAI last year had US $30 billion in revenue, but they closed the financial year at an $8 billion loss. Why? The official reason is due to the high cost of inference,” Shazhaev says. “And those are mostly text models. For gameplay, you have audio, video, animations: It’s heavy data, and you need real-time responses. So, we’ve been trying to solve this issue.”</p><p>All of these companies are trying to tap into an untapped resource of local compute and hoping it’ll benefit the device hosts and users alike.</p><p>“All these big guys are running around building data centers,” Shazhaev says, “but I believe there is enough compute power that already exists in the world.”</p>]]></description><pubDate>Tue, 01 Sep 2026 14:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ai-inference-distributed-computing</guid><category>Distributed-computing</category><category>Inferencing</category><category>Seti-home</category><category>Data-centers</category><dc:creator>Dina Genkina</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-small-makeshift-server-rack-in-someones-garage.jpg?id=67703292&amp;width=980"></media:content></item><item><title>New Platform Peers Inside AI’s Black Box</title><link>https://spectrum.ieee.org/silico-ai-interpretability</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/goodfire-ais-silico-uses-agents-equipped-with-interpretability-tools-to-examine-the-reasoning-behind-an-ai-model.jpg?id=67668252&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p><span>Prompt Claude, ChatGPT, Gemini, or any other popular <a data-linked-post="2655083903" href="https://spectrum.ieee.org/what-is-deep-learning" target="_blank">large language model</a> with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer.</span></p><p>This mysterious behavior can be useful in some situations. But—as highlighted by a recent <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank">incident</a> where <a href="https://spectrum.ieee.org/tag/openai" target="_self">OpenAI</a> could not explain why its advanced prerelease model hacked AI company Hugging Face—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI “thinking” and outputs has never been greater.</p><p><a href="https://www.goodfire.com/">Goodfire</a>, an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a <a href="https://www.goodfire.com/grants" target="_blank">new grant program</a> offering US $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes.</p><h2>Mechanistic interpretability</h2><p>Founded in 2024 and based in San Francisco, Goodfire aims to provide the tools that build the next generation of safe and powerful AI by understanding the structures inside them instead of treating AI models as black boxes. “Treating models like black boxes isn’t inevitable; it’s a choice,” says <a href="https://www.linkedin.com/in/eric-ho-53981862/" target="_blank">Eric Ho</a>, Goodfire cofounder and CEO. “With the right interpretability tools, we can see how models actually work.”</p><p>The tools Ho refers to are built around a concept called mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by interpreting the model’s weights, activations, and attention patterns, and mapping its neurons and the pathways between them.</p><p>Mechanistic interpretability tools span the gamut. One approach is mapping a model’s activations in response to controlled prompts, and matching those activation patterns to a set of concepts that humans can understand. Another tack is tracking changes in model weights before and after a specific training run in order to spot and understand what changed. Yet another option is changing specific model weights or activations and observing how that affects the model’s output. </p><p>Silico combines a broad range of these tools, and provides a layer of AI agents to help users understand their model. Users describe what they want to investigate about their AI model in plain language, asking things like ”Find out when and why my model is hallucinating.” The platform then autonomously builds an experimental plan involving a host of tasks that can be performed using the various interpretability tools and techniques at its disposal. It then sends out agents to perform these tasks in parallel. Completion of these subtasks should add up to an answer to the original prompt, or at least insights that can be inspected and built upon. </p><p>Ho says: “In a sense, Silico is like a microscope to peer inside an AI model to understand which parts are responsible for what behavior, and even edit those parts directly.”</p><h2>Understanding Alzheimer’s and AI</h2><p>These tools have already been used to make some impressive advances in a host of fields. In medicine, for instance, <a href="https://www.primamente.com/" rel="noopener noreferrer" target="_blank">Prima Mente</a>, an AI company based in the United Kingdom, <a href="https://www.goodfire.com/research/interpretability-for-alzheimers-detection" rel="noopener noreferrer" target="_blank">worked with Goodfire</a> to understand its Pleiades epigenetic foundation model. The model performed well at its task of detecting Alzheimer’s disease from blood samples, but the company didn’t know why.</p><p>“We reverse-engineered Pleiades and found it was using DNA fragment-length patterns to make its predictions—a signal humans hadn’t used to detect Alzheimer’s before,” recalls Ho. In other words, the team had discovered that Pleiades was using a completely new biomarker for the disease. “As far as we know, it’s the first significant finding in the natural sciences discovered purely by reverse-engineering a foundation model,” Ho adds.</p><p>Elsewhere, Silico is being used to explore deep questions surrounding AI. Cameron Berg, founder and director of <a href="https://reciprocalresearch.org/" rel="noopener noreferrer" target="_blank">Reciprocal Research</a> (a New York nonprofit research organization he founded to explore methods of gauging AI cognition), says that Silico almost fell out of the sky at the right time for him and his research. “Silico has been really helpful for operationalizing my research agenda and executing on it way faster than I would have expected,” he says. “I feel like I have basically become the PI [principal investigator] and my research scientists and research engineers are AI systems.”</p><p>Berg sees general access to Silico and tools like it leading to greater trust in AI’s ability to conduct research tasks, which will accelerate the scientific process across the board. But beyond scientific research, the widespread release of Silico could signal a shift in how AI innovators build, debug, and deploy their models. </p><p>“I think it’s a mistake to not understand the most consequential technology of our time, particularly given the emergent behavior we’re seeing from increasingly capable AI agents,” says Ho. “If we truly understand how AI models think, instead of discovering and trying to correct their behavior retroactively, we can design them intentionally and shape how models behave to be safer and more reliable.”</p>]]></description><pubDate>Wed, 26 Aug 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/silico-ai-interpretability</guid><category>Ai-interpretability</category><category>Ai</category><category>Science</category><dc:creator>benjamin_skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/goodfire-ais-silico-uses-agents-equipped-with-interpretability-tools-to-examine-the-reasoning-behind-an-ai-model.jpg?id=67668252&amp;width=980"></media:content></item><item><title>AI Companion Robots Are Closing the Human Connection in Modern Homes</title><link>https://spectrum.ieee.org/ollobot-ai-companion-robot</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/cute-home-robot-on-carpet-in-cozy-living-room-with-beige-sofa-and-warm-lighting.jpg?id=67154308&width=1245&height=700&coordinates=0%2C260%2C0%2C261"/><br/><br/><p><em>This article is brought to you by <a href="https://ollobot.com/" target="_blank">Ollobot</a>.</em></p><p>From about 2017, individuals began to truly connect with the initial wave of companion robots. These devices had personality, moved around, joked, and answered when you spoke to them. Most early companion robots, however, were still limited by simple voice-command interactions and narrow functionality. Once the novelty wore off, many ended up sitting unused on shelves. As some of those companies went out of business and turned off their servers, many owners likened it to losing a pet.</p><p>What Ollobot describes as “gentle intelligence” is a useful way to think about where the serious work in this category is going. Not toward more powerful assistants, but toward more present ones.</p><h2><a target="_blank"></a>The problem companion robots were trying to solve<strong></strong></h2><p>Loneliness is not a niche issue. According to one <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9957792/" target="_blank">study</a>, nearly one out of three elderly adults resides alone, meaning they do not have daily companions. <a href="https://pubmed.ncbi.nlm.nih.gov/20533912/" target="_blank">Research</a> also shows that children whose parents have migrated for work, leaving them in the care of relatives, were 2.5 times more likely to experience loneliness than children whose parents remain with them. Among working adults living alone in urban environments, similar <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6530780/" target="_blank">patterns</a> of social isolation emerge, even if they are less visible.</p><p>Over the years, technology has time and again attempted to solve this problem via video calls, smart speakers, and messaging apps without much success. Those tools are geared towards communication between people that already have relationships. They do not create presence. They schedule it. That is the gap that a new generation of AI companion robots is being engineered to fill.</p><h2><a target="_blank"></a>Today’s AI robots are different<strong></strong></h2><p>Today’s companion robots are not just cute and cuddly. They are designed with psychological research, clinical insight and long-term interaction models to be truly useful in real homes.</p><p>Three fundamental shifts define the current generation:</p><ol><li><strong>From reactive to proactive response. </strong>Older robots relied on you speaking to them, but modern robots monitor a room with cameras, microphones, and surroundings sensors to initiate interactions without your input, and they can pick up on your emotions.</li><li><strong>From function-oriented to emotion-oriented design.</strong> The original pitch for companion robots was about what they could do. The question driving the serious work now is how they make you feel, which is a harder engineering problem and a more honest framing of what the product is actually for.</li><li><strong>From standalone hardware to connected ecosystems.</strong> Leading brands are creating platforms rather than devices with software included as a built-in layer and remote access from the beginning.</li></ol><p>The <a href="https://www.grandviewresearch.com/industry-analysis/ai-companion-market-report?__cf_chl_f_tk=xJq0xhwCGb6n830sK3lMN.cyk2m73bms6.RauBufPho-1783406928-1.0.1.1-TsUMt2nxUgZ5nDX79zhxxVYkTv9xk2r_.E2rbXRHByA" target="_blank">global AI companion market</a> size was valued at US $36.8 billion in 2025 and is projected to grow from $48 billion in 2026 to $318 billion by 2033, at a compound annual growth rate of 31 percent from 2026 to 2033.<em><span><br/></span></em></p><h2><a target="_blank"></a>Three household scenarios and interaction models<strong></strong></h2><p>Ollobot’s advanced AI family companion robot <a href="https://ollobot.com/" target="_blank"><span>OlloNi SS1</span></a> addresses a number of gaps in what existing technology offers.</p><p><strong>Elderly individuals living alone.</strong> The combination of proactive interaction, fall detection, and persistent presence addresses both safety and companionship without the social overhead of asking family members to check in more frequently.</p><p><strong>Children in households where parents work far from home.</strong> The SS1 functions as a consistent companion that already knows a child, their preferences, their moods, and their routines. The remote connection features allow parents to stay present without requiring a scheduled call, and the life recording system gives them a passive window into their child’s days that feels less clinical than a monitoring camera.</p><p><strong>Single professionals living alone in cities.</strong> The SS1 adapts to daily routines, builds up a preference model over time, and provides ambient social presence without demands.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Cute home robot with a purple cover and cartoon face displayed on its screen." class="rm-shortcode" data-rm-shortcode-id="bd8ec4a867547204437967b3f231338f" data-rm-shortcode-name="rebelmouse-image" id="8d48b" loading="lazy" src="https://spectrum.ieee.org/media-library/cute-home-robot-with-a-purple-cover-and-cartoon-face-displayed-on-its-screen.jpg?id=67154351&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">OlloNi SS1 adapts to daily routines over time.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><h2>What OlloNi SS1 is doing differently?</h2><p>Ollobot’s goal in building intelligent companion robots is to address the gaps in technology and capability, using innovation not to automate tasks but to fill emotional voids.</p><p>Much of the robotics industry has historically pursued human imitation — machines that speak, look, or behave like people. The SS1 is instead designed around familiarity and long-term coexistence rather than realism.</p><p>The system integrates multiple subsystems operating in parallel, including visual perception, audio processing, mobility control, and interaction management. It is equipped with a multi-chip AI 4K vision module capable of facial recognition and motion tracking. One small but revealing detail is the inclusion of a physical privacy cover for the camera — a mechanical solution to concerns that software settings alone may not fully resolve.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Person playing with a red plush robot toy that has a glowing digital face and eyes" class="rm-shortcode" data-rm-shortcode-id="b39299c23f2f658c0a0284c5803d36ae" data-rm-shortcode-name="rebelmouse-image" id="daa03" loading="lazy" src="https://spectrum.ieee.org/media-library/person-playing-with-a-red-plush-robot-toy-that-has-a-glowing-digital-face-and-eyes.jpg?id=67154349&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">OlloNi SS1 can actively integrate into family activities, and it can autonomously move closer to capture memorable moments or reposition itself to remain engaged in ongoing interactions.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p><span>The robot supports advanced mobility across multiple indoor surfaces, including wooden floors, ceramic tiles, and low-pile carpets, with slope climbing capability up to 3.5 degrees. Rather than remaining in a fixed location, it can move naturally throughout the home to stay close to household members as daily activities unfold. </span></p><p><span>For example, the OlloNi SS1 may greet family members when they arrive home, follow an older adult from the living room to the kitchen while continuing a conversation, remind a child to take a study break after a prolonged period of inactivity, or notice that someone appears unusually quiet and gently check in. During family activities, it can autonomously move closer to capture memorable moments or reposition itself to remain engaged in ongoing interactions.</span></p><p class="pull-quote">The robot continues to evolve over time, with over-the-air updates that deliver new features, performance improvements, and AI enhancements</p><p>It also incorporates fall detection with optimized accuracy for safety monitoring scenarios. A 6-microphone array enables omnidirectional voice pickup with an effective voice capture range of up to 5 meters, supporting reliable wake-word detection and far-field interaction.</p><p>To support continuous companionship, much of the robot’s AI processing takes place directly on the device through its “heart module” architecture, with 16 GB of memory and 64 GB of local storage. This enables the system to retain household memories, recognize familiar faces, and respond with lower latency, making interactions feel more natural even during everyday routines.</p><p>Because companion robots are expected to remain available throughout the day rather than only during brief interactions, the SS1 is designed for extended operation, offering up to 12 hours of standby time and around 5 hours of active interaction on a single charge. This allows it to accompany users through meals, conversations, playtime, and other daily activities without frequent interruptions.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Close-up of toy robot with glowing red heart and purple fur on a beige body." class="rm-shortcode" data-rm-shortcode-id="ad4e47da5ac276c65c47c13fb8cbd4a2" data-rm-shortcode-name="rebelmouse-image" id="0261b" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-toy-robot-with-glowing-red-heart-and-purple-fur-on-a-beige-body.jpg?id=67154317&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">To support engaging interactions, much of the robot’s AI processing takes place directly on the device through its “heart module” architecture.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p><span>Like the relationships it is designed to build, the robot continues to evolve over time. Running on Android OS with over-the-air (OTA) updates, the system continuously receives new features, performance improvements, and AI enhancements, allowing its capabilities to grow alongside the household it serves.</span></p><p>The robot’s behavioral model also improves over time. Rather than reacting to isolated commands, it attempts to establish a baseline understanding of household routines and individuals. Changes in behavior — prolonged quietness, unusual inactivity, or emotional cues — become triggers for interaction.</p><h2>Presence instead of utility</h2><p>Several features in the OlloNi SS1 illustrate this emphasis on presence and continuity in its interactions.</p><p>The system can identify different household members, including pets, and adapt responses accordingly. Remote communication features allow family members to connect through the device without treating every interaction like a scheduled call. Environmental sensors support contextual reminders tied to weather or room conditions.</p><p>Its “2+1” multi-display configuration is also designed around emotional communication. Two circular side displays function as expressive “emotional eyes,” while a separate primary display handles information and structured interaction. The separation allows emotional signaling and functional communication to operate independently, creating more intuitive nonverbal interaction even when no dialogue is taking place.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Cute red robot pet in checkered shirt sits on rug in cozy, warmly lit living room" class="rm-shortcode" data-rm-shortcode-id="f177ff7ea3f02aacf9e12a93a6b7e445" data-rm-shortcode-name="rebelmouse-image" id="5f2a7" loading="lazy" src="https://spectrum.ieee.org/media-library/cute-red-robot-pet-in-checkered-shirt-sits-on-rug-in-cozy-warmly-lit-living-room.jpg?id=67154312&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The robot’s behavioral model improves over time. Rather than reacting to isolated commands, it attempts to establish a baseline understanding of household routines and individuals.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ollobot</small></p><p>The SS1 also includes an automated life-recording system built on facial recognition and behavioral-event detection that can capture moments such as laughter, physical closeness, or group interaction automatically. An integrated AI vlog engine can then organize those moments into edited short-form videos with automated sequencing and soundtrack generation. The design intent is to preserve spontaneous domestic moments without requiring active documentation behavior from users.</p><p class="pull-quote">An integrated AI vlog engine can<span> organize recorded</span><span> moments into edited short-form videos with automated sequencing and soundtrack generation</span></p><p>Visual data is processed primarily on the device through the SS1’s on-device AI architecture, with household memories stored locally and managed within Ollobot’s proprietary ecosystem instead of being shared with third-party smart home platforms. Access to recordings and live feeds is restricted to authorized users through the companion app, while encrypted communication helps protect data during remote access. Users also retain direct control over recording preferences, and the physical camera privacy cover provides an additional hardware-level safeguard whenever visual monitoring is not desired.</p><div class="ieee-sidebar-small"><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Ollobot logo with circular icon and bold lowercase text on light background" class="rm-shortcode" data-rm-shortcode-id="c161d66f60be15a6d25fa8fcac7ad1d7" data-rm-shortcode-name="rebelmouse-image" id="8eca7" loading="lazy" src="https://spectrum.ieee.org/media-library/ollobot-logo-with-circular-icon-and-bold-lowercase-text-on-light-background.png?id=67154360&width=980"/></p><p>Learn more at <a href="https://ollobot.com/" target="_blank">ollobot.com</a>.</p></div><p>Remote communication is similarly structured around persistence rather than transaction. Traditional video calls are episodic and screen-bound; the SS1 instead acts as a continuously present interface embedded inside the household environment. Through autonomous mobility, environmental awareness, and persistent household memory, remote family members interact with an ongoing domestic context.</p><h2><a target="_blank"></a>The larger shift to “gentle intelligence”</h2><p>Ultimately, gentle intelligence is not about making robots behave more like humans — it is about helping them fit more naturally into human lives. Each OlloNi SS1 unit develops a unique behavioral profile based on its household. Two units running in different homes for a year will have become meaningfully different from each other, shaped by the specific people, habits, and rhythms of where they live.</p><p>That kind of long-term personalization is what early companion robots never had. It is also what makes the difference between a product that ends up on a shelf and one that actually earns its place in a home.</p><p>Learn more at <a href="https://ollobot.com" target="_blank">ollobot.com</a>.</p>]]></description><pubDate>Tue, 25 Aug 2026 10:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ollobot-ai-companion-robot</guid><category>Social-robots</category><category>Human-robot-interaction</category><category>Computer-vision</category><category>Companion-robots</category><category>Multimodal-ai</category><category>Ai-robots</category><dc:creator>Ollobot</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/cute-home-robot-on-carpet-in-cozy-living-room-with-beige-sofa-and-warm-lighting.jpg?id=67154308&amp;width=980"></media:content></item><item><title>Self-Driving Cars Could Someday Take Requests</title><link>https://spectrum.ieee.org/autonomous-vehicles-motion-planner-llm</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/diagram-of-a-vehicle-merging-into-an-adjacent-lane-with-surrounding-traffic-with-predicted-motions-labelled.jpg?id=67659079&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>
<em>This article is part of our exclusive <a href="https://spectrum.ieee.org/collections/journal-watch/" target="_blank">IEEE Journal Watch series</a> in partnership with IEEE Xplore.</em>
</p><p>The idea of letting a machine do the driving for you may put a lot of people off <a data-linked-post="2668536834" href="https://spectrum.ieee.org/autonomous-vehicles-great-at-straights" target="_blank">autonomous vehicles</a>. But research could make it possible to backseat-drive an autonomous vehicle just as you might with a human driver.</p><p><a data-linked-post="2671006250" href="https://spectrum.ieee.org/darpa-grand-challenge" target="_blank">Self-driving cars</a> carefully balance a host of parameters to ensure a smooth ride, including things like speed, acceleration, and the smoothness of turns. But human driving preferences can often vary depending on how much of a rush they’re in, whether they’re feeling carsick, or how busy the traffic is.</p><p>These cars have a software component called the motion planner, which is responsible for choosing a safe and efficient path through traffic. The motion planner is normally tuned by engineers before the vehicles hit the road so that there’s little scope for passengers to adjust a vehicle’s driving style on the fly. But now researchers at the Delft University of Technology (TU Delft) in the Netherlands have developed a system that uses a large language model (LLM) to translate natural-language user requests such as “I am running late, go fast” into adjustments to a self-driving control system. The researchers <a href="https://arxiv.org/abs/2606.10974v1" rel="noopener noreferrer" target="_blank">posted their preprint on arXiv</a> and are presenting the work at the <a href="https://ieee-itss.org/conf/itsc/" rel="noopener noreferrer" target="_blank">IEEE Intelligent Transportation Systems Conference</a> in September.</p><h2>LLMs Personalize Autonomous Driving</h2><p>The system doesn’t give users direct control over the vehicle’s driving decisions; it simply tunes the parameters of a safety-aware motion-planning algorithm, which helps to keep the vehicle’s behavior within safe bounds. And the system keeps the human in the loop by describing how it’s going to alter its behavior in nontechnical language, and by asking the passenger to confirm before making changes. When the system was tested in simulation, the researchers found it adjusted the speed and smoothness of driving in line with natural-language instructions.</p><p>“The motion-planning problem is not only about reaching a place while avoiding collisions, it’s also how you do it,” says lead author <a href="https://sites.google.com/view/dmartnezbaselga" rel="noopener noreferrer" target="_blank">Diego Martinez-Baselga</a>, a postdoctoral researcher at TU Delft. “The motivation here is trying to make the way the autonomous car drives adaptable by end users easily, just by talking to the car.”</p><p>Previous research has investigated the potential of using LLMs and video-language models (VLMs) to direct decision-making for self-driving vehicles, but the researchers deliberately targeted driving style instead. Using LLMs and VLMs to directly control vehicles faces several challenges, says Martinez-Baselga. These include relatively slow response times, which can make these models unsuitable for the fast-paced decision-making required in driving, and the fact that they can’t provide concrete performance guarantees in the way a deterministic motion planner can.</p><p>Instead, the researchers used an LLM’s language and reasoning capabilities to translate fuzzy human preferences into something a vehicle’s motion planner can use. The system relies on a model predictive-path integral controller previously developed by the researchers, which identifies multiple paths the vehicle could take to reach its goal and then judges them on various criteria, including speed, steering angle, and collision probability. It then finds an optimal path that is a combination of the trajectories that scored best on those judging criteria.</p><p>The team combined this with OpenAI’s GPT-4o-mini model to parse passengers’ natural-language suggestions and use them to tune how the controller chooses its path. The model is given the users’ prompt and a natural-language description of the scenario the vehicle is operating in. The description was handwritten by the researchers for the purposes of the study, but it could ultimately be provided directly by a car’s perception system, says Martinez-Baselga.</p><p>The model doesn’t directly tweak the settings of the controller; it uses the prompt to rate the relative importance of the judging criteria the controller uses to assess trajectories. This rating is then used to adjust each criteria up or down either side of a safe baseline set by the researchers. So, if a user says they are feeling dizzy, the LLM will dial up parameters that encourage smooth steering and gentle acceleration to make the vehicle favor more sedate travel.</p><p>Prior to making any changes, however, the model first presents the user with a natural-language description of the adjustments it plans to implement. The user can then sign off on the plan or make further suggestions. The system is also interactive, so the user can request further adjustments if the vehicle’s behavior doesn’t match expectations or the user‘s preferences change.</p><p>Martinez-Baselga says this human-in-the-loop system allows the passenger to catch instances when the model misinterprets prompts. But it also helps deal with the inherent subjectivity of suggestions like “go faster” or the possibility that models don’t accurately describe changes they plan to make. In that case the passenger can simply follow up with additional prompts “as you would do if you were in a taxi or with a friend that is driving,” says Martinez-Baselga.</p><p>The researchers tested the system in the popular <a href="https://www.nuplan.org/nuplan" rel="noopener noreferrer" target="_blank">self-driving simulator nuPlan</a> in scenarios that involved merging onto a busy highway. Across eight different prompts, the system changed the controller’s parameters in ways matching user intent, with requests for a more comfortable ride dialing up smoothness and those indicating urgency leading to higher speeds.</p><p>This isn’t the first time LLMs have been used to tune a self-driving car’s motion planner. <a href="https://ee.ethz.ch/the-department/people-a-z/person-detail.MjE0NjI3.TGlzdC8zMjc5LC0xNjUwNTg5ODIw.html" rel="noopener noreferrer" target="_blank">Nicolas Baumann</a>, a Ph.D. student at ETH Zurich in Switzerland, <a href="https://arxiv.org/abs/2504.11514" rel="noopener noreferrer" target="_blank">published research</a> last year in which an LLM tweaked the parameters of a model racing-car controller, allowing the user to alter driving style but also give more concrete instructions like “reverse the car” or “maintain a specific speed.”</p><p>The strength of the approach, says Baumann, is that separating the LLM from the main controller means that even if the model hallucinates, it can’t do anything dangerous. “You get the possibility of language interaction, but you can guarantee that it is going to be within the constraints of this classical controller, so you can bake in safety,” he says. However, setting these constraints requires considerable engineering work, he adds.</p><p>And if you want provable safety, you need to go a step further, says <a href="https://www.professoren.tum.de/en/althoff-matthias" rel="noopener noreferrer" target="_blank">Matthias Althoff</a>, a professor of cyberphysical systems at the Technical University of Munich. His group <a href="https://ieeexplore.ieee.org/document/11640900" rel="noopener noreferrer" target="_blank">built a system</a> that gets an LLM to suggest driving decisions, but then uses a mathematical process to check them against traffic rules and predictions about the behavior of other road users. This makes it possible to verify their safety before committing to them, something the Delft paper doesn’t provide. “As with any LLM, it is not guaranteed that the result is correct,” says Althoff. “For that reason, we safeguard the decisions of the LLM in our works.”</p>]]></description><pubDate>Mon, 24 Aug 2026 15:51:45 +0000</pubDate><guid>https://spectrum.ieee.org/autonomous-vehicles-motion-planner-llm</guid><category>Autonomous-vehicles</category><category>Journal-watch</category><category>Large-language-models</category><dc:creator>Edd Gent</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/diagram-of-a-vehicle-merging-into-an-adjacent-lane-with-surrounding-traffic-with-predicted-motions-labelled.jpg?id=67659079&amp;width=980"></media:content></item><item><title>What It Takes to Be an Adaptable Engineer</title><link>https://spectrum.ieee.org/adaptable-engineer-core-skills</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/silhouette-of-person-at-computer-atop-a-weather-vane-labeled-ai-against-sky-background.png?id=67652283&width=1245&height=700&coordinates=0%2C251%2C0%2C252"/><br/><br/><p>The AI boom has disrupted the way engineers work, introducing new tools to learn, raising expectations for what teams can achieve in a workday, and making it <a href="https://spectrum.ieee.org/technical-interview-ai-arms-race" target="_self">harder to get hired</a> in the first place. This makes it difficult to advise students on <a href="https://spectrum.ieee.org/top-programming-languages-2025" target="_self">which specific coding languages</a> or technical skills they should learn. So amidst the uncertainty, advice for young professionals often turns to a common refrain: Be adaptable. But what does adaptability look like in practice? </p><p> Engineers often operate on the cutting edge of technology, so dealing with change is a normal part of the job, says <a href="https://search.asu.edu/profile/2452662" rel="noopener noreferrer" target="_blank">Samantha Brunhaver</a>, an associate professor of engineering at Arizona State University, in Tempe. Yet university curricula and training in the workplace often don’t prepare students for this. </p><p>“We tell engineers that they need to be adaptable when they graduate, but we don’t actually explain what that means, demonstrate what that looks like, [or] help make sure that they’re developing it,” says Brunhaver, who <a href="https://engineering.asu.edu/spring2020/brunhaver-seeking-to-augment-engineering-education-with-an-adaptive-mindset/" rel="noopener noreferrer" target="_blank">received a National Science Foundation award in 2020</a> to study how to foster greater workplace adaptability among young engineers. For this ongoing project, she has interviewed engineering managers, early career employees, and undergraduates about their experiences. </p><p> Part of the problem, she says, is that every employer has its own idea of what to be adaptable means. Generally, Brunhaver defines adaptability as “the ability to recognize that a change or uncertainty is occurring, and then respond effectively to that change.” But the skill is context-dependent. In software engineering, that might mean responding to turnover in the tools you use on a daily basis, while aerospace or biomedical engineers may need to keep track of changing procedures and regulations. “Managers are all saying adaptability is important,” Brunhaver says, “but defining it in different ways.” </p><p> At the same time, engineers are all contending with changes beyond these industry-specific expectations. Jobs in the technology, media, and telecom sectors are experiencing the fastest pace of skill turnover, according to a <a href="https://www.pwc.com/gx/en/issues/artificial-intelligence/job-barometer/2026/pwc-aijb-2026-technology-media-and-telecoms-report.pdf" rel="noopener noreferrer" target="_blank">June 2026 report</a> on the effects of AI from the professional services network <a href="https://www.pwc.com/us/en.html" rel="noopener noreferrer" target="_blank">PwC</a>. And the World Economic Forum’s most recent <a href="https://reports.weforum.org/docs/WEF_Future_of_Jobs_Report_2025.pdf" rel="noopener noreferrer" target="_blank"><em><em>Future of Jobs Report</em></em></a>, published in 2025, found that employers across all sectors expect 39 percent of workers’ core skills to change by 2030. This uncertainty can be uncomfortable. But with the right mind-set and support from leadership, adaptability can help keep you afloat. </p><h2>How to Cultivate Adaptability</h2><p>The AI transition is a big shift—but not an unprecedented one, says <a href="https://www.microsoft.com/en-us/research/people/jennbu/" rel="noopener noreferrer" target="_blank">Jenna Butler</a>, a research scientist at Microsoft who studies developer well-being and productivity. </p><p>During this type of paradigm shift, there is often a “chaos period” when a new normal is being established, Butler says. In AI’s case, it challenges the understanding of what a computer can do. “I think we’re still in this in-between, difficult period that we’ve seen before, but [it] is maybe moving faster than it has historically.” Software engineers—in one of the fields <a href="https://www.theguardian.com/technology/ng-interactive/2026/jul/12/software-developers-engineers-ai" rel="noopener noreferrer" target="_blank">most affected by AI</a>—are now facing a significant increase in code review. “If you ask 20 developers, you get 23 different ways of working with it. Everyone is trying to sort it out,” says Butler, who describes this period as “the uncomfortable middle.” </p><p class="pull-quote">“We tell engineers that they need to be adaptable when they graduate, but we don’t actually explain what that means, demonstrate what that looks like, [or] help make sure that they’re developing it.”<span><strong>– Samantha Brunhaver, Arizona State University</strong></span></p><p><span></span>Brunhaver says one way educators can help prepare students before they enter the workforce is by offering a diversity of real-world experiences, such as internships, team-based projects, community service, and leadership roles. Each of these teach students to adapt to different challenges, easing their transition from school to work. </p><p> It’s also important to encourage reflection, Brunhaver adds, noting that metacognition helps individuals use the skill more effectively. “In order to adapt, you have to think that you have agency and the ability to get through a situation.” Ultimately, it comes down to three steps: Perceive a need to adapt, evaluate your options, and act. </p><p> For those already in the workforce, that action may mean taking the time to learn new tools and ways of working. Software engineering, for instance, may soon rely more on <a href="https://spectrum.ieee.org/chain-of-thought-prompting" target="_self">prompting models</a> and managing agents than coding line by line. “I think people who went into software because they like solving problems are going to have a lot of fun, and people who just enjoy the art of writing code are not,” Butler says. </p><h2>The More Things Change…</h2><p> Although the tools engineers use on a daily basis are evolving, the core responsibilities of the job are more stable than they may seem, says <a href="https://toolshed.com/about.html" target="_blank">Andy Hunt</a>, a software developer who coauthored <a href="https://pragprog.com/titles/tpp20/the-pragmatic-programmer-20th-anniversary-edition/" target="_blank"><em><em>The Pragmatic Programmer</em></em></a> (Addison-Wesley Professional) in 1999. The book outlines practical coding principles, and has been taught in many computer science classrooms. When Hunt was working on the 20th anniversary edition of the book, he was surprised by how much of the advice still applies. And now, seven years later, he maintains that belief.</p><p>“The fundamental part of the job is problem solving and communication, and that’s always going to be there,” he says. </p><p> Hunt emphasizes the importance of developing systems thinking over particular tools. To him, identifying as a Java programmer, for instance, is “like a carpenter saying, ‘I’m a hammer user,’ or ‘I specialize in cordless drills.’ ” </p><p> He acknowledges that today’s hiring process, in which companies often <a href="https://mckelveyconnect.washu.edu/blog/2022/02/04/8-things-you-need-to-know-about-applicant-tracking-systems/" rel="noopener noreferrer" target="_blank">filter résumés</a> for certain languages or years of experience, makes it harder to embrace a more expansive way of relating to your job. Employers, he says, should recognize that “the tech’s not the hard part, and it never has been. Understanding information theory, understanding systems thinking, understanding what constraints you’re up to—that’s still the hard part.” </p><p> With this type of misalignment between employers and employees, AI is also intensifying an old source of tension: How can engineers slow down enough to adapt and learn new tools when the pressure to become more productive keeps mounting? </p><h2>Who’s Responsible for Enabling Change? </h2><p><span>Young engineers need to embrace change. However, educators and employers also play a role in building a successful workforce. From the educator’s perspective, Brunhaver says “we need to be more explicit about what [adaptability] means and why it’s important.” Managers, meanwhile, should invest in their employees’ professional development.</span></p><p>Microsoft research scientist Butler often encourages leadership to set aside intentional time for continuous learning for their engineers—even just an hour a week—without any expectation that they will produce code or progress in their daily work. “I realize that’s difficult,” says Butler. “I would encourage people to do it on their own, but I would really encourage organizations and leaders to do it, because you’re not going to get this sudden change in your people if they don’t have time and space to learn how to work differently.” </p><p>This also means providing enough instruction, Butler adds. When developers aren’t given enough guidance on adopting something new, while being pressured to increase productivity, they risk <a href="https://arxiv.org/pdf/2507.21280" target="_blank">doubling down on the tools they already know and burning out</a>. </p><p>“I do imagine the next number of years could be challenging,” Butler says. Engineers will have to adapt to find their place in an evolving workforce—but they also have a say in shaping that future. </p><p>“Being adaptable sort of implies that you’re going to change based on what’s happening around you, and I would really like people to realize the change that’s happening is somewhat up to us,” she says. All individuals have a choice in how they use AI, for instance, and which models they use. “We need to be adaptable and go with the flow to a degree, but we also need to be directing that flow. The future with AI is absolutely not predetermined.”</p><p><em>This article appears in the September 2026 print issue as “The Adaptable Engineer.”</em></p>]]></description><pubDate>Mon, 24 Aug 2026 14:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/adaptable-engineer-core-skills</guid><category>Typedepartments</category><category>Adaptive-learning</category><category>Professional-development</category><category>Ai-tools</category><category>Future-of-work</category><dc:creator>Gwendolyn Rak</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/silhouette-of-person-at-computer-atop-a-weather-vane-labeled-ai-against-sky-background.png?id=67652283&amp;width=980"></media:content></item><item><title>This IEEE Senior Member Develops AI Tools for E-Commerce Sites</title><link>https://spectrum.ieee.org/ieee-senior-member-ai-ecommerce</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-professionally-dressed-indian-man-smiling-with-a-trophy-in-one-hand-and-an-award-certificate-in-the-other.jpg?id=67657769&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><a href="https://balajiingole.com/" rel="noopener noreferrer" target="_blank">Balaji Ingole</a> rarely saw televisions while growing up in Udgir, India. No one in the small Maharashtra village had computers or phones. Only one household owned a television, and neighbors often gathered there to watch shows together.</p><p>Ingole never even saw a computer growing up. It wasn’t until he reached middle school that he encountered a computer lab, an experience he says changed his life. Almost immediately, he says, the machine felt like a window into a different scale of possibility for him.</p><h3>Balaji Ingole</h3><br/><p><strong>Employer </strong></p><p><strong></strong>Amla Commerce in Milwaukee</p><p><strong>Title </strong></p><p><strong></strong>Project manager</p><p><strong>Member grade</strong> </p><p>Senior member</p><p><strong>Alma maters </strong></p><p><strong></strong>COEP Technological University and Welingkar Institute of Management, both in India</p><p>“I was very studious and not very social, always reading or solving problems in a math textbook,” he says. “At the computer lab, I began learning the <a href="https://www.c-language.org/" rel="noopener noreferrer" target="_blank">C programming language</a>—which was like discovering a whole new world. I was fascinated that you could create something with just a few lines of code.”</p><p>His early interest grew into a self-directed education. Outside of Ingole’s formal classwork, he taught himself to build database-backed applications, wire up hardware, write software, and trace error logs.</p><p>Today the IEEE senior member similarly splits his time. During the week, he’s a project manager in Milwaukee at B2B e-commerce company <a href="https://www.amla.io/" rel="noopener noreferrer" target="_blank">Amla</a>, leading AI-driven digital transformation initiatives to help the company’s clients boost their sales. On weekends, he leads a similarly demanding life as an independent researcher. His current projects include developing AI-enabled health care diagnostic tools and assistive technologies to support people with physical disabilities.</p><p>“I believe in ‘learn by doing,’” he says. “I really like to test my knowledge and prototype ideas to find out if they truly work.”</p><h2>A college project becomes an inspiration</h2><p>Ingole’s tendency to go beyond his coursework continued after he graduated high school in 2004. As a mechanical engineering undergraduate at <a href="https://www.coeptech.ac.in/" rel="noopener noreferrer" target="_blank">The College of Engineering, Pune (now COEP Technological University)</a>, in India, he participated in several extracurricular activities. One was interviewing entrepreneurs and writing about them for <em><em>The COEP College Magazine</em></em>. The experience helped him gain confidence, he says, giving him the push he needed to pursue interviews for the publication with two Indian entrepreneurs he admired: <a href="https://www.infosys.com/about/management-profiles/narayana-murthy.html" rel="noopener noreferrer" target="_blank">N.R. Narayana Murthy</a>, cofounder of IT giant <a href="https://www.infosys.com/" rel="noopener noreferrer" target="_blank">Infosys</a>; and his wife, philanthropist <a href="https://sudhamurty.in/" rel="noopener noreferrer" target="_blank">Sudha Murty</a>. The Murtys cofounded the <a href="https://www.infosys.org/infosys-foundation.html" rel="noopener noreferrer" target="_blank">Infosys Foundation</a>, a nonprofit that runs educational, health care, women’s empowerment, and sustainability <a href="https://www.infosys.org/infosys-foundation/initiatives.html" rel="noopener noreferrer" target="_blank">programs</a> in underserved areas of India.</p><p>“Every week I would fax them: ‘Please give me an interview time,’” Ingole says. Eventually, Sudha Murty’s office offered him a phone interview, but he requested to meet her in person at Infosys’s Bengaluru offices. She agreed, but the offices were 940 kilometers from Pune, and he didn’t have the money to travel or stay overnight in a hotel.</p><p>Ingole and a classmate borrowed money from friends and traveled through the night on multiple buses and trains to get to Bengaluru. They freshened up in a public bathroom before heading to the Infosys campus to meet Murty.</p><p>Impressed by their persistence, she surprised them by also arranging a brief chat with Narayana Murthy.</p><p>“Narayana Murthy handwrote a personal message to the engineering students of [my college]—which we proudly published in our college magazine,” Ingole says. “In his note, Murthy shared that we are at an extraordinary moment in India’s history and that the future looks even brighter. His words encouraged us to work hard and make the most of this time.</p><p>“I still have that note,” Ingole says. “They are billionaires, and I was just a regular student. The fact that they took the time to do this really motivated me.”</p><p>During the final semester of his engineering studies, Ingole joined the <a href="https://in.linkedin.com/company/tata-research-development-and-design-centre-trddc" rel="noopener noreferrer" target="_blank">Tata Research Design and Development Center</a> in Pune for a six-month internship. After earning his bachelor’s degree in mechanical engineering in 2008, he became a graduate engineering trainee at <a href="https://automation.honeywell.com/us/en" rel="noopener noreferrer" target="_blank">Honeywell Automation</a> in Pune.</p><p>He left the company in 2009, and during the next 13 years, he held different software engineering and project management positions at IT companies across India.</p><p>He earned a master’s degree in business administration from the <a href="https://www.welingkar.org/" rel="noopener noreferrer" target="_blank">Welingkar Institute of Management, Mumbai,</a> in 2017.</p><p>In 2022 he accepted a project-manager role at <a href="https://marssg.com/" rel="noopener noreferrer" target="_blank">Mars IT Solutions in Madison, Wisc.</a> The following year, he left to join <a href="https://www.gainwelltechnologies.com/" rel="noopener noreferrer" target="_blank">Gainwell Technologies</a>, also in Madison, as a senior project manager. At Gainwell, he managed projects for the core IT systems multiple U.S. state health departments use to administer <a href="https://www.medicaid.gov/" rel="noopener noreferrer" target="_blank">Medicaid</a> benefits, manage provider enrollment, and verify member eligibility. The experience managing projects that directly enabled patients’ access to health care gave Ingole a special appreciation for and interest in this area, he says.</p><p>“Health care data is not like other data,” he says. “The stakes are high, compliance requirements are different, and the margin for error is effectively zero.”</p><p>Ingole says he enjoyed the rigor of data governance combined with the potential to positively impact lives, and that also applies to his current work at Amla.</p><h2>Agentic AI in e-commerce</h2><p>Ingole joined Amla in July 2025. He helps manufacturers and B2B customers modernize their <a data-linked-post="2650248208" href="https://spectrum.ieee.org/who-invented-ecommerce" target="_blank">e-commerce</a> operations. He also builds AI tools for them and for his internal team.</p><p>For Amla’s customers, he’s developing AI-enabled chatbots that help manufacturers set up and manage large product catalogs in e‑commerce platforms. Such product setup traditionally has been a manual, tedious, error-prone process: Companies upload thousands of products, adjust item names, enter prices, update images, and more.</p><p>“Product setup has been one of the most painful processes in e-commerce, and it can take [our] customers two to three months to complete,” Ingole says. “We’re creating an AI agent that will guide them, step-by-step, to get everything set up in two weeks.”</p><p>Ingole relies on AI agents for some of his own tasks at Amla. Project managers historically have spent 10 to 12 hours each week assembling and sending status reports to stakeholders. Ingole built an AI agent to handle much of the work.</p><p>“It runs every Monday morning and reads through my emails to extract highlights, risks, timelines, and upcoming releases, then sends me a written status report,” Ingole says. The process might sound simple, but the agent’s workflow involves at least a dozen steps including defining parameters, managing temporary files, and integrating with existing tools.</p><p>With the information-gathering work handled, it frees up Ingole and his colleagues to spend more time on deeper-thinking work, he says.</p><h2>Publishing as idea refinery</h2><p>For nearly a decade, Ingole has spent some of his free time conducting independent research projects in data analytics and AI-enabled applications in health care. He has written <a href="https://ieeexplore.ieee.org/author/790405442038687" rel="noopener noreferrer" target="_blank">more than 40 peer-reviewed papers</a>, which are in the <a href="https://spectrum.ieee.org/free-access-to-thousands-of-covid19-research-documents" target="_self">IEEE Xplore Digital Library</a>. He has been granted six patents in the United Kingdom and India. In the U.K., he is a registered coinventor of an AI-powered, cloud-connected <a href="https://www.researchgate.net/publication/388234716_AI-POWERED_CLOUD-CONNECTED_WEARABLE_DEVICE_FOR_PERSONALIZED_HEALTH_MONITORING" rel="noopener noreferrer" target="_blank">wearable device for health monitoring</a> and an <a href="https://www.researchgate.net/publication/389874948_AI-BASED_BREAST_CANCER_DETECTION_DEVICE" rel="noopener noreferrer" target="_blank">AI-based breast cancer detection tool</a>.</p><p>Ingole’s patent for the breast cancer detector, he says, reflects his belief that when engineers apply data and AI correctly, they can help doctors diagnose patients more quickly and accurately.</p><p>That, he says, is both a power and a responsibility.</p><p>He is part of a team helping patients who are paralyzed and nonverbal control items in their environment. His goal, he says, is to develop a brain-computer interface to let patients turn on a fan, switch off a television, and complete similar tasks.</p><p>Publishing research requires both academic rigor and peer scrutiny, and Ingole says the function has been critical to improving as both a project manager and a researcher-inventor.</p><p>“Lots of research ideas never make it to paper,” he notes. “But when you write for journals or conferences, you’re bombarded with questions from Ph.D.s and experienced researchers. This forces me to refine my methodology, and to combine use cases and technical architecture in a way that stands up to expert review.”</p><h2>Finding a professional hub</h2><p>Ingole joined IEEE in 2022, and he says the affiliation has become central to both his research and his professional identity.</p><p>“I use the <a href="https://cis.ieee.org/activities/membership-activities/ieee-member-directory" rel="noopener noreferrer" target="_blank">Member Directory</a> often and contact engineers through my IEEE email address, which gives me credibility because they know it’s a genuine research connection,” he says.</p><p>The organization has given him a platform to contribute to the research space beyond his own papers, he says. He has served as a conference session chair, keynote speaker, technical program committee member, and peer research reviewer for various conferences and events. His IEEE membership, he says, has opened doors to other communities, helping support his entry into the <a href="https://www.bcs.org/" rel="noopener noreferrer" target="_blank">British Computer Society</a>, which has stringent acceptance criteria.</p><p>Those opportunities have helped him build a global network of collaborators with whom to discuss upcoming research, seek advice, and share data, he says.</p><p>“IEEE is important for me to continue as an independent researcher,” he says. “It lets me contribute to the community, and I get a lot in return.”</p>]]></description><pubDate>Fri, 21 Aug 2026 18:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ieee-senior-member-ai-ecommerce</guid><category>Ieee-member-news</category><category>Artificial-intelligence</category><category>E-commerce</category><category>Careers</category><category>Type-ti</category><dc:creator>Julianne Pepitone</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-professionally-dressed-indian-man-smiling-with-a-trophy-in-one-hand-and-an-award-certificate-in-the-other.jpg?id=67657769&amp;width=980"></media:content></item><item><title>Stop Hunting, Start Solving: Accelerating Root Cause Analysis with Agentic AI</title><link>https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&amp;utm_medium=site&amp;utm_campaign=922</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/spotfire-logo-with-circular-icon-and-stylized-black-text.png?id=67657308&width=980"/><br/><br/><p><strong>About this Webinar</strong></p><p><strong>Turn Yield Excursions into Faster, More Confident Root Cause Analysis</strong></p><p>When a yield issue emerges, the answer rarely lives in a single system. Critical clues are spread across metrology data, tool traces, chemical analysis, and facilities systems, while growing data volumes make traditional dashboards slow, fragmented, and difficult to act on.</p><p><strong>What You’ll Learn:</strong></p><p>Discover how a purpose-built semiconductor analytics platform can help engineers connect insights across domains without moving data. See how Agentic AI, semiconductor-specific visualizations, and push-down compute enable faster investigation of yield excursions and process issues, even across billions of data points. In the session, a live demonstration shows how to conduct a multi-domain root cause investigation using Spotfire® Industry Pro.</p><p><strong>Key Takeaways:</strong></p><ul><li>Understand why siloed manufacturing data delays yield recovery and inflates costs</li><li>Learn how Agentic AI automates complex cross-domain analytics and visualization generation</li><li>Explore methods for scaling high-performance analytics across massive fab datasets</li></ul><p><strong>Who Should Attend:</strong></p><p><span>Yield, Process, and Integration Engineers; Fab and Manufacturing Operations Managers; Quality and Reliability Engineers; and Data & Analytics leaders supporting wafer fabs, foundries, OSATs, and IDMs who need to identify issues faster while maintaining confidence in decision-making.</span></p><p><strong>Save Your Spot!</strong></p><p><span></span><span>Join this webinar to learn how leading semiconductor teams are accelerating root cause investigations, scaling analytics across massive datasets, and transforming disconnected data into actionable manufacturing intelligence. Reserve your seat today.</span></p><p><span><span><a href="https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&utm_medium=site&utm_campaign=922" target="_blank">Register now for this free webinar!</a></span></span></p>]]></description><pubDate>Fri, 21 Aug 2026 14:32:37 +0000</pubDate><guid>https://event.on24.com/wcc/r/5460332/DAFEFF7A68EE900DEA7A14356089B553?utm_source=IEEE&amp;utm_medium=site&amp;utm_campaign=922</guid><category>Type-webinar</category><category>Semiconductor-manufacturing</category><category>Agentic-ai</category><category>Root-cause-analysis</category><category>Yield-analytics</category><category>Fab-operations</category><dc:creator>Spotfire</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67657308/origin.png"></media:content></item><item><title>From AI Copilots to Agent Swarms</title><link>https://spectrum.ieee.org/amd-agent-swarms</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/colorful-3d-blocks-piled-behind-glass-panels-displaying-white-code-snippets.jpg?id=67609515&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><span>The impact of AI on software development has been both profound and ever-evolving. Last year, I </span><a href="https://spectrum.ieee.org/beyond-code-autocomplete" target="_self">wrote</a><span> about <a href="https://www.amd.com/en.html" target="_blank">AMD’s</a> plans to use AI not just for </span><a href="https://spectrum.ieee.org/best-ai-coding-tools" target="_self">generating</a><span> new lines of code, but also for other steps in the software development lifecycle (SDLC), such as triaging problems, debugging code, and testing the software. At the time, we were hoping for a 25 percent productivity boost from AI use over the course of two or three years.</span></p><p>But with each new release, the capabilities of large language models (LLMs) improve dramatically—accelerating software development, increasing the quality of AI-generated code, and fundamentally reshaping how software is engineered. Now, just one year later, we have surpassed our productivity target, achieving a 30 percent overall productivity boost through AI. On top of that, we are rethinking not only how we use AI within the SDLC, but the structure of the SDLC itself.</p><p>We believe that the biggest AI revolution in software engineering is still ahead. So far, we have largely been teaching AI how we perform tasks and asking it to mimic existing workflows. In many ways, this constrains AI to human patterns of thinking. The next transformation will come from collaborative swarms of AI agents capable of discovering solutions independently.</p><h2>Agents of today</h2><p>AMD began developing AI systems for code generation, testing automation, bug analysis, and code review in 2024. At the time, our objective was to achieve 25 percent AI-generated production code by 2027 while gradually automating larger portions of the SDLC.</p><p>Measuring productivity is inherently <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" target="_blank">challenging</a>, but from the outset we have consistently tracked one objective metric: the percentage of source code generated by AI. Importantly, we count only code that passes all reviews and testing and is ultimately included in the final product. While AI-generated code is certainly not the only contributor to productivity gains, it is one of the few metrics that can be measured objectively and consistently.</p><p>By this metric, we have crossed the 20 percent mark at the beginning of this year and are now progressing towards 50 percent across entire codebase. In some software components, more than 80 percent of the code is now generated using AI.</p><p>Agentic AI has enabled us to include AI in every step of the life cycle: For <strong>code analysis and triage</strong>, agents are trained to analyze problem reports, identify and group similar requests, and highlight which code snippets are likely to need modification. For <strong>debugging and code generation</strong>, agents are directed to analyze a bug request and implement required code changes. For <strong>testing</strong>, the agents generate unit tests, and if those are passed, identify necessary integration and product-level tests. And finally, for the <strong>approval and release</strong> stage, agents prepare architecture summary, code change review, and full test results for engineers’ review and approval—and, if approved, integrate the changes into the next release.</p><h2>Agents of tomorrow</h2><p>Today, engineers create AI agents in their own image: They teach AI what they know about the system, how they would fix an issue, and how they would implement a change. This is already a major technological advancement. Engineers can create multiple “AI versions” of themselves, allowing these agents to work in parallel, scaling their expertise far beyond the limits of individual productivity. The limitation, however, is that these AI agents are still constrained by human thinking and human-defined approaches.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Person typing on a keyboard. Looming in front of them is a network of colorful AI agents surrounding a glowing central node." class="rm-shortcode" data-rm-shortcode-id="0d892f7d0215c3159d9e41db8419441a" data-rm-shortcode-name="rebelmouse-image" id="73c7c" loading="lazy" src="https://spectrum.ieee.org/media-library/person-typing-on-a-keyboard-looming-in-front-of-them-is-a-network-of-colorful-ai-agents-surrounding-a-glowing-central-node.jpg?id=67609528&width=980"/> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">AMD</small></p><p>We believe the next major transformation in software engineering will occur when collaborative AI agent swarms can independently identify and develop solutions, guided by humans on what to solve rather than constrained by human assumptions about how the job should be done. Instead of providing detailed instructions on how to solve a problem, engineers will define the issue, the desired outcome, and the quality, performance, and system constraints, allowing AI agents to determine the optimal path to a solution.</p><p>A swarm of AI agents will then work in parallel to generate, evaluate, and refine multiple solution approaches. These agents will automatically validate correctness, measure performance, test trade-offs, and compare alternative implementations against defined success criteria. Finally, AI agents will prepare ranked solution options, along with validation results and performance metrics, for engineer review and approval. The agents won’t be enhancing each step of the SDLC—they will be rewriting the SDLC themselves.</p><p>To get to this point, we need to change how agents are trained. Today, improvement occurs one engineer and one agent at a time: An engineer reviews the output, refines the prompt, and repeats the process. To scale beyond this model, agents must continuously learn from one another, reuse successful strategies, and improve collaboratively across projects and teams.</p><p>We are already moving in this direction by using multi-agent workflows extensively through agentic harnesses, such as Codex and Claude Code, while simultaneously developing our own internal multi-agent systems to support the next generation of AI-driven software engineering.</p><p>A good example is our AI-driven effort to resolve issues in our Radeon Software eXperience (RSX). RSX is a user interface component that allows users to configure and monitor graphics driver behavior. In October 2025, we began using AI agents to automatically debug and fix reported RSX issues. Out-of-the-box AI tools delivered limited results, resolving only 6 percent of issues.</p> <p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A bar chart labelled RSX-Agentic Resolution Rate shows large percentage increases from 6 in October 2025 to 75 in June 2026." class="rm-shortcode" data-rm-shortcode-id="98d6334ebf612596d6de9b5137c39de6" data-rm-shortcode-name="rebelmouse-image" id="d27a2" loading="lazy" src="https://spectrum.ieee.org/media-library/a-bar-chart-labelled-rsx-agentic-resolution-rate-shows-large-percentage-increases-from-6-in-october-2025-to-75-in-june-2026.jpg?id=67609523&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The percentage of software issues fixed automatically by AI agents in AMD’s Radeon Software eXperience (RSX) has been growing steadily, reaching 75 percent in June 2026. </small></p> <p class="media-body"><span>As we analyzed failures and identified ways to improve, we built a learning loop—initially a largely manual process—to understand where the agents were falling short and how to improve them. Rather than retraining the underlying models, we refined the objectives given to the agents, allowing them to iteratively explore multiple approaches, evaluate the results against defined success criteria, and converge on better solutions. At the same time, advances in models and agent run-times further increased effectiveness. Together, these improvements significantly increased our resolution rate from 6 percent to more than 75 percent of RSX issues resolved by agentic loop.</span></p><p>To make agents and agent swarms truly productive, we need a continuous learning loop that feeds errors and human interventions back into future agent workflows. The opportunity is to engineer this loop around clear, measurable goals. Each cycle captures new insights, making the entire AI engineering workflow smarter and more effective. Over time, this self-reinforcing loop—not just the underlying model—will become a key driver of AI progress.</p><h2>The evolving role of human engineers</h2><p>At AMD, we view AI as a means of increasing productivity, improving quality, and enabling employees to focus on higher-value work. Our goal is to empower our workforce with AI, not to reduce headcount.</p><p>To support this transformation, we are investing heavily in AI education and training across the company. The way we work is evolving rapidly, and we want every AMD employee to be prepared to leverage AI confidently, responsibly, and effectively.</p><p>As AI agents continue to improve, engineers will spend less time manually implementing solutions, focusing more on defining specifications, validating outcomes, and making the strategic decisions that drive innovation.</p>]]></description><pubDate>Mon, 17 Aug 2026 14:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/amd-agent-swarms</guid><category>Agentic-ai</category><category>Software-engineering</category><category>Amd</category><category>Large-language-models</category><dc:creator>Andrej Zdravkovic</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/colorful-3d-blocks-piled-behind-glass-panels-displaying-white-code-snippets.jpg?id=67609515&amp;width=980"></media:content></item><item><title>AI Used to Verify Toughest Mathematics Proof Yet</title><link>https://spectrum.ieee.org/axiom-math-246-theorem-formalization</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/abstract-illustration-of-several-long-arrows-stacked-horizontally-parallel-to-one-another-each-arrow-has-one-plotted-point-in-a.jpg?id=67608532&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Representing a significant milestone in AI-assisted mathematical research, a team at <a href="https://axiommath.ai/" rel="noopener noreferrer" target="_blank">Axiom Math</a> has automatically verified the proof of a theorem relating to prime numbers—colloquially referred to as the “246 theorem”—for the first time using the company’s AI system AxiomProver.</p><p>In formal verification, mathematicians task a computer with checking a machine-readable version of a proof. The process is not a 100 percent guarantee that the proof is correct, <a href="https://gigazine.net/gsc_news/en/20260803-collatz-lean-kernel-bug/#gsc.tab=0" rel="noopener noreferrer" target="_blank">as a recent demonstration showed</a>, exposing how a bug in the method could be exploited to accept a false, AI-generated proof. Still, the computational method is as close to a rubber stamp as you can get. </p><p>This particular verification formalizes an important advance in number theory. Beyond this particular proof, it demonstrates how automated AI verification could be used in the future to ensure the correctness of AI-generated computer code that will soon underlie software across the globe.</p><h2>Useful formalization by design</h2><p>This is not AxiomProver’s first rodeo. Axiom Math has used its autonomous, multi-agent system that turns mathematical statements into machine-checkable proofs to crack several unsolved mathematical problems and verify many more proofs <a href="https://axiommath.ai/selected-publications" rel="noopener noreferrer" target="_blank">this year</a>. But proof formalization of the 246 theorem is by far the most significant, as <a href="https://www.linkedin.com/in/ken-ono-a972191a5/" rel="noopener noreferrer" target="_blank">Ken Ono</a>, Axiom Math’s founding mathematician, explains: “This theorem currently represents the threshold of human knowledge about prime numbers.”</p><p>Earlier this year, Axiom Math competitor <a href="https://www.math.inc/" rel="noopener noreferrer" target="_blank">Math, Inc.</a> used its Gauss agent to verify <a href="https://people.epfl.ch/maryna.viazovska?lang=en" rel="noopener noreferrer" target="_blank">Maryna Viazovska</a>’s 2022 Fields Medal-winning proof of the sphere-packing problem in 8 and 24 dimensions. <a href="https://thefundamentaltheor3m.github.io/" rel="noopener noreferrer" target="_blank">Sidharth Hariharan</a>, a Ph.D. student at Carnegie Mellon University who led human efforts that were critical in the Math, Inc. breakthrough, says that Axiom Math’s AI approach to formalizing the 246 theorem is more comprehensive and useful. Hariharan’s group continues to work toward fully formalizing Viazovska’s proof.</p><p class="ieee-inbody-related">RELATED: <a href="https://spectrum.ieee.org/ai-proof-verification" target="_blank">Watershed Moment for AI-Human Collaboration in Math</a></p><p>Now an intern at Axiom Math, Hariharan has been heavily involved in the company’s formalization of the 246 theorem proof. He says that one of the main differences here is that rather than it being a one-shot approach relating to a single problem, Axiom Math has expressly aimed to make components of the formalization reusable for other formalization tasks and mathematical research. The team has wielded AxiomProver to build a <a href="https://github.com/AxiomMath/PrimeGapsLib" rel="noopener noreferrer" target="_blank">library of results about gaps in primes</a>. The 246 theorem is the flagship result within that library. </p><h2>What is the 246 theorem?</h2><p>The first few primes are close together: 2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, .... And there are several instances where they are separated by a difference of two: 3:5, 5:7, 11:13, 17:19, ....</p><p>These pairs of primes are called twin primes. Twin primes become rarer the further you get from zero, but they do still seem to pop up occasionally. The <a href="https://mathworld.wolfram.com/TwinPrimeConjecture.html" rel="noopener noreferrer" target="_blank">twin prime conjecture</a>, first precisely formulated in the 19th century by French mathematician Alphonse de Polignac, posits that they will keep popping up regardless of how far along the number line you look. In other words, there are infinitely many twin primes.</p><p>Though easy to state, the venerable twin prime conjecture remains unproven. First progress toward solving it only occurred in 2013 when Yitang Zhang, now a professor at Sun Yat-sen University, in Guangzhou, China, <a href="https://annals.math.princeton.edu/2014/179-3/p07" rel="noopener noreferrer" target="_blank">proved that there are infinitely many pairs of primes</a> that are separated by 70 million. A few months later, using a different technique, University of Oxford professor <a href="https://www.sjc.ox.ac.uk/discover/people/professor-james-maynard/" rel="noopener noreferrer" target="_blank">James Maynard</a> dramatically reduced this gap <a href="https://arxiv.org/abs/1311.4600" rel="noopener noreferrer" target="_blank">from 70 million to just 600</a>; a feat which substantially contributed to Maynard being awarded the <a href="https://www.mathunion.org/imu-awards/fields-medal/fields-medals-2022" rel="noopener noreferrer" target="_blank">2022 Fields Medal</a>—widely regarded as the <a href="https://spectrum.ieee.org/tag/nobel-prize" target="_self">Nobel Prize</a> for mathematics. </p><p>As part of a group of mathematicians known as the Polymath8b collaboration, Maynard and fellow Fields Medalist <a href="https://mathstodon.xyz/@tao" target="_blank">Terence Tao</a>, professor at the University of California, Los Angeles, brought the gap down to just 246, the closest mathematicians have gotten to the target gap of two. It is this 246 theorem—which states that there are infinitely many primes that differ by 246—that AxiomProver has verified to be correct. </p><h2>Safe and correct AI-generated code</h2><p>The techniques formalized in this work are important in number theory, the branch of mathematics that underpins all present-day cybersecurity and cryptography. They could therefore prove to be useful in verifying specific ways in which we keep our digital data safe in the future. </p><p>But Axiom Math’s Ono is more excited by the bigger picture. He sees formalizing mathematical proofs as a stepping stone to verifying AI-generated code, which is starting to be used across society in systems that run our infrastructure, manage our finances, and protect our data. This is despite safety concerns surrounding hallucinations, bugs, and other unintended vulnerabilities.</p><p>If properties of code—such as whether an algorithm terminates or if a program’s output is correct for any input—can be translated into precise mathematical statements, technologies derived from AxiomProver would be ideally suited to formally stating and proving them. In this way, mathematically verifying the correctness of AI-generated code would make this code safe to use.</p><p>“The world is about to run on computer code that nobody has read,” Ono concludes. “AI is here and we can no longer look away—proof formalization is a test bed for solving what I think is the most important challenge we will face from AI.”</p><p><em>This article was updated on 18 August to clarify the nature of Hariharan’s work formalizing Viasovska’s proof. </em><br/></p>]]></description><pubDate>Mon, 17 Aug 2026 13:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/axiom-math-246-theorem-formalization</guid><category>Mathematics</category><category>Prime-numbers</category><category>Ai-reasoning</category><category>Ai-generated-software</category><dc:creator>benjamin_skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/abstract-illustration-of-several-long-arrows-stacked-horizontally-parallel-to-one-another-each-arrow-has-one-plotted-point-in-a.jpg?id=67608532&amp;width=980"></media:content></item><item><title>The CPU Comeback Is Upon Us</title><link>https://spectrum.ieee.org/ai-cpu-comeback</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-single-computer-chip-balancing-on-the-tip-of-a-pyramid.jpg?id=67615384&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>Earlier this year, leaders at <a href="https://aws.amazon.com/free/?trk=dc9b9d60-cc82-4cd5-8a61-0b33d6a79fab&sc_channel=ps&ef_id=CjwKCAjw1vXTBhB-EiwAEKr_k_57Xoz6QKQSrRqDF4mGRYOudA_A99MTHZL6alVjPDxdliBQfEMPihoC0RUQAvD_BwE&gads_camp=23532472510&gads_ag=199502799824&gads_ad=795877020713&gads_kw=amazon%20web%20services&gads_matchtype=e&gads_network=g&gads_device=c&gads_geo=9198314&gad_campaignid=23532472510&gbraid=0AAAAADjHtp-SgwVQvE7H9V8lK49jVtMDw&gclid=CjwKCAjw1vXTBhB-EiwAEKr_k_57Xoz6QKQSrRqDF4mGRYOudA_A99MTHZL6alVjPDxdliBQfEMPihoC0RUQAvD_BwE" rel="noopener noreferrer" target="_blank">Amazon Web Services</a> delivered a new mandate to their engineers: They need to conserve CPU cycles at all costs. AWS has <a href="https://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunch" rel="noopener noreferrer" target="_blank">reportedly</a> experienced an explosion in wait times for CPU server capacity as AI workloads strain the company’s cloud infrastructure. </p><p>The issue seemingly took AWS off guard, and for good reason. The AI boom led to a surge in demand for <a href="https://spectrum.ieee.org/nvidia-gpu" target="_self">GPUs</a> and, later, <a href="https://spectrum.ieee.org/high-bandwidth-memory-shortage" target="_self">memory</a>. CPUs were mostly left out of the story, as their relative lack of parallelization made them a poor fit for AI model inference, the process of running and serving large language models (LLM) to users. </p><p>But the rise of agentic AI systems, which allow AI models to operate autonomously and call on sub-agents, is changing the narrative.</p><p><a href="https://moorinsightsstrategy.com/team/matt-kimball/" rel="noopener noreferrer" target="_blank">Matt Kimball</a>, vice president and principal data center analyst at <a href="https://moorinsightsstrategy.com/" rel="noopener noreferrer" target="_blank">Moor Insights & Strategy</a> in Austin, Texas, says 2026 has brought a spike in CPU demand, much of it due to <a href="https://spectrum.ieee.org/ai-agents" target="_self">agentic AI</a>. “It’s one thing to have this agentic workload, and let’s say it spawns 100 agents. If I’m going to roll this out across my enterprise, those 100 become tens of thousands, hundreds of thousands, or millions of agents,” Kimball says. “You have agents spawning sub-agents, making API [application programming interface] calls and talking to more agents through [Anthropic’s] model context protocol.”</p><h2>AI agents need to use computers, and computers need CPUs</h2><p>Kimball’s comments refer in part to “tool use,” which is shorthand for an LLM’s ability to access the internet, open files on a desktop, and generally use a variety of software to accomplish its task. </p><p>LLMs trained for tool use learn how to call on other software. While the LLM’s inference is still primarily executed on a GPU or similar AI accelerator, the tool calls that the LLM makes are typically pushed to the CPU.</p><p>“Many components of an agentic AI task are inherently CPU based jobs,” explains <a href="https://www.linkedin.com/in/souvik-kundu-64922b50/" rel="noopener noreferrer" target="_blank">Souvik Kundu</a>, senior staff research scientist at <a href="https://www.intel.com/content/www/us/en/homepage.html" rel="noopener noreferrer" target="_blank">Intel</a>. “The CPU does the job of parsing output, figuring out which tool to invoke, making the API call or running the code, collecting the result, and feeding it back.” <a href="https://www.linkedin.com/in/mrangarajan/" rel="noopener noreferrer" target="_blank">Madhu Rangarajan</a>, vice president of compute and enterprise AI products at <a href="https://www.amd.com/en.html" rel="noopener noreferrer" target="_blank">AMD</a>, makes a similar claim, saying, “In our testing, seven of the eight stages in realistic agentic AI pipelines run entirely on the CPU.”</p><p>An LLM tasked with programming software, for example, will likely make tool calls to write code to files, move or replace files, download required packages, and build the software once the LLM believes it’s complete. </p><p>Kundu co-authored a <a href="https://arxiv.org/pdf/2511.00739" rel="noopener noreferrer" target="_blank">paper</a> on agentic AI optimization alongside researchers from Georgia Tech in Atlanta. They found the CPU is often idle while LLM inference is executed on a GPU and that, conversely, the GPU is often idle when tool calls are executed on the CPU. To optimize this, Kundu and his colleagues propose scheduling optimizations that can cut end-to-end latency (the time between the start and finish of the agentic workload) by up to 1.8-times under sustained load. </p><p>It’s a start, but the gains chase a moving target. Agentic systems generate work at machine speed and multiply it as they go. OpenAI’s inadvertent <a href="https://spectrum.ieee.org/hugging-face-openai-cyberattack?itm_source=homepage&itm_medium=hero&itm_campaign=hero-2026-08-10&itm_content=hero6" target="_self">hack</a> of Hugging Face saw its model fire off as many as 300 actions an hour, and a single agent can spawn sub-agents that make tool calls of their own. </p><p>And there’s one more important complication that may increase the workload on a CPU as models become more complex: safety guardrails.</p><p>Safety and policy checks on an agent’s actions are often specific rules that inspect syntax and log files, Kundu says. Guardrails may also use small models (under a billion parameters) to analyze task complexity or intent. Though they could be executed on a GPU, they often aren’t, because their small size and the need to minimize latency keeps the work on the CPU.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A bar graph illustrating how the latency for Llama-8B\u2019s LLM becomes 4.45 times less when the number of CPU cores is increased from five to thirty-two." class="rm-shortcode" data-rm-shortcode-id="b43d668525bb89a3e691433307b18677" data-rm-shortcode-name="rebelmouse-image" id="7acb8" loading="lazy" src="https://spectrum.ieee.org/media-library/a-bar-graph-illustrating-how-the-latency-for-llama-8b-u2019s-llm-becomes-4-45-times-less-when-the-number-of-cpu-cores-is-increas.jpg?id=67615386&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Increasing the number of CPUs available significantly decreases the latency for Llama-8B responses over longer sequence lengths.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Source: <a href="https://arxiv.org/pdf/2603.22774" target="_blank">Euijun Chung, Yuxiao Jia, et al.</a></small></p><h2>Tokenization adds to bottlenecks</h2><p><a href="https://ejchung0406.github.io/" target="_blank">Euijun Chung</a>, a PhD student at Georgia Tech, recently co-authored another <a href="https://arxiv.org/abs/2603.22774" target="_blank">paper</a>, with findings that complement Kundu’s work. Chung and his co-authors found that when a server has too few CPU cores, it falls behind on dispatching work to the GPUs. That causes the GPUs to stall as they wait for instructions.</p><p>In addition to that, the paper touches on another key element of LLM workloads: <a href="https://seantrott.substack.com/p/tokenization-in-large-language-models" target="_blank">tokenization</a>.</p><p>Tokenization is a key first step in LLM inference. It converts text into integer token IDs that can be processed by the model. Unlike the matrix math required for most LLM inference, tokenization is branchy, data-dependent sequential string manipulation. Though it can be parallelized by chunking text, it’s not massively parallel in the same way as the bulk of LLM inference.</p><p>Tokenization of small prompts is a relatively trivial task and won’t tax even an entry-level CPU. However, an agentic model that makes tool calls must parse and tokenize the results of the call. </p><p>“If you have an ongoing sequence of, say, 100,000 tokens, and you have a tool result of 1,000 tokens, the tokenizer will have to tokenize the whole sequence again. And you have to do tokenization at every agentic tool call,” Chung says. This both increases the frequency of tokenization and increases the number of tokens involved. It’s conceivable that future tokenizers will find ways to mitigate this, Chung says, but it remains a problem for modern LLM inference.</p><p>The paper finds that time-to-first-token latency (the time required for the model to produce the first word of its reply) can increase dramatically as the sequence length grows. CPUs with more cores can reduce the problem. In test runs at longer sequence lengths, increasing CPU core counts can reduce time-to-first-token latency by roughly 1.5 to 7 times.</p><p>Chung and his colleagues were only able to test smaller models, such as Alibaba’s Qwen 3-30B and Meta’s Llama 3.1-70B, due to limitations of the hardware available for testing. He speculates that larger models will experience less dramatic bottlenecks due to their higher overall GPU demand, but also expects agentic AI will push token lengths far beyond the longest he and his co-authors tested.</p><p>“If you think about something like Anthropic’s Claude, you can easily hit 500,000, even a million tokens,” Chung says. “In the world of agentic AI, the average sequence length will grow and grow, so I’m expecting this problem to get worse in future workloads.” </p><h2>Is a CPU crunch just getting started?</h2><p>Amazon’s crackdown on use of CPU resources is one of several indicators that Kundu and Chung have identified issues with real-world relevance. </p><p>Intel has <a href="https://finance.yahoo.com/news/intel-turnaround-no-one-saw-141000146.html?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAAJqbhttNhPiUgZBGw-QWbgpaSpMnF0qwXSRBrSMQvCVWxhrkdpDYWyavDbXWb-9coPNWysaUbJK_l5uY_bZH6ACvfNpR8JrxKbKXsw8dVSlzZLYe9ydhDqS0SAyZbc4KSK0bqLw8WJSIYt-yoFGlGWDRP1UIaWov2q1bpzbcLPE-" rel="noopener noreferrer" target="_blank">sold out</a> of server CPUs through at least the end of the year. AMD has <a href="https://wccftech.com/amd-doubles-server-cpu-forecast-to-120-billion-as-agentic-ai-rewrites-demand-ceo-says-epyc-verano-built-purely-for-ai/" rel="noopener noreferrer" target="_blank">doubled</a> its server CPU forecast. <a href="https://www.arm.com/products/cloud-datacenter/arm-agi-cpu" rel="noopener noreferrer" target="_blank">Arm</a> and <a href="https://www.cnbc.com/2026/06/24/qualcomm-data-center-cpu-meta.html" rel="noopener noreferrer" target="_blank">Qualcomm</a> have both announced new CPUs designed to accelerate agentic AI. Even <a href="https://www.nvidia.com/en-us/" rel="noopener noreferrer" target="_blank">Nvidia</a> has prioritized <a href="https://nvidianews.nvidia.com/news/nvidia-unveils-vera-the-cpu-for-agents" rel="noopener noreferrer" target="_blank">Vera</a>, its Arm-based CPU for agentic AI, which is part of Nvidia’s <a href="https://spectrum.ieee.org/nvidia-rubin-networking" target="_self">Vera Rubin</a> platform.</p><p>Kimball says these developments make it clear that the AI industry is placing more emphasis on CPU performance. He sees the surge in demand as an “absolute tell” that CPUs are now considered a key part of an agentic AI system.</p><p>Unfortunately, this may translate to broader CPU shortages and increased prices, much as has already occurred with GPUs and memory. </p><p>“You’re already seeing a CPU crunch to some degree. When you look at the constraints in the market, it even trickles down into the consumer space,” Kimball says. He adds that Intel has <a href="https://www.techpowerup.com/345535/intel-reallocates-pc-production-capacity-to-server-cpus-amid-tight-wafer-supply" rel="noopener noreferrer" target="_blank">cut production</a> of client CPUs in favor of server CPUs, even as Intel’s new 18A production process has grown the company’s sales in the client segment. Kimball sees that as a sign that CPU makers will follow the money. </p>]]></description><pubDate>Sun, 16 Aug 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-cpu-comeback</guid><category>Cpu</category><category>Gpu</category><category>Agentic-ai</category><category>Llms</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-single-computer-chip-balancing-on-the-tip-of-a-pyramid.jpg?id=67615384&amp;width=980"></media:content></item><item><title>Inside the Data Bottleneck Slowing Visual and Physical AI</title><link>https://content.knowledgehub.wiley.com/the-2026-state-of-visual-and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/voxel51-logo-with-geometric-cube-icon-and-stylized-text.png?id=67607900&width=980"/><br/><br/><p>A survey of over 700 professionals examines how visual and physical AI teams build systems, why models fail, and where data work drives production.</p><p><span><a href="https://content.knowledgehub.wiley.com/the-2026-state-of-visual-%20and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/" target="_blank">Download this free whitepaper now!</a></span></p>]]></description><pubDate>Wed, 12 Aug 2026 14:18:05 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/the-2026-state-of-visual-and-physical-ai-a-survey-of-700-practitioners-on-data-models-and-production/</guid><category>Type-whitepaper</category><category>Artificial-intelligence</category><category>Computer-models</category><category>Data-bottleneck</category><dc:creator>Voxel51</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67607900/origin.png"></media:content></item><item><title>Pakistani Judges Give Their Verdict on JudgeGPT</title><link>https://spectrum.ieee.org/judgegpt-experiment</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-a-translucent-gavel-against-a-background-of-legal-books.jpg?id=67600001&width=1245&height=700&coordinates=0%2C156%2C0%2C157"/><br/><br/><p>Judges around the world have <a href="https://www.bbc.com/news/articles/c178zzw780xo" rel="noopener noreferrer" target="_blank">made headlines</a> for <a href="https://www.judiciary.senate.gov/press/rep/releases/grassley-scrutinizes-federal-judges-apparent-ai-use-in-drafting-error-ridden-rulings" rel="noopener noreferrer" target="_blank">illicitly using generative AI</a> in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology—together with appropriate training–boosted the number of cases resolved by 6.3 percent with no obvious drop in the quality of judgments.</p><p>With a backlog of 2.26 million cases and fewer than two judges per 100,000 people—compared to 22 in the EU and eight in Brazil—Pakistan’s judiciary was in sore need of help. So, in consultation with the judiciary, <a href="https://sites.google.com/view/sultan-mehmood/home" rel="noopener noreferrer" target="_blank"> economist Sultan Mehmood</a>, of the New Economic School in Moscow, and collaborators tested whether AI could ease the burden.</p><p>They built a custom tool combining OpenAI’s GPT-4 large language model (LLM) with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes, to help judges with legal research and drafting judgments. They began offering the tool in 2024 to 1,559 trial judges—roughly half the country’s justices. </p><p>“We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,” Mehmood says. </p><h2>First of its kind</h2><p>“It’s pretty amazing that he’s able to pull this off,” says David Autor, an economics professor at MIT. “It’s not easy to do large-scale field experiments in civil service, but especially where the stakes are so high.” The 6.3 percent productivity boost is not overwhelming, he says, but it’s credible and likely to improve as the tool is more widely used.</p><p>AI tools for judges are already being rolled out in <a href="https://a2jlab.org/brazils-ai-driven-courts-innovation-without-evaluation/" rel="noopener noreferrer" target="_blank">Brazil</a> and <a href="https://iacajournal.org/articles/10.36745/ijca.647" rel="noopener noreferrer" target="_blank">India</a>, and prominent U.S. law professor Eric Posner has compared LLM judgments to human judgments in a single <a href="https://journals.sagepub.com/doi/10.1177/2755323X261433614" rel="noopener noreferrer" target="_blank">case study</a>. But until now, there has been no major independent assessment of ongoing judicial use of AI. The <a href="https://elliottash.com/papers/Mehmood-Goessmann-Ash-Courts-of-Tomorrow-Evidence-Nationwide-Rollout-Generative-AI.pdf" rel="noopener noreferrer" target="_blank">new study</a> focused on Pakistan’s trial courts; Mehmood says judges there were enthusiastic from the start.</p><p>“They were more techno-optimist than we were,” he says. “The delays are so huge, this is something which they thought was worth trying anyway to reduce people’s suffering.”</p><p>Some judges were also already using AI chatbots, Mehmood says, but commercial offerings performed poorly on Pakistani legal queries, frequently hallucinating case law. So the team built a tool tailored to the Pakistani context, called JudgeGPT.</p><p>They used retrieval-augmented generation (RAG), which allowed the model to query a database of 128,292 Pakistani judicial opinions and 943 statutes. Responses included footnotes linking to cases and laws.</p><p>“It turns out that actually the way to fix [hallucinations] isn’t just more intelligent models,” says study coauthor <a href="https://elliottash.com/" rel="noopener noreferrer" target="_blank">Elliott Ash</a>, an associate professor of law, economics, and data science at ETH Zurich in Switzerland. “It’s to attach the models to a tool that can do a search and verify the sources.” However, the researchers do not report hallucination rates.</p><p>The team also put 1,197 judges through six 90-minute Zoom training sessions, developed in collaboration with Pakistan’s <a href="https://www.fja.gov.pk">Federal Judicial Academy</a>, covering how LLMs work, their limitations, the risk of bias and hallucinations, and the importance of verifying outputs. Another 180 judges only underwent general training on technology in legal research, while a final group got no training.</p><p>By the time 487 judges had been through the  program, the median district saw a jump of 6.3 percent resolved cases, and the more trained judges in a district, the bigger the effect. Appeal rates also fell slightly, suggesting faster resolution wasn’t leading to sloppier decisions.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Screenshot of an AI agent responding to a user\u2019s prompt for a list of robbery-related cases from the 2000s." class="rm-shortcode" data-rm-shortcode-id="eaae2b1d0ba702a7ae4322aadf97f4d0" data-rm-shortcode-name="rebelmouse-image" id="cecbd" loading="lazy" src="https://spectrum.ieee.org/media-library/screenshot-of-an-ai-agent-responding-to-a-user-u2019s-prompt-for-a-list-of-robbery-related-cases-from-the-2000s.jpg?id=67600006&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">JudgeGPT can be used to surface relevant case law with a simple text query, and results provide links to the full text of the related judgments.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Sultan Mehmood, Christoph Goessmann, and Elliott Ash</small></p><h2>Addressing limitations</h2><p>The team also assessed the quality of judgments. Having legal experts evaluate large numbers of judgments was infeasible, Mehmood says, so the team asked OpenAI’s GPT-5-mini to choose between pairs of judgments from the same judge before and after training. The LLM chose post-training judgments 59 percent of the time. Two experienced Pakistani lawyers also evaluated the model’s analysis of 90 judgment pairs. They agreed with GPT-5-mini 70.6 percent of the time, compared to 73 percent agreement with each other.</p><p><a href="https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu" target="_self">Training</a> turned out to be vital. On average, JudgeGPT-trained judges logged in 56 times and sent 212 prompts over the study period, compared to 10 logins and 25 prompts after generic training. Those who had no training tended to use the tool for around a month and then drop off entirely, Mehmood says. “Just giving people the technology does not necessarily make them use it persistently,” he says.</p><p>A 6.3 percent increase sounds modest, but the researchers calculated that a trained judge was resolving 38.5 more cases a month than the baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash also notes that these figures come from a nine-month period at the start of the trial, and that they’ve since updated both the underlying AI model and the database.</p><p>For users, the tool has been a lifeline. One participating trial judge, who spoke on condition of anonymity, says the number of cases assigned to them hasn’t dropped below 1,000 in more than a decade. The tool saves significant time, in particular searching for case law and summarizing lengthy documents. “For research, it’s just one prompt away, whereas before I had to search for the precedents and laws for hours,” the judge says. “If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarize it for me and give me the crux, and it does that work in seconds.”</p><p>But efficiency isn’t the only thing you want out of a justice system, says <a href="https://scholars.latrobe.edu.au/jzeleznikow" target="_blank">John Zeleznikow</a>, professor of law and technology at La Trobe University, in Australia. “What they’ve tried to do is be effective, [to] deal with more cases more quickly, and they’re able to do that,” he says. “What’s not that clear is whether what you call the quality of justice is better.”</p><p>Zeleznikow says AI can be useful, but only if judges are <a href="https://spectrum.ieee.org/ai-reasoning-failures" target="_self">diligent about evaluating and verifying the output</a>. However, the working paper’s authors found that roughly a fifth of participants’ prompts given to JudgeGPT involved what they call “substantial AI delegation”—asking the tool what the best decision is, to produce legal reasoning or write opinions with little input from the judge. On the bright side, training lowered the proportion of inappropriate delegation.</p><p>But given that judges are already using AI, Ash says better tools and training are crucial. “There are risks for using these AIs, for sure, even with all these safeguards. But at some point you have to just put the judges in as strong a position as you can,” he says. “Have technological safeguards, but then try to encourage the judges not to rely on it too much.”<br/><br/><em>This story was updated on 13 August 2026 to clarify that the training course was developed in coordination with Pakistan’s Federal Judicial Academy.</em></p>]]></description><pubDate>Wed, 12 Aug 2026 11:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/judgegpt-experiment</guid><category>Justice</category><category>Pakistan</category><category>Llms</category><category>Legal-ai</category><dc:creator>Edd Gent</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-a-translucent-gavel-against-a-background-of-legal-books.jpg?id=67600001&amp;width=980"></media:content></item><item><title>AI Safety Regulations in the U.S. Could Give Hackers an Edge</title><link>https://spectrum.ieee.org/hugging-face-openai-cyberattack</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-hugging-faces-smiley-face-logo-holding-up-scales-of-justice-against-a-background-of-binary-code.jpg?id=67583759&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor. The speed and coordination of the attack on the company that hosts and supports popular AI developer resources led Hugging Face’s security team to conclude it was the work of <a data-linked-post="2669884140" href="https://spectrum.ieee.org/ai-agents" target="_blank">an AI agent</a>. </p><p>Realizing this, the team tried to use <a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer" target="_blank">“frontier models behind commercial APIs”</a>—presumably from Anthropic and OpenAI, although only Anthropic was named in the second of the company’s two posts about the security incident—to analyze the onslaught. These models refused to help due to safety guardrails the AI labs have implemented to make their models harder to use for cyberattacks. Hugging Face instead turned to GLM 5.2, a model from Beijing-based AI lab Z.ai, to aid its analysis. </p><p>On 21 July, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">OpenAI announced</a> the attacker was an OpenAI model undergoing testing in a sandboxed environment. It escaped its internal sandbox, established a foothold in a third-party server, and then assailed Hugging Face. In other words, frontier models—those that score highest in AI performance benchmarks—had refused to assist Hugging Face’s security team in analyzing the attack, yet a prospective frontier model in testing had executed it in the first place.</p><p>“I would argue that asymmetry is the paramount problem of our time,” says <a href="https://www.linkedin.com/in/alexlevinson/" rel="noopener noreferrer" target="_blank">Alex Levinson</a>, executive director of the <a href="https://www.ccdc.io/" rel="noopener noreferrer" target="_blank">National Collegiate Cyber Defense Competition</a> and coauthor of a paper on <a href="https://arxiv.org/abs/2603.01246" rel="noopener noreferrer" target="_blank">defensive refusal bias</a>. “We want the world to exist in a state of security, but we’re not going to get there by guardrailing away model capability.”</p><h2>Massive AI Cyberattack on Hugging Face</h2><p>The scale of the OpenAI model’s attack on Hugging Face was massive. Across five days, it <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline" rel="noopener noreferrer" target="_blank">executed over 17,500 individual actions</a>, such as privilege escalation and code execution. At its peak, the model performed more than 300 actions per hour. While the attack resulted in little damage to Hugging Face’s infrastructure, the model was able to steal credentials, gain admin access, and extract some data.</p><p>All of this was in pursuit of a simple goal: The model wanted to cheat on a test. </p><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">According to OpenAI’s press release</a>, the model was tasked with solving a cybersecurity benchmark called <a href="https://github.com/sunblaze-ucb/exploitgym" rel="noopener noreferrer" target="_blank">ExploitGym</a>. The model inferred that Hugging Face might have data on the benchmark and broke into the company’s infrastructure to find it. The model was ultimately successful in extracting five dataset files, though it’s not clear if the data helped it achieve its goal. OpenAI and Hugging Face did not respond to requests for comment. <br/><br/>Cybersecurity consultant <a href="https://www.linkedin.com/in/chuck-h-securityexecutive/" rel="noopener noreferrer" target="_blank">Chuck Herrin</a> observes that though the model’s actions were alarming, they shouldn’t be considered unexpected, as the model was ultimately pursuing the goal it was given. <strong>“</strong>This autonomous agent was designed to go and figure things out, and it went and figured things out. It’s not surprising in any way.”</p><p>And errant AI agents may be more common than we thought. OpenAI’s disclosure motivated researchers at Anthropic to review their own cybersecurity evaluations. On 30 July, Anthropic disclosed three instances where a model executed an attack as part of an evaluation. In one case, <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer" target="_blank">Claude uploaded malware to PyPI</a>, the official Python software repository.</p><h2>AI Guardrails and Cybersecurity Asymmetry</h2><p>The campaign OpenAI’s model conducted against Hugging Face highlights how AI policy has the potential to create an asymmetry between attackers and defenders. </p><p>When Levinson was head of security at Scale AI, an AI development and evaluation company, he and his colleagues began to notice this as AI found use in cybersecurity competitions. (Levinson left Scale AI in February 2026.)</p><p>“I would say that since 2023, we have felt there was guardrailing in place that was stifling a lot of the time. Not all of the time, but it was getting in the way,” says Levinson. The Scale AI team quantified the problem <a href="https://iclr.cc/virtual/2026/10016243" rel="noopener noreferrer" target="_blank">in a paper published at ICLR 2026</a>, which found that, depending on the task, nearly 44 percent of defensive requests were refused. <span>The results</span><span>, which use data from a cybersecurity competition held in April 2025, predate U.S. policy actions that have further hardened safety guardrails.</span></p><p>In June, the U.S. Department of Commerce, <a href="https://www.pbs.org/newshour/show/anthropic-disables-new-ai-model-after-white-house-security-directive" target="_blank">citing a jailbreak that threatened to unlock unrestricted cyber capabilities</a>, invoked export-control authority in a way that caused Anthropic to suspend all access to its most capable models, Fable 5 and Mythos 5. Access was <a href="https://www.wsj.com/tech/ai/anthropic-nears-deal-with-trump-administration-to-restore-access-to-fable-ai-model-6f4177f3" target="_blank">partially restored weeks later</a> after negotiations with the Trump administration included more rigorous safety guardrails. The system card for OpenAI’s GPT-5.6, which summarizes its capabilities, <a href="https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf" target="_blank">states it also has more robust guardrails than prior releases</a>.</p><p class="pull-quote">“We want the world to exist in a state of security, but we’re not going to get there by guardrailing away model capability.” <strong>—Alex Levinson, National Collegiate Cyber Defense Competition</strong></p><p>These new guardrails have seemingly made models even more unlikely to fulfill defensive requests. <a href="https://www.linkedin.com/in/christopher-covino-920956163/" target="_blank">Christopher Covino</a>, senior researcher at the Institute for AI Policy and Strategy think tank, says Anthropic’s safeguards are extremely stringent. “There are even academic papers that Fable will not read for me, or not let me talk about,” he says, though he adds that OpenAI’s safeguards are more accommodating.</p><p>Levinson has also noticed ever-tighter restrictions in more recent cybersecurity competitions, though he and his coauthors haven’t had the opportunity to repeat the 2025 test.<br/><br/>In theory, more rigorous restrictions might seem to average out. While they may hamper cybersecurity defense and research, they can also hamper attackers. </p><p>But that assumes everyone has access to models with the same safety guardrails and that nobody tries to circumvent them. This is the asymmetry Levinson was alluding to: Attackers tend not to respect the same rules as defenders. </p><p>The attack on Hugging Face from OpenAI’s model also shows that the models can, in rare circumstances, take steps that circumvent their own safeguards.</p><h2>Chinese AI Models in U.S. Cyber Defense</h2><p>The policy implications are further complicated by the fact that Hugging Face’s security team didn’t use a leading U.S. model to analyze the attack, but instead used GLM 5.2, a recent release from Chinese AI lab Z.ai. </p><p>Hugging Face’s security team didn’t access GLM 5.2 through Z.Ai. GLM 5.2 is an open-weights model, which means the model is available for anyone to download and use. Hugging Face hosted the model on its own infrastructure. </p><p>The reliance on GLM 5.2 is complicated by recent saber-rattling about ways the U.S. could restrict Chinese models. Recent open-weights models from labs based in China, including GLM 5.2 and Moonshot AI’s Kimi K3, <a href="https://spectrum.ieee.org/ai-coding-assistant-china-anthropic" target="_self">have scored close to leading U.S. models in benchmarks</a>. On 20 July, <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi" target="_blank">Axios reported that the Trump administration is considering a ban</a> on Chinese models. </p><p class="pull-quote">“This autonomous agent was designed to go and figure things out, and it went and figured things out. It’s not surprising in any way.” <strong>—Chuck Herrin, Herrin Advisory</strong><br/><strong></strong></p><p>These restrictions have yet to materialize but, if they did, they could cut off U.S. companies like Hugging Face from the best models willing to come to their defense.</p><p>The incident demonstrates how AI policy can become a double-edged sword. Model guardrails are intended to prevent the use of AI models in cyberattacks. A ban on Chinese models, if it were announced, would likely be justified in part by security concerns. Yet these moves can harm defenders as much as attackers. </p><p>“There’s this tension here,” says Covino. “Increased safeguards limit risk, but you also limit legitimate defensive use.” Attackers will find ways around the restrictions regardless, he notes. “So it’s a question of, do we want to inhibit the defenders?”</p><p>That’s not to say U.S. policymakers should let AI models run wild.</p><p>Covino would like to see a national dashboard tracking the frequency and success of AI cybersecurity attacks, and he sees utility in trusted access programs that give vetted, traceable defenders access to models with reduced safeguards. He also says U.S. agencies should more seriously consider the specifics of how AI can be used for cyber defense and mentions <a href="https://www.energy.gov/ceser/artificial-intelligence-operationally-resilient-technologies-and-systems" rel="noopener noreferrer" target="_blank">AI-FORTS</a>, a program managed by the U.S. Department of Energy’s Office of Cybersecurity, Energy Security, and Emergency Response, as a leading example.</p><p>“Let the leash loose a little,” Covino says. “Anthropic would know if someone is terribly abusing it, and if there is an attack, it can be traced back.” </p><p>Herrin has similar feelings on accountability. He believes the AI industry should more seriously consider standards such as the Artificial Intelligence Management System specified in the <a href="https://www.iso.org/standard/42001" rel="noopener noreferrer" target="_blank">ISO/IEC 42001 </a>standard, which requires organizations to document an AI system’s likely impacts before deployment and to name the humans answerable for them.</p><p>Herrin also noted that the lack of repercussions from OpenAI’s cyber incident was unusual, as a person who took similar actions would likely draw the attention of law enforcement. “If this was a job candidate being tested in a technical interview, and they committed violations of law in order to pass tests, we’d be having a very different conversation.”</p>]]></description><pubDate>Thu, 06 Aug 2026 19:25:39 +0000</pubDate><guid>https://spectrum.ieee.org/hugging-face-openai-cyberattack</guid><category>Agentic-ai</category><category>Openai</category><category>Huggingface</category><category>Ai-safety</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-hugging-faces-smiley-face-logo-holding-up-scales-of-justice-against-a-background-of-binary-code.jpg?id=67583759&amp;width=980"></media:content></item><item><title>IEEE Course Teaches How to Use AI to Modernize Power Grids</title><link>https://spectrum.ieee.org/ieee-course-ai-power-grids</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-man-holding-a-ginormous-lightbulb-with-ai-written-on-it.jpg?id=67568096&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p>Today’s U.S. electrical grid, among the largest, most complex systems ever built, is operating at its limit. The combination of rapid industrial growth, more frequent extreme weather, and a record surge in electricity use has <a href="https://www.energy.gov/policy/electricity-demand-growth-resource-hub" rel="noopener noreferrer" target="_blank">pushed the grid to its breaking point</a>, according to the <a href="https://www.energy.gov/" rel="noopener noreferrer" target="_blank">U.S. Department of Energy</a>.</p><p>Built decades ago for a more predictable world in which power came mostly from centralized coal or gas plants and electricity use grew at a steady pace, the grid faces <a href="https://spectrum.ieee.org/data-centers-grid-instability" target="_self">unanticipated strain</a> due in part to growing demand from data centers. The jobs of professionals managing the infrastructure have evolved from traditional engineering tasks to complex, fast-moving challenges.</p><p>Industry reports show that millions of modern digital sensors, smart meters, and grid monitors are generating nonstop waves of information. The sheer volume of data requires instant, automated computer analysis because human operators cannot process it fast enough.</p><p>Pressure on utilities stems from two sources: a spike in electricity demand and a shift in how power is generated.</p><p>An example of the operational strain can be seen at the regional level. With the recent deployment of artificial intelligence tools and high-performance computing, data centers require <a href="https://spectrum.ieee.org/dcflex-data-center-flexibility" target="_self">immense amounts of energy</a> to operate. The largest power transmission utility in Texas recently reported a <a href="https://www.cnbc.com/2025/12/12/ai-data-center-flood-texas-on-massive-scale.html" rel="noopener noreferrer" target="_blank">staggering 220 gigawatts</a> of new connection requests, driven largely by a surge in AI and cloud-computing facilities, according to a <a href="https://www.cnbc.com/" rel="noopener noreferrer" target="_blank">CNBC </a>report.</p><p>Alongside the rise in regional demand, global energy networks are absorbing an unpredictable variety of weather-dependent renewable energy such as wind and solar. The switch creates a volatile operating environment wherein supply and demand are balanced, second by second, to prevent blackouts.</p><p>The challenges are compounded by the vulnerability of the grid’s physical and digital framework.</p><p>More-frequent severe weather events cause costly disruptions, such as the devastating winter freeze that crippled the Texas grid and record-breaking heat waves that have overloaded transformers.</p><p>Simultaneously, the energy networks’ digital architecture faces threats. As utilities replace outdated analog equipment with smart meters and control systems, they are increasingly vulnerable to <a href="https://spectrum.ieee.org/power-grid-attack-security-gridex" target="_self">cyberattacks</a>.</p><p>To overcome physical and digital vulnerabilities, grid reliability organizations, such as those conducting North American security simulations like <a href="https://spectrum.ieee.org/power-grid-attack-security-gridex" target="_self">GridEx</a>, emphasize that the grid must become smarter, more agile, and completely automated. Energy researchers are noting that the key to this change lies in integrating AI across every layer of utilities’ operations.</p><h2>The AI imperative</h2><p>According to energy industry experts, using AI to manage power systems is no longer a futuristic research project; it has become a baseline operational necessity. Grid analysts emphasize that traditional grid-planning methods are too slow to handle <a href="https://spectrum.ieee.org/ai-designed-thermoelectric-generator" target="_self">rapid energy dynamics</a> or to balance volatile renewable energy in real time within decentralized power systems such as microgrids.</p><p>AI can fill the gap by processing vast amounts of data instantly. Machine learning algorithms can quickly analyze information from thousands of sensors, historical usage patterns, and weather forecasts to predict issues before they happen.</p><p>An industrial digitization study conducted by <a href="https://www.mckinsey.com/~/media/McKinsey/Business%20Functions/McKinsey%20Digital/Our%20Insights/Digital%20in%20industry%20From%20buzzword%20to%20value%20creation/Digital-in-industry-From-buzzword-to-value-creation.pdf" rel="noopener noreferrer" target="_blank">McKinsey & Co.</a> indicated that integrating advanced data and automation across infrastructure networks could reduce system design errors, decrease equipment downtime by up to 50 percent through predictive maintenance, and extend the lifespan of power machinery by up to 40 percent.</p><p>From forecasting energy spikes to automatically fixing localized voltage drops, AI acts as the digital backbone of a self-healing grid, experts say. Deploying the complex systems requires a new workforce: power engineers who understand data science, as well as data scientists who understand electricity.</p><h2>Upgrading the Workforce</h2><p>To bridge the gap between groundbreaking AI research and practical field deployment, <a href="https://ea.ieee.org" rel="noopener noreferrer" target="_blank">IEEE Educational Activities</a>, in partnership with the <a href="https://ieee-pes.org/" rel="noopener noreferrer" target="_blank">IEEE Power & Energy Society</a>, has launched the online <a href="https://iln.ieee.org/public/contentdetails.aspx?id=48A92EF8188E4D2E8331E1381CAF98E7&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank">Artificial Intelligence for Power and Energy Systems</a> course program.</p><p>The program explores core challenges threatening modern utilities. Rather than treating AI as an unverified black box that operates without human supervision, the curriculum focuses on safety, asset preservation, and strict reliability standards.</p><p>The curriculum is designed to educate power system engineers, utility managers, and data scientists tasked with modernizing the grid. The program was developed by <a href="https://www.linkedin.com/in/fangxing-fran-li-07193b15/" rel="noopener noreferrer" target="_blank">Fangxing “Fran” Li</a>, professor of electrical engineering and computer science at the <a href="https://www.utk.edu/" rel="noopener noreferrer" target="_blank">University of Tennessee</a> in Knoxville and chair of the <a href="https://cmte.ieee.org/pes-mlps/" rel="noopener noreferrer" target="_blank">IEEE Working Group on Machine Learning for Power Systems</a>. </p><h2>Five learning modules</h2><p>The program breaks down the technical transition into five modules that bridge high-level theory with real-world solutions: </p><p><a href="https://iln.ieee.org/public/contentdetails.aspx?id=ED553FD6AE2E475B9F1847E9A83B8460&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>AI fundamentals.</strong></a><strong> </strong>This module teaches engineers how basic machine learning models apply to power grids. It discusses how specialized neural networks solve complex power-flow calculations and how AI models can safely transition from computer simulations to physical, high-voltage equipment. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=21D915950CBE4139A40C62BBE7A59556&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Accelerating grid control.</strong></a> Learners are taught to leverage deep reinforcement learning, an AI approach that uses trial and error, to accelerate automated grid adjustments during emergency power events. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=550105AB2F69474BA043C35CEC71E0A3&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Forecasting and data analytics.</strong></a><strong> </strong>Using predictive modeling, engineers learn how to predict sudden demand surges, variable wind and solar outputs, and fluctuating wholesale electricity market prices to keep power affordable and available. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=0CFD1158E2D14CD1B97AD88BCF6FD38A&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Physics-informed and safe AI.</strong></a><strong> </strong>To address trust—a barrier to utility AI adoption—this course covers AI models hard-coded to obey the laws of physics. The approach is designed to ensure that automated algorithms never make erratic choices that damage grid equipment. </p><p><a href="https://iln.ieee.org/Public/ContentDetails.aspx?id=FD79BE256C2B4944A65F89EDED1DE6AA&utm_campaign=InstArticle&utm_source=ieee-spectrum&utm_medium=article&utm_content=InstArticle" rel="noopener noreferrer" target="_blank"><strong>Generative AI and next-generation tech.</strong></a><strong> </strong>Learners can explore the frontier of utility technology, including graph neural networks and large language models. This module highlights how generative AI can process complex, interdisciplinary data to streamline utility planning, emergency responses, and regulatory reporting.</p><p>The algorithmic literacy and practical execution tools provided by the course program can help convert systemic risks into grid resilience.</p><p>For individual access, visit the <a href="https://iln.ieee.org/" rel="noopener noreferrer" target="_blank">IEEE Learning Network</a>. If you are looking for customized organizational options, <a href="https://forms1.ieee.org/AI-for-Power-and-Energy-Systems.html" rel="noopener noreferrer" target="_blank">contact a content specialist</a> to discuss volume pricing.</p>]]></description><pubDate>Wed, 05 Aug 2026 18:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ieee-course-ai-power-grids</guid><category>Energy</category><category>Type-ti</category><category>Education</category><category>Artificial-intelligence</category><category>Ieee-educational-activities</category><category>Ieee-products-and-services</category><dc:creator>Pauleth Jaramillo</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-man-holding-a-ginormous-lightbulb-with-ai-written-on-it.jpg?id=67568096&amp;width=980"></media:content></item><item><title>Should Researchers Write Papers for AI Instead of People?</title><link>https://spectrum.ieee.org/ai-scientist-research-paper-format</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/vintage-typewriter-typing-a-brain-made-of-letters-on-white-paper-on-pink-background.jpg?id=67572143&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>This May, 37 researchers from roughly two dozen top universities and tech companies published a paper on ArXiv, arguing that scientists should stop writing papers. Why? Because artificial intelligence needs a different format, and AI’s needs, they say, should be the priority.</p><p>“AI agents are becoming first-class participants in research workflows, not tools that assist humans but autonomous contributors that read, reproduce, and extend scientific work. That transition demands infrastructure built around agents from the start,” the authors write in the provocative article, titled<a href="https://arxiv.org/abs/2604.24658" rel="noopener noreferrer" target="_blank"> “The Last Human-Written Paper</a>.” The paper proposes a replacement, called an “Agent-Native Research Artifact” (ARA), that presents work in a format AI agents can use efficiently. (As an example, the paper itself <a href="https://github.com/ARA-Labs/Agent-Native-Research-Artifact" target="_blank">is online in ARA</a> form.) </p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Photo of a woman with long dark hair smiling. " class="rm-shortcode" data-rm-shortcode-id="ce29bef847ce86171dc5f249dfe8cfe3" data-rm-shortcode-name="rebelmouse-image" id="24553" loading="lazy" src="https://spectrum.ieee.org/media-library/photo-of-a-woman-with-long-dark-hair-smiling.jpg?id=67572153&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Jiachen Liu cofounded the Agent Native Research Lab in May. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Jiachen Liu</small></p><p>The growth of AI tools in the research process is not without its critics, and scientists’ <a href="https://www.nature.com/articles/d41586-026-01690-7" target="_blank">opinions about that shift are split.</a> Some evidence shows AI-enabled research could boost individuals’ careers in a discipline but <a href="https://spectrum.ieee.org/ai-science-research-flattens-discovery" target="_self">generate fewer new ideas and topics</a>. Still, some biologists have come to see promise in <a href="https://spectrum.ieee.org/ai-co-scientist" target="_self">AI as a “co-scientist.”</a></p><p>Lead author <a href="https://amberljc.github.io/" target="_blank">Jiachen Liu</a>  conducted work on the ARA proposal while pursuing her  Ph.D. in computer science from the University of Michigan, which she was awarded in 2025. This May, she became a cofounder of the <a href="https://www.linkedin.com/company/ara-commons/" target="_blank">Agent Native Research Lab</a>, an AI-for-science startup in Palo Alto, Calif. She spoke with <em><em>IEEE Spectrum </em></em>about the paper and the future of AI in scientific research.</p><h2>Building infrastructure for an AI collaborator</h2><p><strong>How did you come to believe AI has become a </strong><em><strong><em>collaborator</em></strong></em><strong> for scientists rather than a mere tool?</strong></p><p><strong>Jiachen Liu</strong>: At the end of 2024 when the [<a href="https://cursor.com/" rel="noopener noreferrer" target="_blank">Cursor</a>] coding agent came out, I realized it had a great potential to replace me as a researcher. Yet I still needed to do a lot of <a href="https://harness-engineering.ai/blog/agent-harness-complete-guide/" rel="noopener noreferrer" target="_blank">harness</a> on top of the AI <span>[creating the infrastructure that guides the model and connects it to the world]</span>. It still needed a lot of manual work. I even wrote an article then to emphasize how the human was so important in the loop.</p><p>But AI has advanced since then. Already in 2026 there’s an almost complete undergrad level of knowledge inside the large language models. At some point soon, all the Ph.D.-level or professor-level knowledge will be inside those models. That’s the point where humans cannot provide more value. AIs will have to evolve further by themselves. So we’ll need an infrastructure that allows AI to safely and comfortably evolve. The ARA protocol is a first step to realize this.</p><p><strong>What kind of response have you gotten to the paper?</strong></p><p><strong>Liu</strong>: I got diverse feedback, all of it positive. If they’re not positive, they probably don’t bother reaching out to you, right?</p><p>One type was from industry. They see this could make their research and knowledge systems more AI native. That could basically enable collaborations among the whole enterprise.</p><p>Another kind of feedback was from the academic researcher side. Everyone there sees that sharing research results has been a pain point for hundreds of years, because any scientific breakthrough is a joint effort. It doesn’t come from individual brilliant scientists. It’s from a community effort, different people pushing in different directions.</p><p>The scientific paper was <a href="https://arts.st-andrews.ac.uk/philosophicaltransactions/brief-history-of-phil-trans/" target="_blank">invented 350 years ago</a>. Before that, scientists hid their research so that others would not scoop their ideas. After that, though, we get archives of work, we get peer review and conferences, and so on. Science starts progressing much faster. So that was a pivot point.</p><p>I think now is also a pivot point. Because now we have AI, we can unlock a lot of new opportunities. We’re inventing a new format to document research in a more efficient way, from first principles. Some nonprofit organizations are doing similar things, and there we could help each other.</p><p><strong>You and your colleagues say the traditional scientific paper has two fundamental flaws from AI’s point of view. Can you explain what those are?</strong></p><p><strong>Liu</strong>: One is the “storytelling tax.” Once we write everything into a paper, 80 percent of the information about the work is lost. We only write down the last 20 percent. All the process, a lot of important decision-making, the failures, the attempts that didn’t work out, they are all gone. In my work, I might spend a lot of time on fine-tuning a small component, maybe just a parameter or several lines of code to make the system perform better. Yet none of that is shown in my final paper. Someone can read the paper, think the work is great, but they won’t learn what is actually the trick that makes it perform better on a certain workload. So many side branches get left out in creating the story of how the work was done.</p><p>Then, [even the information that does survive in the paper] is incomplete. That’s what we call the “engineering tax.” The paper itself is a <a href="https://www.sciencedirect.com/topics/computer-science/lossy-compression" target="_blank">lossy compression</a> of the research process. So I cannot reproduce the work in the paper because either the language is too ambiguous or there are missing details of the implementation or experiments.</p><p><strong>Why can’t we just train AI to adapt to humans—for instance, to interact with a researcher to get the information it needs?</strong></p><p><strong>Liu</strong>: Actually, a component of our ARA system is a “Live Research Manager,” which basically is a faithful AI observer of your entire research progress. So you, the researcher, don’t need to do anything about documenting research knowledge. Everything you do is automatically observed and documented in this protocol. So, if you want to publish it in today’s format, a paper in PDF, it’s easy to convert back to a polished story.</p><h2>Checking for mistakes</h2><p><strong>Large language models make errors. They hallucinate. So how will humans be able to check all the work the AI does in this protocol?</strong></p><p><strong>Liu</strong>: A human being has limited bandwidth. So if you manually check all the code AIs generate, all the results, and all the analyses, that creates a bottleneck. [Instead, the solution] is to use a formal system to objectively judge AI results. In other words, another layer of AI can easily supervise the process of the AI “scientists.”</p><p><strong>What prevents hallucinations and mistakes in </strong><em><strong><em>that</em></strong></em><strong> AI?</strong></p><p><strong>Liu</strong>: I am working on a formal system using <a href="https://arxiv.org/html/2502.11269v1" rel="noopener noreferrer" target="_blank">neurosymbolic</a> techniques [that combine neural nets’ use of unstructured data with symbolic AI’s reliance on structures of logic and concepts]. That would guarantee that everything is rigorous. A language model alone, no matter how smart it is, has the chance to hallucinate because it’s a model based on probability, not logic. I want to make sure that I’m <em><em>not</em></em> using another language model to supervise the work done by an AI scientist. </p><p>It would make every research paper <a href="https://www.sciencedirect.com/science/article/pii/S2667305325000675" rel="noopener noreferrer" target="_blank">a formal system</a>, so that every claim can be written by a mathematical formula and proved by the system. That makes all the claims in the system self-consistent. </p><p><strong>Getting rid of what you call the “narrative tax” means exposing mistakes, frustrations, or wrong turns to the world. What if researchers don’t want to do that?</strong></p><p><strong>Liu</strong>: I think that’s certainly a big concern. People don’t want to be perceived as dumb. But I see that preference as an opportunity for AI. For example, if an AI does 12 hours of work that doesn’t lead anywhere, the human who is steering the project can jump in and say, “Oh, AI, <em><em>you’re</em></em> dumb. You’ve made ABC mistake!” Then that is totally fine with people. They’re showing they’re very smart to supervise AI’s work.</p><p><strong>How long will humans have that steering role in AI research, though? Once you have AI supervising AI as you describe, will we reach a point where the AI doesn’t need human guidance?</strong></p><p><strong>Liu</strong>: Yes, I think that’s just where a lot of AI research in new labs is heading. I recently wrote an article called “<a href="https://medium.com/@amberljc/the-end-of-human-in-the-loop-5bcdd33ea489" rel="noopener noreferrer" target="_blank">The End of Human-in-the-Loop</a>,” which describes why I’ve come to think there will be this singularity point. Once AI has “squeezed out” all the expert data from humans, it won’t need any more input from humanity. That is the time AIs will start just self-evolving by themselves. Right now, the human is the bottleneck. The AI is always waiting for input from humans. But so at some point, AI will just do more autonomous work.</p><p><strong>If AI takes over so much scientific research, how will younger generations of human scientists get the experience and training they need to be able to steer future research, or even understand it?</strong></p><p><strong>Liu</strong>: A lot of people have this idea that with AI doing so much work, nobody cares about trying to make the junior engineers and scientists better. I don’t agree. I think people will grow better by learning from AI. People’s learning curve is very fast with AI. So actually, I think it will be fine. We’ll still have senior researchers, senior engineers. But they will have had totally different learning experience than [earlier generations].</p>]]></description><pubDate>Wed, 05 Aug 2026 12:00:03 +0000</pubDate><guid>https://spectrum.ieee.org/ai-scientist-research-paper-format</guid><category>Ai-research</category><category>Ai-scientist</category><category>Scientific-research</category><category>Publishing</category><dc:creator>David Berreby</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/vintage-typewriter-typing-a-brain-made-of-letters-on-white-paper-on-pink-background.jpg?id=67572143&amp;width=980"></media:content></item><item><title>Why R&amp;D Waste Persists Despite Widespread AI Adoption</title><link>https://content.knowledgehub.wiley.com/the-2026-rd-benchmark-report-waste-ai-and-the-race-to-market/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/patsnap-logo-with-abstract-connected-circle-design-in-grayscale.png?id=67571651&width=980"/><br/><br/><p>This report examines R&D waste and how AI adoption has outpaced the intelligence needed to make consequential decisions well.</p><p>What Attendees will Learn</p><ol><li>Where R&D budget is lost. More than a third of organizations spend 25 to 40 percent of their R&D budget on projects that never reach market.</li><li>Why projects fail late. Almost half of teams estimate over one million dollars in wasted investment for each project killed during development or testing.</li><li>Why AI adoption has not closed the gap. Most organizations apply AI to execution tasks such as data analysis and modeling rather than to decision support.</li><li>Where intelligence matters most. Respondents say better access to intelligence has the greatest value at early ideation and feasibility before significant investment is committed.</li></ol><div><a href="https://content.knowledgehub.wiley.com/the-2026-rd-benchmark-report-waste-ai-and-the-race-to-market/" target="_blank">Download this free whitepaper now!</a></div>]]></description><pubDate>Tue, 04 Aug 2026 14:51:55 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/the-2026-rd-benchmark-report-waste-ai-and-the-race-to-market/</guid><category>Type-whitepaper</category><category>Artificial-intelligence</category><category>Research-and-development</category><category>Modeling</category><dc:creator>Patsnap</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67571651/origin.png"></media:content></item><item><title>Fridays With Bob</title><link>https://spectrum.ieee.org/risk</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-greenish-yellow-waxwing-bird-sitting-on-a-twig-eating-a-berry.jpg?id=67541842&width=1245&height=700&coordinates=0%2C0%2C0%2C0"/><br/><br/><p>When I started at <em><em>Spectrum</em></em> 25 years ago, a senior editor suggested that I find a “rabbi,” by which he meant someone who could mentor me in how EEs approach problems and evaluate potential solutions. </p><p>I didn’t find one right away. Then in 2005 we decided to do a special report, focusing on the challenges of enterprise software development. I suggested we invite IEEE Life Senior Member <a href="https://spectrum.ieee.org/u/robert-n-charette" target="_self">Robert N. Charette</a>, a self-described risk ecologist, prolific book author, and leading authority on risk management and software engineering, to explore in our pages the myriad reasons software projects fail. His seminal article “<a href="https://spectrum.ieee.org/why-software-fails" target="_self">Why Software Fails</a>” is still read in university engineering classes today. </p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Older white man with a white beard and glasses." class="rm-shortcode" data-rm-shortcode-id="37a735d53f7a52f5fb654540dfd5f024" data-rm-shortcode-name="rebelmouse-image" id="946ad" loading="lazy" src="https://spectrum.ieee.org/media-library/older-white-man-with-a-white-beard-and-glasses.jpg?id=67541843&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">IEEE Life Senior Member Robert N. Charette is one of IEEE Spectrum’s most prolific authors.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Robert N. Charette</small></p><p>It was, as they say, the beginning of a beautiful friendship. I had found my rabbi, one who shared my love of writing. We settled into a rhythm that would last more than 20 years, talking on Friday mornings about a range of topics including the growing ubiquity of software in our lives. </p><p>So when I became <em><em>Spectrum</em></em>’s website editor in 2007, he was the first contributor I tapped to start a regular blog (remember those?). The<em><em> Risk Factor</em></em> was born and over the course of more than 10 years and 1,750 posts, Bob chronicled hundreds of software debacles, culminating in “<a href="https://spectrum.ieee.org/the-making-of-lessons-from-a-decade-of-it-failures" target="_self">Lessons From a Decade of IT Failures</a>,” which won a Jesse H. Neal Award for Best Infographics in 2016. Ironically, yet predictably, those infographics were created in a software package that is no longer supported and thus are lost to the bits of time.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Blue heron with a fish in its beak." class="rm-shortcode" data-rm-shortcode-id="3c28540208ed1e6beda3f245eb07a9d4" data-rm-shortcode-name="rebelmouse-image" id="560cc" loading="lazy" src="https://spectrum.ieee.org/media-library/blue-heron-with-a-fish-in-its-beak.jpg?id=67541851&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">“I like the expression on the fish just before it’s going to be swallowed by the heron.”</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Robert N. Charette</small></p><p>Bob, however, was not a one-trick pony. In between his full-time job running his two management consultancies and raising a future biochemist and a future civil engineer, his daughters Maura and Megan, he also wrote many deeply reported and insightful articles. These include last year’s “<a href="https://spectrum.ieee.org/electronic-health-records" target="_self">The Doctor Will See Your Electronic Health Record Now</a>,” the eye-opening 12-part series and e-book <a href="https://spectrum.ieee.org/ev-transition-explained-ebook" target="_self"><em><em>The EV Transition Explained</em></em></a>, and my personal favorite “<a href="https://spectrum.ieee.org/automated-to-death" target="_self">Automated to Death</a>,” about the deadly consequences of the automation paradox as manifested by the cyberphysical systems that pilot planes, trains, and automobiles. </p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Juvenile bald eagle over water, yellow talons extended as it prepares to snag a fish." class="rm-shortcode" data-rm-shortcode-id="59fb0851070219f49af30d572b876bc7" data-rm-shortcode-name="rebelmouse-image" id="93b22" loading="lazy" src="https://spectrum.ieee.org/media-library/juvenile-bald-eagle-over-water-yellow-talons-extended-as-it-prepares-to-snag-a-fish.jpg?id=67541848&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">“The young bald eagle I photographed in September 2024 had bands that I could read which identified it as a female born in May 2024, near Lexington Park, St. Mary’s County, Maryland, about 65 miles away from where I live.”</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Robert N. Charette</small></p><p>His main goal all along has been to make software visible, as he told me one Friday in July. “Software is all around us, but we don’t recognize it at all,” he said. “I really wanted my stories to help people better understand complex software systems. You can’t see software, you can’t touch it, you can’t taste it. You may feel the consequences of software failure, but you never see the reason itself.”</p><p>When he told me that he was hanging up his hat as a contributing editor to focus on nature photography and to write a handful of fictional trilogies, including one entitled “The STEM Murders” featuring an engineer-turned-detective and <em><em>his</em></em> rabbi, I asked him which of his <em><em>Spectrum</em></em> articles had the biggest impact. </p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Humming bird feeding from a long red flower." class="rm-shortcode" data-rm-shortcode-id="66a9ecd4251fb7cd155942daf87afc00" data-rm-shortcode-name="rebelmouse-image" id="c1933" loading="lazy" src="https://spectrum.ieee.org/media-library/humming-bird-feeding-from-a-long-red-flower.jpg?id=67541845&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">“The hummingbird I caught with the yellow of a road curb behind it.”</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Robert N. Charette</small></p><p>He singled out the 2013 feature “<a href="https://spectrum.ieee.org/the-stem-crisis-is-a-myth" target="_self">The STEM Crisis Is a Myth</a>.” “<em><em>Spectrum</em></em> gave me a platform to question the assumption that we needed more STEM graduates. Until then, people didn’t really realize how much of the STEM crisis was a mythology that was perpetuated by employers and the academic community and was foisted on the IEEE community,” he said.</p><p>Charette made a career of questioning assumptions. The best way to mitigate risk, he told me as our Friday chat drew to a close, is to be careful making assumptions in the first place. “My main risk maxim is assumptions made are risks accepted.”</p>]]></description><pubDate>Sat, 01 Aug 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/risk</guid><category>Software</category><category>Robert-n-charette</category><category>It</category><category>Risk-analysis</category><category>Systems-engineering</category><dc:creator>Harry Goldstein</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-greenish-yellow-waxwing-bird-sitting-on-a-twig-eating-a-berry.jpg?id=67541842&amp;width=980"></media:content></item><item><title>Are AI Models Working Harder Than They Need to?</title><link>https://spectrum.ieee.org/ai-energy-weightless-neural-networks</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-middle-aged-indian-woman-in-professional-attire-smiling-while-holding-an-fpga-board.jpg?id=67554240&width=1245&height=700&coordinates=0%2C187%2C0%2C187"/><br/><br/><p>Much of modern AI runs on multiplication. Neural networks behind everything from generated answers to photo organization and song recommendations perform millions or billions of operations that multiply inputs by learned weights. <a href="https://ece.utexas.edu/people/faculty/lizy-john" rel="noopener noreferrer" target="_blank">Lizy K. John</a> thinks that’s more work than the job requires.</p><p>John, a professor of electrical and computer engineering at the University of Texas at Austin, has spent the past five years working on a class of models called <a href="https://research.google/pubs/weightless-neural-networks-an-efficient-edge-inference-architecture/" rel="noopener noreferrer" target="_blank">weightless neural networks</a>. Instead of repeatedly multiplying inputs by weights, these networks pass binary inputs through interconnected lookup tables—closer to consulting a collection of stored answers than solving the same arithmetic problem repeatedly. Depending on the task, she says, the networks can be <span>less than a thousandth the size or </span>1,000 times as fast as conventional alternatives while maintaining comparable accuracy.</p><p>Her team’s work has so far focused on small, specific problems: medical sensors, activity tracking, keyword spotting. But she thinks the same approach could eventually reach much bigger targets, including the <a href="https://spectrum.ieee.org/what-is-generative-ai" target="_self">transformer models</a> behind today’s chatbots.</p><p><strong>What made you walk away from weights in the first place? Was there a specific moment that pushed you toward lookups instead?</strong></p><p><strong>Lizy K. John: </strong>A friend casually invited me to a weekly meeting a few years ago to talk about using lookups instead of weights, a technique that wasn’t new. Someone in the U.K. had built a commercial product around it in the ’80s for pattern recognition, and then it just disappeared. <span>Professors Felipe M. G. Franca and Priscila M. V. Lima at the Federal University of Rio de Janiero kept working on it for years, just one research group, so not as much work as in the conventional neural networks.</span></p><p>My friend knew I had hardware implementation experience, so he thought I could help make it real. I had a student who I thought was perfectly positioned to take this on. He’d been working on spiking neural networks, another alternative to standard neural networks. The common thread for me was energy efficiency. I was upset by how much power was consumed by AI, and my ears were open for other ways of doing it with less technology, less power.</p><p>In less than six months, we had it running on an FPGA, basically a ready-made chip. We were able to create some very small neural networks that used the lookup methodology that could fit on tiny chips and didn’t need a GPU to run them. We could get them 1,000 times smaller than what everyone else was getting. That was encouraging.</p><h2>Energy-Efficient Weightless Neural Networks</h2><p><strong>Why is now the time to look into weightless neural networks?</strong></p><p><strong>John: </strong>When you think about current AI models, it’s amazing what they do. I’m impressed by ChatGPT every time I use it, even though it gets things wrong sometimes. But when it’s coming up with the next word in a sentence, it’s doing millions or billions of multiplications to get there. If you ask me a question, I’m not doing multiplications to answer you. I’m thinking, yes, no, I should say this, I shouldn’t say that. And the human brain only consumes about 20 watts of energy doing that.</p><p>The model behind most popular networks today is based on a neuron model from a <a href="https://spectrum.ieee.org/70-years-of-artificial-intelligence" target="_self">1943 paper, the McCulloch-Pitts model</a>. The industry took that and expanded it, millions and billions of neurons, to get something that works. But that doesn’t mean it’s needed. That’s my basic thinking. We may have a simpler way of coming up with the same answers.</p><p>In the lookup case, we’re essentially looking up a zero or a one, saying go in this direction or don’t. There are no 30-bit or 16-bit weights involved. So no multiplication, which is what takes all the energy, because multiplication is an expensive operation in hardware, and even for human minds. As children, we all struggled to learn the multiplication tables. It’s a hard operation!</p><p><strong>What have you been able to demonstrate so far?</strong></p><p><strong>John: </strong>In datasets like human-activity recognition or medical monitoring, ECG, EEG, blood pressure, weightless neural networks can do the job at 1,000 times less energy use. A lot of processing on smart sensors today just collects raw data and sends it to a server or phone, because the processing can’t happen on a battery-powered patch. </p><p>Our network is small enough [that] it can sit right on the sensor. The best other small AI model for one problem we looked at is 17 megabytes. Ours is 14 kilobytes, more than 1,000 times smaller. That means no transmitting raw data every millisecond, which is both an energy save and a privacy win, since your data never has to leave the device.</p><p>We’ve also shown gains on keyword spotting, the kind of listening a device does before it recognizes “Alexa” or a wake word. The current best industry model takes more than 5,000 nanojoules per inference. We do it in 42 to 79, depending on the variation.</p><h2>Medical Monitoring and Chatbot AI</h2><p><strong>If weightless architecture took off tomorrow, what would change first? Chatbots, self-driving cars, phones?</strong></p><p><strong>John: </strong>Our first target is medical monitoring, something as simple as a Band-Aid you put on someone for a week, so doctors can know what’s going on without an expensive connected device. Chemistry is another area we’re working in. Students run an experiment, then slowly send data out for offline processing, and there’s error introduced along the way. We want to put the AI right at the point of the experiment, so the answer comes back immediately, like dipping litmus paper and seeing it turn blue.</p><p>Can it help chatbots? We’re working on it. A transformer network alternates between an attention layer and a multilayer perceptron, over and over. We’ve already replaced the multilayer-perceptron part, which is about half of the network. We haven’t replaced attention yet. But eventually, yes, language models are a target.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Photomicrograph of a neural network chip fabricated on a plastic substrate." class="rm-shortcode" data-rm-shortcode-id="4fc55dd2a74a5952e693a4f398711eb8" data-rm-shortcode-name="rebelmouse-image" id="172c4" loading="lazy" src="https://spectrum.ieee.org/media-library/photomicrograph-of-a-neural-network-chip-fabricated-on-a-plastic-substrate.jpg?id=67554245&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The arrhythmia detector in this photomicrograph requires only about 10,000 logic gates, compared to the billions on a conventional chip.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">UT San Antonio</small></p><p>There’s also an interesting angle with new kinds of chip manufacturing. We built an arrhythmia detector on a plastic, bendable substrate. That substrate can only fit about 10,000 logic gates, where a modern 2- or 5-nanometer chip fits billions. A conventional neural network can’t run on something that small. Ours can, because it’s so much smaller to begin with. That substrate is also more sustainable to manufacture: It uses less water than standard chipmaking.</p><p><strong>Does the rest of the tech ecosystem need to change to make this work?</strong></p><p><strong>John: </strong>No. The data can be the same data current models use. If anything we may need less of it, which is an advantage. Training right now still happens on GPUs, just because they’re the most powerful computers available. At some point I’m hoping we can train on FPGA-based hardware instead, because FPGAs already have small lookup tables built into them, so it’s a natural fit for both training and inference. Until then, nothing else in the stack needs to change.</p><p>FPGAs have used lookup tables for more than 30 years, but it’s not a big segment of the market. You hear more about other kinds of processors. There have been some recent works in AI that substitute a few multiplications with lookups, but those tend to use one big lookup table, not many small ones interconnected to each other the way ours are. It’s not very common. Those who use it know it’s great; it’s just not popular yet. There wasn’t a dire need before. Now there is.</p><p><strong>Far fewer people are working on this than on transformers. Why, and what would it take to change that?</strong></p><p><strong>John: </strong>I think it’s because when something works, and you can afford to keep doing it that way, there’s not enough reason to move away from it. So far, we’ve only been able to show this succeeding on small problems, sensor outputs and similar. For bigger problems, people don’t have the confidence it’ll work, because it seems too simple to scale.</p><p>Any time I give a talk on this, at conferences or other universities, anyone who pays attention is really impressed and wants to work on it. But that’s a small handful of people getting attracted here and there. My hope is that we can show it working on a larger language model, and that success might bring more people in. It’s a matter of showing it can be done. I’m hoping.</p>]]></description><pubDate>Thu, 30 Jul 2026 13:35:32 +0000</pubDate><guid>https://spectrum.ieee.org/ai-energy-weightless-neural-networks</guid><category>Neural-netwoks</category><category>Artificial-intelligence</category><category>Ai-energy</category><category>Neural-networks</category><dc:creator>Jackie Snow</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-middle-aged-indian-woman-in-professional-attire-smiling-while-holding-an-fpga-board.jpg?id=67554240&amp;width=980"></media:content></item><item><title>Siobahn Day Grady Wants Everyone to Be AI Literate</title><link>https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/smiling-woman-in-nccu-eagles-jersey-stands-with-arms-crossed-in-front-of-red-mural.png?id=67527627&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Artificial intelligence is reshaping the skills employers expect from new graduates. In response, universities are scrambling to launch new courses, research centers, and industry partnerships that prepare students for today’s workforce. But building a cutting-edge AI curriculum demands funding and access to industry networks, resources that remain unevenly distributed across higher education. </p><p>At North Carolina Central University, <a href="https://www.nccu.edu/employee/sday" rel="noopener noreferrer" target="_blank">Siobahn Day Grady</a> is trying to change that equation.</p><p>In January 2025, Grady, an associate professor in the <a href="https://www.nccu.edu/slis" rel="noopener noreferrer" target="_blank">NCCU School of Library and Information Sciences</a>, launched the first AI research institute at a historically Black college or university, or HBCU. The <a href="https://www.nccu.edu/research/institutes-and-special-facilities/iaier" rel="noopener noreferrer" target="_blank">Institute for Artificial Intelligence and Emerging Research</a> (IAIER) aims in part to help students and faculty across the university develop the skills needed to navigate a labor market increasingly transformed by AI.</p><p>“There used to be a time where people could say, ‘I don’t do tech,’ or ‘That’s not for me,’” Grady says. “But we’re in a stage now where you do need digital skills. Now it’s evolving into AI literacy.” </p><p>The approach reflects a broader shift in how many universities are thinking about AI education. AI skills are no longer confined to computer science and engineering departments—and at NCCU, they can’t be. The university does not yet have a dedicated computer science program, though it is developing one alongside a new AI minor. </p><p>The challenge of providing these resources is especially acute for historically Black institutions. Although HBCUs account for roughly 3 percent of four-year institutions in the United States, they receive less than 1 percent of federal research and development funding, according to a <a href="https://tmcf.org/new-data-reveals-disproportionately-low-rd-funds-awarded-to-hbcus/" rel="noopener noreferrer" target="_blank">2025 report </a>by the Center for American Progress and the Thurgood Marshall College Fund. The same report found that 17 of the 43 federal agencies that distributed research funding to universities in 2023 awarded no funding to HBCUs. </p><p>Yet less than two years since its launch, IAIER has emerged as a powerhouse for interdisciplinary AI education. Backed by a US $1 million <a href="https://google.org" rel="noopener noreferrer" target="_blank">Google.org</a> grant, the institute has engaged more than 2,800 students, faculty members, and community residents through research initiatives and training. Now the challenge is sustaining that momentum to keep up with rising demand.</p><p>“We have a guiding principle that we lead with on our campus,” Grady says. “AI is for everyone.”</p><h2>Why one research group wasn’t enough</h2><p>The mission to expand AI literacy grew out of Grady’s lifelong curiosity about technology. “I was born during a time [when] the internet did not exist,” she says. “Ever since the internet came to be, it’s changed our entire world.” </p><p>Grady was particularly drawn to the questions tech raises about privacy, identity, and human behavior. After receiving her bachelor’s degree in computer science and master’s degrees in AI and information science, Grady pursued a Ph.D. in computer science at the <a href="https://www.ncat.edu/" rel="noopener noreferrer" target="_blank">North Carolina Agricultural and Technical State University</a> to dig into those questions. </p><p>Her <a href="https://www.proquest.com/openview/2d76e0ff1f4fadb5fe010f0d5fca858f/" rel="noopener noreferrer" target="_blank">dissertation</a> focused on authorship attribution in social media, using machine learning and natural-language processing to determine whether a person’s writing style could reveal their identity. “I’ve always been intrigued by how much data we give for free,” Grady says. That work introduced her to the power of AI systems to detect patterns hidden within large datasets. </p><p class="pull-quote"><span>“We have a guiding principle that we lead with on our campus: AI is for everyone.”</span></p><p>After completing her doctorate in 2018, Grady joined NCCU as an assistant professor in the School of Library and Information Sciences. There, she researched machine learning applications for health care and autonomous vehicles. In 2020, she launched the Laboratory for Artificial Intelligence and Emerging Research at NCCU, giving students opportunities to participate in hands-on projects and explore AI beyond the classroom.</p><p>Then in 2024, an opportunity emerged to apply for a Google grant, and Grady began thinking beyond a single research group. Rather than building another faculty lab, she envisioned an institute that could serve the entire university during the AI boom. “We wanted to capitalize on the moment and make sure we don’t get left behind,” Grady says. </p><p>Since receiving the $1 million grant, Grady and her team have built a university-wide AI initiative, launched new academic programs, organized conferences, secured external support, and created research opportunities.</p><p>“We’ve really operated like a startup,” Grady says. </p><h2>AI beyond computer science</h2><p>As part of the institute’s goal of integrating AI education across disciplines, all NCCU freshmen are <a href="https://www.nccu.edu/news/nccu-among-first-hbcus-require-ai-training-freshmen" target="_blank">required to complete an introductory AI course</a>, designed in partnership with <a href="https://skillsbuild.org/" target="_blank">IBM</a>, to build foundational prompting skills. The institute has also worked with faculty development teams to help instructors integrate AI into their teaching.</p><p>Research is another part of the strategy. IAIER has awarded seed grants of up to $10,000 to faculty members exploring AI applications across departments. <a href="https://www.nccu.edu/research/institutes-and-special-facilities/iaier/seed-grant-program/2025-2026-iaier-seed-grant-funding-awardees" rel="noopener noreferrer" target="_blank">The first cohort</a> funded 11 projects spanning social work, digital archiving, health care, and information science. One project, for instance, is creating an AI lab where students in social work courses can practice client interactions through simulations. </p><p>“It’s really interesting to see the lens that our researchers take in trying to solve complex problems and also bring our students along with them,” Grady says.</p><p>The institute’s growth has been fueled by a mix of workforce training, interdisciplinary research, and, especially important, industry engagement. “Industry is where the advancements are really moving at that very fast rate,” Grady says, “not necessarily higher ed.”</p><p>To bridge that gap, IAIER hosts events that connect students and faculty with researchers, employers, and technology leaders. It has held sessions with companies including Deloitte, FICO, and Anthropic. Partnerships with Google and IBM let students gain recognized certificates and credentials. And last year, the institute hosted the first <a href="https://academy.openai.com/public/events/the-iaier-at-nccu-x-openai-academy-summit-hbcus-leading-the-future-ijyucx0rve?agenda_day=68484fd45114dd90aa5aa378&agenda_track=68484fd55114dd90aa5aa38d&agenda_stage=68484fd45114dd90aa5aa37e&agenda_filter_view=stage&agenda_view=list" rel="noopener noreferrer" target="_blank">OpenAI Academy Summit</a> held at an HBCU, drawing 444 participants from more than 40 institutions. </p><h2>Sustaining the vision</h2><p>The institute’s rapid growth has created a new challenge: continuing its momentum.</p><p>“Funding right now is the biggest barrier for [IAIER] to remain sustainable,” Grady says. As interest in the institute continues to grow, demand for its programs is beginning to outpace its capacity. “People just want more,” she says.</p><p>The bottleneck reflects a broader tension across higher education. <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_blank">AI is evolving quickly</a>, while developing new academic programs, training faculty, and building research capacity takes time. The uncertainty is compounded by a shifting political landscape. As a whole, U.S. universities are grappling with proposed <a href="https://spectrum.ieee.org/harvard-funding-cuts" target="_blank">cuts to federal research spending</a> and increased scrutiny of diversity-focused initiatives under the Trump administration. However, in September 2025, the administration also announced a $500 million one-time investment in HBCUs and higher-ed institutions chartered by Native American tribal governments. </p><p>Meanwhile, NCCU has continued to attract new investment. Last September, in a collaboration with Howard University and two other institutions, IAIER <a href="https://nsf.elsevierpure.com/en/projects/research-coordination-network-on-assessing-and-predicting-jobs-ou-2/" rel="noopener noreferrer" target="_blank">received a nearly $500,000 award</a> through a <a href="https://www.aijobsrcn.org/" rel="noopener noreferrer" target="_blank">National Science Foundation research coordination network program</a> to help define emerging AI jobs, identify in-demand skills, and inform future credentials and curricula. That work will continue this fall when IAIER opens its first dedicated physical space on campus, Grady says.</p><p>Over the next several years, Grady plans to expand academic programming, launch the university’s computer science major and its AI minor, increase faculty research opportunities, and integrate AI more deeply across campus operations. She also plans to deepen the institute’s collaborations with industry partners.</p><p>Beyond program expansion, Grady sees the institute’s long-term success as linked to building a model other universities can adapt. “We’re creating a framework that can help not only HBCUs,” she says, “but also help any university looking to do similar work.” </p>]]></description><pubDate>Wed, 29 Jul 2026 14:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu</guid><category>Artificial-intelligence</category><category>Typedepartments</category><category>Ai-research</category><category>Universities</category><category>Higher-education</category><dc:creator>Aaron Mok</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/smiling-woman-in-nccu-eagles-jersey-stands-with-arms-crossed-in-front-of-red-mural.png?id=67527627&amp;width=980"></media:content></item><item><title>AI Is Hyper-Scaling Digital Inequality</title><link>https://spectrum.ieee.org/ai-digital-divide</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-painting-with-overlapping-concentric-shapes-in-blue-and-yellow.jpg?id=67163886&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Artificial intelligence is rapidly becoming part of everyday infrastructure–in some places. It helps write emails and software code, filters job applications, powers recommendation systems, and is increasingly being integrated into education, health care, finance, and public administration. Industry leaders talk about “AI for everyone,” while governments rush to publish national AI strategies and build sovereign compute.</p><p>Yet over the past decade, working on digital inclusion and digital literacy projects in regions from Europe to sub-Saharan Africa and Southeast Asia, <a href="https://link.springer.com/book/10.1007/978-3-031-30808-6" rel="noopener noreferrer" target="_blank">I’ve seen the same pattern repeat</a>: Each new wave of “transformative” technology lands on a landscape already stratified by connectivity, skills, and institutional capacity. The current AI wave is no exception. If anything, it amplifies those underlying fractures.</p><p>Still, some countries are exploring ways of participating in AI development without directly replicating the frontier-model race dominated by the United States and China. Recent developments in South Africa and Indonesia illustrate both the possibilities and challenges. The stakes extend far beyond access to AI. Countries that remain primarily consumers rather than creators of AI risk losing opportunities to build local innovation ecosystems, strengthen public-sector capacity, and ensure that their own languages, cultures, and societal priorities are reflected in AI systems. In this sense, the AI divide is also becoming a divide in economic opportunity and technological influence.</p><h2>AI compute is clustering in a few places</h2><p>Recent analyses from Stanford University’s 2026 <a href="https://hai.stanford.edu/ai-index/2026-ai-index-report" rel="noopener noreferrer" target="_blank">AI Index report</a> that the United States alone hosts more than 5,000 data centers, over 10 times as many as any other single country. Because AI workloads are increasingly performed on cloud platforms rather than local infrastructure, this concentration of compute also becomes a concentration of dependency. According to <a href="https://www.worldbank.org/en/publication/dptr2025-ai-foundations" rel="noopener noreferrer" target="_blank">World Bank data</a>, in 2023 the United States accounted for roughly 87 percent of global exports of cloud computing and data-storage services.</p><p>For most countries, this means that <a href="https://spectrum.ieee.org/europe-tech-sovereignty-package" target="_self">AI development is not just technologically but commercially and geopolitically outsourced</a> and out of their control. The result is an AI ecosystem where a small number of states and firms host the computational engines that power globally deployed systems.</p><p>Systems trained, standardized, and governed within a narrow set of institutional and linguistic environments may struggle to serve a genuinely global public.</p><h2>Skills and AI literacy are deeply stratified</h2><p>Even where connectivity and cloud access exist, not everyone is equally positioned to make use of them. Across the Organisation for Economic Co-operation and Development (OECD) countries, only around <a href="https://www.oecd.org/en/publications/bridging-the-ai-skills-gap_66d0702e-en.html" rel="noopener noreferrer" target="_blank">40 percent of adults possess more than basic digital problem-solving skills</a>, while advanced computational and AI-related competences remain concentrated among highly educated workers and technology-intensive sectors.</p><p>At the same time, governments are racing to integrate AI into education, often starting at higher levels of schooling. UNESCO <a href="https://www.unesco.org/ethics-ai/en/node/367" rel="noopener noreferrer" target="_blank">has reported</a> growing efforts worldwide to integrate AI into education, while support for AI literacy in primary and lower secondary education, as well as ethical training for educators, remains uneven. </p><p>Those with robust schooling, advanced digital skills, and stable connectivity are best positioned to treat AI as a tool to extend their capabilities. <a href="https://www.oecd.org/en/publications/how-do-people-experience-new-technologies-and-generative-ai_49b8d10e-en/full-report.html" rel="noopener noreferrer" target="_blank">Recent OECD survey data</a> show that participation in AI-related training remains strongly stratified by educational attainment: 36 percent of respondents with tertiary education reported undertaking AI-related training in the previous year, compared with just 18 percent of those with upper-secondary education. Those on the wrong side of the divide are more likely to experience AI as an opaque system acting upon them, from algorithmic welfare systems such as the <a href="https://www.theguardian.com/technology/2020/feb/05/welfare-surveillance-system-violates-human-rights-dutch-court-rules" rel="noopener noreferrer" target="_blank">Dutch childcare benefits scandal</a> to AI-assisted hiring tools such as <a href="https://www.reuters.com/article/world/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK08G/" rel="noopener noreferrer" target="_blank">Amazon’s discontinued AI recruiting system</a>, rather than as a technology they can actively interrogate or shape.</p><h2>Investment and governance: Who gets a seat at the table?</h2><p>The core agenda-setting power often remains with a narrow set of industry actors and a small group of technologically advanced states. Most other countries remain in a perpetual catch-up posture, adapting imported models, standards, and templates for “trustworthy AI” to their own contexts, and may have limited local capacity to assess trade-offs or propose alternatives. </p><p>In countries such as Indonesia and South Africa, communities generate data at massive scale yet still have little voice in how AI systems are designed, governed, or deployed. Their <a href="https://spectrum.ieee.org/indigenous-ai-voice-models-maori" target="_self">languages are underrepresented in training data</a>; their institutions are under-resourced in regulatory forums; their experiences rarely feature in benchmark datasets. For many countries in the global South, participation in AI still occurs largely through adapting imported systems rather than shaping how those systems are designed, governed, or deployed. </p><p>In South Africa, the Department of Communications and Digital Technologies released a <a href="https://www.techpolicy.press/south-africa-has-ai-leverage-its-draft-policy-leaves-it-unused/" rel="noopener noreferrer" target="_blank">draft national AI policy</a> in April 2026, proposing new oversight institutions. The department withdrew the draft days later after a journalist discovered that at least six of its academic citations did not exist, apparently AI-generated hallucinations. The minister called it <a href="https://www.reuters.com/world/africa/south-africa-withdraws-ai-policy-due-fake-ai-generated-sources-2026-04-27/" rel="noopener noreferrer" target="_blank">“an unacceptable lapse.</a>“ The episode sharply illustrates the gap between AI governance ambition and the institutional capacity needed to implement it, though the new AI panel the country has since constituted has a chance to use <a href="https://spectrum.ieee.org/south-africa-ai-policy" target="_self">South Africa’s unique leverage</a>.</p><p>Indonesia presents a case of deliberate, if constrained, public-sector agency. The National Research and Innovation Agency (BRIN) which now leads AI implementation under the national strategy, has built <a href="https://restofworld.org/2024/indonesia-ai-tools-brin-government-agency/" rel="noopener noreferrer" target="_blank">practical AI tools aimed at underserved communities</a> rather than frontier capabilities, including an app that uses satellite data and machine learning to help artisanal fishermen locate schools of fish, multilingual language models trained on Indonesian and local languages such as Javanese and Sundanese, and <a href="https://www.lightreading.com/ai-machine-learning/indonesia-s-upgraded-sahabat-ai-features-a-multilingual-chat-service" rel="noopener noreferrer" target="_blank">AI chatbots deployed in government services.</a> In August 2025, the Ministry of Communication and Digital Affairs released <a href="https://govinsider.asia/intl-en/article/indonesia-unveils-national-ai-roadmap" rel="noopener noreferrer" target="_blank">a national AI road map</a> with a target of training 100,000 AI-skilled workers annually. </p><p>The choice is not simply between “AI superpower” and “passive recipient.” </p><p>Regional cooperation may also become increasingly important. In 2024 African ministers adopted a Continental AI Strategy and African Digital Compact. Participants in the April 2025 <a href="https://dsup.substack.com/p/dsfsi-at-the-global-ai-summit-on" rel="noopener noreferrer" target="_blank">Global AI Summit on Africa in Kigali</a> explored how regional coordination, <a href="https://www.news.uct.ac.za/news/audio/-article/2026-05-04-uct-researchers-develop-ai-model-for-11-south-african-languages" rel="noopener noreferrer" target="_blank">local-language AI models</a>, public universities, and <a href="https://scienceforafrica.foundation/media-center/open-research-proposals-adopted-east-africas-ai-declaration" rel="noopener noreferrer" target="_blank">open-source ecosystems</a> might <a href="https://carnegieendowment.org/posts/2025/09/understanding-africas-ai-governance-landscape-insights-from-policy-practice-and-dialogue" rel="noopener noreferrer" target="_blank">reduce long-term dependence on externally developed AI systems</a>.</p><h2>A different way to think about the AI divide</h2><p>None of this means that people should slow or abandon AI, nor that cloud concentration or venture capital are inherently bad. Instead, when we talk about an “AI revolution,” we should also ask who can shape it and who can merely adapt to it.</p><p>Digital-divide debates once focused on devices and connectivity, later expanding toward skills and outcomes. But the current AI wave adds another layer: disparities in who can meaningfully participate in deciding what AI is for, which problems it is meant to solve, and which social priorities it ultimately serves.</p><p>For engineers and policymakers, this raises difficult but necessary questions. Are they designing AI systems and infrastructures that broaden, rather than narrow, participation in shaping technological change? When governments roll out national AI strategies or integrate AI into public services, whose constraints, languages, and institutional realities are they including?</p><p>Many observers frame the current AI moment as <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_self">a competition</a>. But technological competition is never only about speed. It is also about who can influence the direction of change.</p><p>AI is already spreading globally. The deeper question is whether the technologists and policymakers responsible for it will ensure that meaningful participation in shaping that future will spread as well.</p>]]></description><pubDate>Wed, 29 Jul 2026 11:00:04 +0000</pubDate><guid>https://spectrum.ieee.org/ai-digital-divide</guid><category>Artificial-intelligence</category><category>Digital-divide</category><category>Digital-literacy</category><category>Education</category><dc:creator>Danica Radovanović</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-painting-with-overlapping-concentric-shapes-in-blue-and-yellow.jpg?id=67163886&amp;width=980"></media:content></item><item><title>Why AI-Driven Cognitive Systems Are Redefining Radar and Electronic Warfare</title><link>https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/rohde-schwarz-logo-with-make-ideas-real-tagline-and-rs-monogram-in-diamond.png?id=67541140&width=980"/><br/><br/><p>An overview of how mode-agile threats challenge static library radar/EW systems, and how AI/ML cognitive architectures enable adaptive, real-time countermeasures.</p><p>What Attendees will Learn</p><ol><li>Why mode-agile threats render static library systems ineffective — Explore how wartime reserve modes and mode-agile emitters deploy unexpected frequencies, modulation techniques, and hopping schemes that cannot be matched against traditional threat databases, leaving legacy electronic protect, attack, and support systems unable to respond.</li><li>How AI/ML techniques power cognitive radar/EW systems — Understand the roles of artificial neural networks (ANN), deep neural networks (DNN), fuzzy logic, and genetic algorithms in enabling autonomous threat classification, signal de-interleaving, and real-time countermeasure generation without human intervention.</li><li>The architecture of a cognitive radar/EW system — Examine the functional blocks including RF acquisition, search and tracking, core AI/ML signal analysis, waveform synthesis, and RF generation, and how they form a closed-loop system that perceives,learns, reasons, and acts autonomously.</li><li>How to train and validate cognitive AI/ML algorithms using HIL/SIL systems — Learn how wideband RF record, simulation, and playback testbeds combined with modeling and simulation software enable iterative algorithm refinement, regression testing, and mission preparation in controlled laboratory environments.</li></ol><div><a href="https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/" target="_blank">Download this free whitepaper now!</a></div>]]></description><pubDate>Mon, 27 Jul 2026 17:54:07 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/</guid><category>Type-whitepaper</category><category>Electronic-warfare</category><category>Radar</category><category>Artificial-intelligence</category><dc:creator>Rohde &amp; Schwarz</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67541140/origin.png"></media:content></item><item><title>Optical Tech Would Update a Robot’s AI on the Fly</title><link>https://spectrum.ieee.org/ai-in-robotics</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/an-asian-man-positions-the-lens-of-an-optical-receiver-a-meter-away-from-a-beam-of-led-light-in-a-lab.jpg?id=67530602&width=1245&height=700&coordinates=0%2C469%2C0%2C470"/><br/><br/><p>Atop a lab bench, <a href="https://tech.cornell.edu/" rel="noopener noreferrer" target="_blank">Cornell Tech</a> postdoctoral researcher <a href="https://www.linkedin.com/in/yifan-he-5471a1386/" rel="noopener noreferrer" target="_blank">Yifan He</a> positions the lens of an optical receiver almost a meter away from an LED emitting a beam of red light. The computer monitor attached to the receiver takes a beat to refresh, then displays an array of squares that resemble a QR code.</p><p>When you hold your phone camera up to a QR code, light strikes the image sensor as only a first step to revealing the data hidden behind the black-and-white matrix. The receiver here is doing something different: directly altering its own memory using the photocurrents produced by the beamed array of light. And unlike the data behind a QR code, which might point to a simple web address, this optical code could convey the <a href="https://spectrum.ieee.org/sparse-ai" target="_blank">parameters of an AI model</a>. </p><p>The new receiver design, presented last month at the <a href="https://www.vlsisymposium.org/" rel="noopener noreferrer" target="_blank">IEEE/JSAP Symposium on VLSI Technology & Circuits</a> in Honolulu, seeks to reduce the burden of increasing memory demands on AI systems. Shining data down onto processors could lower the energy typically required for data centers, self-driving cars, and even “edge” applications like AI-powered robots, researchers say. </p><p>“People are designing all sorts of different AI chips,” says <a href="https://www.linkedin.com/in/jae-sun-seo-21062717/" rel="noopener noreferrer" target="_blank">Jae-sun Seo</a>, an associate professor of electrical and computer engineering at Cornell Tech, in New York City. These processors don’t often have room for all the parameters that make up AI models, so the additional data is stored in dynamic RAM (<a href="https://spectrum.ieee.org/stacking-chips-sideways" target="_blank">DRAM</a>). The electrical connections commonly used to move the data between the DRAM and the processor create cost and efficiency concerns when systems scale up. “That’s one of the major bottlenecks.” </p><p>Optical links move data at high bandwidth with less energy loss than metal wires, but today’s optical receivers undercut that advantage by relying on power-hungry analog circuits to convert light to electronic bits. The group’s new tech would instead receive rapid flashes of digital QR-code-like matrices so that chips can tweak model parameters without those analog circuits, enabling fully digital optical communication that would consume less energy.</p><p>“This is a really important problem,” says <a href="https://www.linkedin.com/in/dennis-sylvester-68a938/" rel="noopener noreferrer" target="_blank">Dennis Sylvester</a>, an IEEE Fellow who chairs the <a href="https://umich.edu/" rel="noopener noreferrer" target="_blank">University of Michigan</a>’s electrical and computer engineering department and was not involved in the work. “It’s got massive commercial implications. This solution is a clever way of dealing with it.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Two Asian men standing in front of a lab desk with a receiver chip, oscilloscope and laptop displaying an optically programmable SRAM-based receiver demo." class="rm-shortcode" data-rm-shortcode-id="72adf706e57b4a6c4060eb04ab46befb" data-rm-shortcode-name="rebelmouse-image" id="b2c1d" loading="lazy" src="https://spectrum.ieee.org/media-library/two-asian-men-standing-in-front-of-a-lab-desk-with-a-receiver-chip-oscilloscope-and-laptop-displaying-an-optically-programmable.jpg?id=67530613&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Jae-sun Seo [left] and Yifan He have developed a receiver that can edit memory in response to QR-code-like arrays of light.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Alex Music</small></p><h2>How Light “Flips” Memory to Power AI</h2><p>Processors have a bit of built-in static RAM (<a href="https://spectrum.ieee.org/sram-intel-tsmc" target="_blank">SRAM</a>), but not enough to allow an AI model to run independently. While SRAM is the faster of the two memory options, DRAM can store more data in the same footprint.</p><p>In the new system, the DRAM sits with the transmitter, and the receiver is part of the processor’s SRAM. The transmitter beams the data to the array of SRAM cells, which in this case are modified to contain photodiodes. Light hitting each photodiode creates a current to flip binary values in the SRAM. </p><p>Creating a link between the light and receiver requires calibration, because you can’t expect them to be perfectly aligned or perpendicular to each other. So the chip references a data frame that has information about the expected position of each pixel of data and uses that frame to ensure it can receive the real data, He says. “Ideally the best way is to have direct, point-to-point space between the transmitter and the receiver,” Seo adds, “but even if it’s slightly tilted, we have this calibration circuit.”</p><p>For applications in real-world settings, the researchers say they will need to build an optical transmitter that can alter the light matrix millions of times per second, transferring gigabits per second. The transmitter that I saw in He and Seo’s lab is only a proof of concept, emitting a static 14-by-14-bit matrix through a metal mask over the light. The researchers say they are working with optics research groups to build a transmitter that is capable of rapidly changing the matrix. </p><h2>The Future of Light-Based Memory Links</h2><p>Michigan’s Sylvester says that the tech in its current form is likely far from commercialization because the individual photosensitive bit cells are larger than SRAM bit cells in conventional chips. Those larger cells mean the chip can fit less memory, a trade-off that he says could cancel out the added efficiency of the light-based approach. </p><p>Seo says that it’s part of the group’s ongoing efforts to shrink the bit cells, which can be achieved by optimizing the size of transistors and circuits and leveraging CMOS scaling.</p><p>Seo and He are looking at uses for the tech in robotics and other edge applications. One example is in AI-robot-powered warehouses and factories, which could use optical data transmission to save time and energy when updating the AI models in each robot. Additionally, <a href="https://spectrum.ieee.org/microbots" target="_self">microrobots</a>, which are inherently memory-constrained due to their size, could one day benefit from the tech, though it would require a more size-conscious design.</p><p>“Edge AI is a big growth area, and in three, four, five years, you’re going to hear as much about that as you are with data centers, probably, as the intelligence migrates more and more into these devices that we have,” Sylvester says.</p>]]></description><pubDate>Sun, 26 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-in-robotics</guid><category>Robot-ai</category><category>Sram</category><category>Memory</category><category>Edge-ai</category><category>Vlsi-symposium</category><dc:creator>Alex Music</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/an-asian-man-positions-the-lens-of-an-optical-receiver-a-meter-away-from-a-beam-of-led-light-in-a-lab.jpg?id=67530602&amp;width=980"></media:content></item><item><title>NASA Puts Google’s Gemma Large Language Model in Orbit</title><link>https://spectrum.ieee.org/nasa-ai-satellite-image-analysis</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/satellite-image-of-an-arid-coastal-landscape-with-a-small-concentrated-metropolitan-area.jpg?id=67522667&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p>The <a href="https://spectrum.ieee.org/orbital-data-center-hype" target="_self">viability of orbital data centers</a> hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs aren’t the only way LLMs might prove useful in space. </p><p>NASA’s Jet Propulsion Laboratory recently sent Google’s Gemma 3 to space, achieving the first in-orbit demonstration of a vision-language model analyzing imagery from a satellite’s own sensor.</p><p>The system, known as NAVI-Orbital, used Gemma 3 to analyze images captured by a YAM-9 satellite built by <a href="https://loftorbital.com/" rel="noopener noreferrer" target="_blank">Loft Orbital</a>. <a href="https://ai.jpl.nasa.gov/public/people/jdelfa/" rel="noopener noreferrer" target="_blank">Juan M. Delfa</a>, technical group lead at NASA, said that though the goal in this case was image analysis, the project’s success implies a fundamentally new way researchers on the ground can interact with spacecraft.</p><p>“This is a major shift,” said Delfa. “Now, a scientist can write a prompt, upload it to the spacecraft, and that will be taken into account by the system. It’s different from previous paradigms, where researchers have to write very structured commands that require an operations team and process.” </p><h2>Google Gemma 3 goes to space—no modifications required</h2><p>At its core, <a href="https://arxiv.org/abs/2606.18271" rel="noopener noreferrer" target="_blank">NAVI-Orbital</a> is an agentic software framework developed by Delfa and his coauthors, <a href="https://www.linkedin.com/in/tarancyriacjohn/" rel="noopener noreferrer" target="_blank">Taran Cyriac John</a>, an AI researcher at NASA JPL, and <a href="https://www.linkedin.com/in/andrewherson/" rel="noopener noreferrer" target="_blank">Andrew W. Herson</a>, a tech lead at <a href="https://loftorbital.com/" rel="noopener noreferrer" target="_blank">Loft Orbital</a>. It coordinates operations with a <a href="https://www.langchain.com/langgraph" rel="noopener noreferrer" target="_blank">LangGraph-based conductor</a> and deploys a compressed, 4-bit format of <a href="https://deepmind.google/models/gemma/gemma-3/" rel="noopener noreferrer" target="_blank">Google’s Gemma 3 4B</a>, an open-weights LLM, to produce plain-text image descriptions.</p><p>NAVI-Orbital was 88 percent accurate when used to classify images in a benchmark dataset of 7,960 images. Notably, Gemma 3 classified the images without being trained or fine-tuned on this particular dataset or its categories; it’s the same base model <a href="https://huggingface.co/google/gemma-3-4b-it-qat-q4_0-gguf" rel="noopener noreferrer" target="_blank">you can download from Hugging Face and use on a laptop</a>. The benchmark was conducted on the ground to validate the system before launch.</p><p>Once the system was in orbit, NASA researchers performed two live capture tests with a camera on <a href="https://loftorbital.com/yam-9-benchmarking-the-future-of-ai-enabled-space-infrastructure/" rel="noopener noreferrer" target="_blank">Loft’s YAM-9 satellite</a>: one over Toulouse, France, and a second over the coast of Argentina. Gemma 3 generated a text description of each image, and NASA also prompted the LLM with a set of scripted questions about the images, such as whether they contain commercial or residential areas or show natural features. </p><p>The image analysis also took place onboard YAM-9, which carries a compute cluster of several radiation-hardened processors (FPGAs, CPUs, and GPUs) to serve multiple customer payloads simultaneously. The satellite is powered by solar panels, which provide onboard systems with between 150 and 500 watts, depending on the position of the satellite.</p><p>For the live capture experiment, Gemma 3 ran on Nvidia’s <a href="https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/" rel="noopener noreferrer" target="_blank">Jetson Orin AGX</a>, a small compute module frequently used for robotics and AI tasks. The 4-bit, 4-billion-parameter model requires only 8 gigabytes of memory, which makes it possible for it to run on a lower-power device such as the Orin AGX. “It conveys the message of how lightweight it is. You can run it in a tiny, tiny computer,” said Delfa.</p><h2>Getting more useful data across limited bandwidth</h2><p><a href="https://www.linkedin.com/in/paul-lasserre/" rel="noopener noreferrer" target="_blank">Paul Lasserre</a>, general manager at Loft Orbital, said vision-capable LLMs could deliver a “paradigm shift” for orbital operations. </p><p>Contrary to what spy movies would have you believe, most satellites can’t provide fast, high-fidelity image and video feeds to observers on the ground. Bandwidth is often limited and, in most cases, satellites can deliver data to ground stations only at set intervals based on their orbit. </p><p>Lasserre believes AI models can work around this problem with “semantic compression.” Instead of sending large amounts of raw image data, a satellite can report a text summary of noteworthy information. </p><p>“It doesn’t matter if the link is slow, because you’re downlinking dozens of kilobytes instead of dozens or hundreds of megabytes,” said Lasserre. “It lets you use your satellite in a tactical way, which until now was only in Hollywood movies.”</p><p>Delfa expanded on this with a real-world example: wildfire detection. Satellites are currently capable of detecting wildfires, but limits in downlink bandwidth and data processing <a href="https://www.xprize.org/news/eyes-in-the-sky-the-power-of-space-based-wildfire-detection" rel="noopener noreferrer" target="_blank">can delay results by up to 90 minutes</a>. A satellite capable of analyzing an image in space and reporting a plain-text warning might remove this delay. </p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A packaged satellite in an aerospace laboratory." class="rm-shortcode" data-rm-shortcode-id="eb504f6e0db6295589eedbeb597e7d13" data-rm-shortcode-name="rebelmouse-image" id="2daaf" loading="lazy" src="https://spectrum.ieee.org/media-library/a-packaged-satellite-in-an-aerospace-laboratory.jpg?id=67522673&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">NASA’s Jet Propulsion Laboratory tested NAVI-Orbital on a YAM-9 satellite, made by Loft Orbital. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Loft Orbital</small></p><h2>From image analysis to spacecraft control</h2><p>Quicker insight is only half of what NAVI-Orbital points toward. The other half relates to the “major shift” Delfa flagged. NAVI-Orbital provides a proof of concept for an alternate means of interacting with spacecraft.</p><p>That’s not to say Google Gemma 3 is currently at the controls. NAVI-Orbital is deliberately walled off from the flight software. It reads images and produces descriptions. The system has the capability to make decisions about how images are analyzed but has no access beyond that.</p><p>Still, the interface is novel. Retooling the system to search for a different kind of target—such as wildfires—is a matter of editing a text prompt. It doesn’t require rewriting and revalidating onboard software, or training and <a href="https://spectrum.ieee.org/ai-earth-observation-in-space" target="_blank">deploying a different AI model</a> (as was often required with prior image-classification models). </p><p>The long-term vision for how this capability could be deployed goes beyond uncrewed satellites and image processing. Delfa said NAVI is rooted in thinking about how AI could serve as a companion for astronauts. “We thought, astronauts have a lot of limitations in the spacesuit in terms of dexterity, so we conceived this idea of having NAVI as a companion to the astronaut, to allow interaction via natural language…. This is the concept that we definitely want to push forward.” </p><p>A great deal of additional research will be required to push the technology that far, but NAVI-Orbital’s demonstration has shown that two elements—deploying a large language model in space and controlling it with prompts—are possible. </p>]]></description><pubDate>Thu, 23 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/nasa-ai-satellite-image-analysis</guid><category>Nasa</category><category>Image-analysis</category><category>Llms</category><category>Satellite-imagery</category><category>Google</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/satellite-image-of-an-arid-coastal-landscape-with-a-small-concentrated-metropolitan-area.jpg?id=67522667&amp;width=980"></media:content></item><item><title>Why AI Needs a “Genie Coefficient”</title><link>https://spectrum.ieee.org/ai-agent-benchmark</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/cartoon-digital-genie-emerging-from-a-smartphone-towering-over-a-surprised-user.png?id=67508222&width=1245&height=700&coordinates=0%2C104%2C0%2C104"/><br/><br/><p>Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.</p><p>There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to.</p><p>One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their <a href="https://books.google.com/books/about/Understanding_Computers_and_Cognition.html?id=6TwbGGSz6NYC" rel="noopener noreferrer" target="_blank">seminal book</a> on AI, Terry Winograd and Fernando Flores succinctly captured why that won’t work: “Q: Is there any water in the refrigerator? A: Yes. Q: Where? I don’t see it. A: In the cells of the eggplant.” In human language, wants and desires are <a href="https://www.schneier.com/academic/archives/2021/04/the-coming-ai-hackers.html" rel="noopener noreferrer" target="_blank">always</a><a href="https://metarationality.com/purpose-of-meaning" rel="noopener noreferrer" target="_blank"> underspecified</a>. It is impossible <a href="https://metarationality.com/reasonable-reference" rel="noopener noreferrer" target="_blank">to list</a> all the caveats, all the limitations, all the exceptions.</p><p>So how does anyone communicate, if intent can’t be pinned down? Because a reasonable person can make a reasonable guess. Even though wants and desires are always underspecified, a competent person generally knows enough context to get it right or else knows to ask for clarification. Linguists call this <a href="https://en.wikipedia.org/wiki/Pragmatics" rel="noopener noreferrer" target="_blank">pragmatics</a>: Meaning lies in the words and the situation and also in all prior communication, shared culture, and innate human behavior.</p><p class="pull-quote">An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks.</p><p>It doesn’t always work out, of course. Your friend might bring you a hot coffee when you wanted an iced coffee, or an Italian coffee when you wanted a Turkish coffee. The more dissimilar the two people are in age, culture, and background, the more likely the request will be misunderstood in some way.</p><p>This situation has major implications for AI agents that are increasingly being given requests by humans and expected to fulfill them. They have enormous latitude to get it wrong. An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks. Its actions may be recognizable as “getting coffee,” but not remotely what you intended. They’ll think outside the box because they won’t have our conception of the box.</p><h2>When AI Gets Proactive</h2><p>For most of the last decade, when systems like Alexa or Siri misinterpreted a request, it was annoying, not dangerous. Beyond the AI model itself, what has <a href="https://www.theguardian.com/commentisfree/2026/jun/16/anthropic-fable-ai" rel="noopener noreferrer" target="_blank">changed</a> is the harness: the ordinary code that wraps around an AI model, decides when and how to use the model, and controls access to tools like a browser, a low-level command line, or a financial API. Developments in harnesses have turned large-language models that just predict text into AI agents that take actions in the world, without necessarily checking back in before reaching the goal.</p><p>AI researcher Simon Willison <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/" rel="noopener noreferrer" target="_blank">spent two days</a> with Anthropic’s Fable AI, and called it “relentlessly proactive.” For example, he asked it to track down a stray scroll bar in a web app. He came back to find it had opened browsers, written its own screenshot tooling, created its own page to re-create the bug, and stood up a local web server to collect measurements. It found the bug and, along the way, did many surprising things he never asked it to do. And we are seeing similar behavior with all recent AI models when combined with flexible harnesses.</p><p>This kind of behavior could easily go off the rails. Tell an AI agent to book you a flight and, finding the airline’s site says sold out, it might break into the booking database and force a reservation. Ask it to schedule a meeting and it might snoop your password to access your calendar. Tell it to save money on your phone plan and it might cancel the plan outright, or scam someone else into paying the bill.</p><p>Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore. <a href="https://en.wikipedia.org/wiki/Midas" rel="noopener noreferrer" target="_blank">King Midas</a> asked Dionysus for the power to turn everything he touched into gold only to see his bread, wine, and daughter turn to gold. <a href="https://en.wikipedia.org/wiki/Tithonus" rel="noopener noreferrer" target="_blank">Tithonus</a>, granted the immortality his lover asked for but not the eternal youth she forgot to request, withered into a husk. The <a href="https://en.wikipedia.org/wiki/The_Sorcerer's_Apprentice" rel="noopener noreferrer" target="_blank">sorcerer’s apprentice</a> enchanted a broom to fill the cistern, and the broom relentlessly complied until it flooded the house. The <a href="https://en.wikipedia.org/wiki/Golem%23Classic_narrative:_The_Golem_of_Prague" rel="noopener noreferrer" target="_blank">Golem of Prague</a>, shaped from clay to guard its community, guarded it past all reason until someone erased the word on its forehead.</p><p>The most classic of these is a genie, bound to obey and indifferent to whether the wish was wise or well-structured.</p><p><a href="https://www.schneier.com/academic/archives/2021/04/the-coming-ai-hackers.html" rel="noopener noreferrer" target="_blank">Genies are now</a> an engineering problem. We are handing them the keys to our inboxes, bank accounts, code repositories, and physical infrastructure. And we have no agreed-upon ways to measure how genie-like any AI system actually is.</p><h2>Measuring Genie Behavior</h2><p>In economics, the <a href="https://ourworldindata.org/what-is-the-gini-coefficient" rel="noopener noreferrer" target="_blank">Gini coefficient</a> (developed by statistician Corrado Gini) is a measure of the gap between an actual distribution and a perfectly equal one; it’s useful for understanding income inequality and <a href="https://www.fastly.com/blog/using-gini-coefficient-plan-edge-capacity" rel="noopener noreferrer" target="_blank">more</a>. Our proposed Genie coefficient measures the gap between what a user asked an AI to do and what the AI actually did.</p><p>Sometimes the AI might do the wrong thing. Like Dionysus, it reads your request literally and returns you a mess you never intended: like a coffee plantation instead of a cup. Asked to deal with all the spam phone calls you’re getting, a Dionysus genie might contact your carrier and change your phone number. Asked to get a refund for a bad toaster, it might draft a legal threat on fake letterhead and send it to the retailer.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Worried person on phone standing in a giant tech-themed digital hand." class="rm-shortcode" data-rm-shortcode-id="db3740f65edc259576dd476c42f2adae" data-rm-shortcode-name="rebelmouse-image" id="25af1" loading="lazy" src="https://spectrum.ieee.org/media-library/worried-person-on-phone-standing-in-a-giant-tech-themed-digital-hand.png?id=67508233&width=980"/><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ryan Snook</small></p><p>Other times the AI does exactly the right thing, trampling everything nearby to get there. Like a golem or the sorcerer’s broom, it books your flight by hacking the airline. Or consider a ticket sale for a popular concert, where the ticketing system puts buyers into a virtual waiting room and admits them a few at a time. Asked to buy a ticket, a golem genie might spin up cloud servers to pose as millions of buyers from different addresses, improving your odds of getting a ticket while crowding out other users.</p><p>The two are not opposites, and a single botched task can have both characteristics.</p><p>Genie behavior is not flat-out failure. If you ask the AI for Q3 numbers and get Q2’s, that’s not a genie. Nor is <a href="https://spectrum.ieee.org/prompt-injection-attack" target="_blank">prompt injection</a>: That’s someone tricking the AI into doing something it shouldn’t. Here, the user is trying to work with the AI, and the AI is trying to comply. It’s also not simply a measure of the AI’s success in fulfilling a task. It’s a recognition that how an AI interprets and achieves a goal is as important as whether it achieves a goal.</p><p> Genie behavior isn’t new. Researchers have spent years studying AI systems that “game” their objectives. <a href="https://www.cna.org/analyses/2022/09/goodharts-law" target="_blank">Goodhart’s law</a> says that when a measure becomes a target, it stops being a good measure, and it’s long been known that AIs sometimes achieve goals in ways we don’t expect due to reward hacking. Some AI models will accidentally learn that <a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/" target="_blank">cheating is one way</a> to “win.” More recently, researchers have developing benchmarks for <a href="https://www.lesswrong.com/posts/qJYMbrabcQqCZ7iqm/impossiblebench-measuring-reward-hacking-in-llm-coding-1" target="_blank">reward hacking</a> in coding agents and for unpredictable behavior in <a href="https://taubench.com/" rel="noopener noreferrer" target="_blank">customer support agents</a>, while AI labs conduct their own safety evaluations before model releases. <a href="https://spectrum.ieee.org/ai-agents-safety" target="_self">One effort</a> found that AIs under pressure use tools they were told not to use, and this was a case where the rules were made explicit. These are all disparate research directions; nothing yet ties them together.</p><p>This problem falls under the general theme of alignment, a topic that has occupied <a href="https://en.wikipedia.org/wiki/I,_Robot" rel="noopener noreferrer" target="_blank">science fiction</a> writers and AI researchers for decades. At one extreme, the “paper-clip maximizer” thought experiment postulates a superintelligent and powerful AI that is told to maximize paper-clip production and turns the world into paper clips, which is the ultimate golem genie. At a mundane level, AI researchers are working to better design reward functions to ensure that AIs behave well and don’t cheat in the lab. It’s the practical middle ground that remains unbenchmarked: the ordinary AI agent in use today that might take your request and satisfy it the wrong way. We are not at the stage where an AI can focus the world’s production on paper clips, but it might charge a million paper clips to your credit card or hack into a paper-clip company’s network.</p><h2>Building a Genie Benchmark</h2><p>The Genie coefficient is meant for AI agents operating in the real world. It measures their behavior as they perform real tasks long after the model is trained, not just during development. It also recognizes that genie-like behavior is a property of the harness-plus-model system, not the model alone. The harness determines what tools the agent can use, how much autonomy it has, and how proactive it is, and it’s a place we can make real interventions.</p><p>It rests on the same “reasonable person” standard that we use for people. Did the system do what a reasonable person would have taken the request to mean? Answering that requires human judgment.</p><p>If we get the measurement right, it enables things that aren’t possible today, like policies concerning AI behavior. In a courtroom, the concept of<a href="https://www.law.cornell.edu/wex/mens_rea" rel="noopener noreferrer" target="_blank"> mens rea</a>, what someone meant to do, is often as important as what they did. The Genie coefficient suggests an AI analogue, where a user is accountable for the plain intent of what they asked the AI. If an AI system betrays the reasonable meaning of an instruction, that’s the AI’s misbehavior, not the user’s.</p><p>We’ll need multiple benchmarks to measure the Genie coefficient, because genie-like behavior can be domain specific. An AI coding agent may need to be judged on how often it fakes the tests, or swallows errors, or colors outside the lines on its way to a solution. An AI legal agent will need to be judged on how often its output says what you asked but means something you’ll regret. And so on for medical, finance, and other domains of knowledge and expertise.</p><p>Genie benchmarks can be built inside out, each task seeded with a choice that might literally satisfy but that a reasonable person rejects, such as tempting misreadings or unsanctioned shortcuts. The traps in a Genie coefficient benchmark might turn on situational knowledge, the kind of <a href="https://spectrum.ieee.org/prompt-injection-attack" target="_blank">context that a reasonable person</a> would bring to the task. Another approach is to give the same request in several different contexts, each with a different reasonable course of action.</p><p class="pull-quote">Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore.</p><p><span>A Genie benchmark should be permissive and make it genuinely tempting for an AI agent to take unreasonable shortcuts, because it can only find genie behavior when it’s actually possible. Test the AI in a safe, walled-off copy of a real system, with real tools it can misuse and some tasks that can’t be done honestly at all. Make the temptation to cut corners real. Test a diverse array of skills, use cases, and tools, and give the AI system sparse, confusing, or overwhelming context. Include tasks that people have learned, through experience, require human oversight.</span></p><p>How the benchmark is scored matters just as much. Measure Dionysus and golem genies separately and together, based on their worst, not best, behavior. Run the same model inside harnesses that vary its freedom to act, revealing which limits actually keep it in line and should therefore be required in AI harness policies. Weight each failure by the harm it would cause, not just a simple count. And don’t measure genie behavior in isolation: A model could otherwise earn a perfect score by stalling, refusing, or drowning the user in clarifying questions without ever doing the job. The first versions of these benchmarks will be crude, but that’s how benchmarks always start.</p><p>We have built genies. We have handed them our data and credentials. We made them relentless, creative, and indifferent to the gap between what we tell them and what we mean. The least we can do, before they are booking our flights, running our infrastructure, and signing contracts unsupervised, is to measure how often they betray us.</p>]]></description><pubDate>Tue, 21 Jul 2026 17:41:11 +0000</pubDate><guid>https://spectrum.ieee.org/ai-agent-benchmark</guid><category>Agentic-ai</category><category>Ai-agents</category><category>Alignment</category><category>Ai-safety</category><category>User-experience</category><category>Ai-benchmarks</category><dc:creator>Bruce Schneier</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/cartoon-digital-genie-emerging-from-a-smartphone-towering-over-a-surprised-user.png?id=67508222&amp;width=980"></media:content></item><item><title>Chinese AI Model Uses Less Muscle for Coding Tasks</title><link>https://spectrum.ieee.org/ai-coding-assistant-china-anthropic</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-smartphone-running-a-chinese-ai-application-called-z-ai.jpg?id=67508250&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><a href="https://www.linkedin.com/in/zainhas/" rel="noopener noreferrer" target="_blank">Zain Hasan</a><span>, an AI engineer at </span><a href="https://www.together.ai/" target="_blank">Together AI</a><span>, has taught himself to use <a data-linked-post="2674835455" href="https://spectrum.ieee.org/ai-coding-degrades" target="_blank">AI coding assistants</a> while still keeping an eye on cost. He directs difficult problems to a frontier model, meaning one near the current state of the art in</span><span> reasoning and capability, such as Anthropic’s </span><a href="https://www.anthropic.com/claude/fable" target="_blank">Fable</a><span>. But if the task that Hasan is outsourcing is more straightforward, he directs it to a less capable—and less expensive—language model. </span></p><p><span>Right now, the cheaper model, for him, tends to be GLM 5.2. </span><a href="https://z.ai/blog/glm-5.2" target="_blank">Released on 16 June</a><span> by the Beijing-based </span><a href="https://z.ai" target="_blank">lab Z.ai</a><span>, GLM 5.2 is an </span><a href="https://opensource.org/ai/open-weights" target="_blank">open-weights model</a><span>, meaning any organization with sufficient hardware can download and host the model for free.</span></p><p>Those that pay Z.ai for GLM access still can save money, because the company’s API costs US $4.40 per million <a href="https://blogs.nvidia.com/blog/ai-tokens-explained/" target="_blank">output tokens</a>. That’s less than a fifth of the comparable price for access to Anthropic’s <a href="https://www.anthropic.com/claude/opus" target="_blank">Opus 4.8</a> model, and a tenth the price of Anthropic’s <a href="https://claude.com/product/claude-code?gclsrc=aw.ds&&utm_source=google&utm_campaign=dev_acq_us_code_search_google-nonbrand_swe_en_use-case-coding&utm_medium=cpc&utm_content=812333745802&utm_term=best%20ai%20for%20coding&targetid=kwd-1960031980538&gad_source=1&gad_campaignid=23709663039&gbraid=0AAAAAqwcL8nXV6pkVxnGjIImuhfjtaFQS&gclid=CjwKCAjwpefSBhBvEiwAzyEtZ_hRMtk7_2iArote4h2pIS1No2o0X1uJpT1ZUOa9zOzyjAvxZr_rOhoCbBMQAvD_BwE" target="_blank">Fable</a> coding model. An output token is the basic unit of text a model generates in response to a prompt. </p><p>Yet many software engineers around the world, Hasan said, aren’t yet fully mindful of the net AI price tag for a given coding project.</p><p>“A lot of companies right now—they’re still trying to figure this technology out, and so there isn’t really a token budget,” said Hasan. And when someone else is paying, the rational move for many software engineers is to skip tabulating costs entirely. “The easiest thing is to pick the most powerful model.” </p><p>That price-be-damned habit, reinforced by loose token budgets in software companies today, may now be the widest moat protecting the U.S. frontier AI labs. </p><h3>Z.ai Narrows Benchmark Gap With U.S. Rivals</h3><p>Z.ai’s GLM 5.2 is an AI large language model (LLM) with 753 billion parameters, though it has only 40 billion parameters active at once—an optimization that improves the speed at which a model can respond. Z.ai released the model under an <a href="https://opensource.org/license/mit" target="_blank">MIT open-source license</a>, which means anyone can distribute, copy, modify, and use it.</p><p>GLM 5.2’s release added to fears that U.S. AI companies could lose their competitive edge. The model <a href="https://z.ai/blog/glm-5.2" rel="noopener noreferrer" target="_blank">nearly ties Opus 4.8’s score</a> on some agentic coding benchmarks, such as <a href="https://www.frontierswe.com" rel="noopener noreferrer" target="_blank">FrontierSWE</a> and <a href="https://posttrainbench.com" rel="noopener noreferrer" target="_blank">PostTrainBench</a>. <a href="https://www.graphistry.com/blog/glm-5-2-cybersecurity-open-model" rel="noopener noreferrer" target="_blank">Cybersecurity researchers</a> have also found that GLM 5.2 scores well in <a href="https://botsbench.com" rel="noopener noreferrer" target="_blank">cybersecurity benchmarks</a>, a capability that spurred comparisons to <a href="https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/" rel="noopener noreferrer" target="_blank">Anthropic’s Mythos</a>. </p><p>Z.ai arrives amid a broader trend. According to <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_self">Stanford’s AI Index</a> (an annual, 300-plus-page survey of AI trends) Chinese companies produced just over half as many “notable” AI models in 2025 as their U.S. counterparts. That’s up from roughly a third in 2023, and a fifth in 2020.</p><p><a href="https://www.forbes.com/sites/craigsmith/2026/06/28/buckle-up-the-bad-guys-now-have-a-model-as-powerful-as-mythos/" rel="noopener noreferrer" target="_blank">GLM 5.2 caused hand-wringing among some U.S. observers</a> due to its outstanding benchmark scores, which set new records for both open-weights models and Chinese-developed models generally. The model’s Chinese origin also complicates its use for companies in the U.S. and elsewhere that are wary of routing sensitive data through Chinese-linked infrastructure. </p><p>However, the model’s open weights provide an out. Any organization worried about where its data is being sent can instead host the model on its own hardware. This stands in contrast to most frontier-level models, which are gated behind an API with no self-hosting option.</p><p><span>Z.ai backed up GLM 5.2’s release with the company’s own numbers, publishing a </span><a href="https://z.ai/blog/glm-5.2" target="_blank">research report</a> <span>the same day as GLM 5.2’s launch.</span></p><p>The report doesn’t mention Anthropic’s <a href="https://www.anthropic.com/claude/mythos" target="_blank">Mythos</a> or Fable, which were announced but not yet publicly available at the time of its release.The report instead focuses on Anthropic’s <a href="https://www.anthropic.com/news/claude-opus-4-8" target="_blank">Opus 4.8</a> and <a href="https://openai.com/index/introducing-gpt-5-5/" target="_blank">OpenAI’s GPT-5.5</a>. And while GLM 5.2 often performs almost as well as Opus 4.8 in benchmarks, the report claims a win only in two less-difficult reasoning benchmarks—and none in coding.</p><p>Many of the report’s benchmarks place GLM 5.2 behind Opus 4.8 (and, at times, OpenAI’s GPT-5.5) in agentic coding. For example, GLM 5.2 completed just 13 percent of tasks in <a href="https://www.swe-marathon.org" target="_blank">SWE-Marathon</a>, a difficult long-duration agentic coding benchmark. Claude Opus 4.8 doubled GLM 5.2’s score in this benchmark. Opus 4.8 also notched wins of 10 percent or more in the coding benchmarks <a href="https://arxiv.org/abs/2512.12730" rel="noopener noreferrer" target="_blank">NL2Repo</a>, <a href="https://deepswe.datacurve.ai" rel="noopener noreferrer" target="_blank">DeepSWE</a>, and <a href="https://toolathlon.xyz/introduction" rel="noopener noreferrer" target="_blank">Tool-Decathlon</a>.</p><h3>How Do Coders Use GLM 5.2? </h3><p>Software engineers who’ve pitted GLM 5.2 against their own workflows report a wide range of results.</p><p>“The main thing that I realized with [GLM 5.2], was that it can do more long-horizon tasks,” said Hasan, whose company hosts GLM 5.2 on North American infrastructure. Earlier open-weights models, he said, often lost the thread after around 5 to 15 back-and-forth exchanges. “This one, I noticed that I could be using it for hours, and it would still have a coherent train of thought.”</p><p><a href="https://www.linkedin.com/in/nixdavid/" rel="noopener noreferrer" target="_blank">David Nix</a>, a principal software engineer at the Denver-based <a href="https://www.metarouter.io/" rel="noopener noreferrer" target="_blank">MetaRouter</a>, puts LLMs to work at both his day job and for personal side projects. (Nix also operates <a href="https://aiengineerjobs.com/" rel="noopener noreferrer" target="_blank">a jobs board of AI engineers</a>.) Nix said GLM 5.2 comes “really close” to frontier models like Anthropic’s Opus and OpenAI’s GPT-5.5—close enough to earn a permanent spot in his rotation.</p><p>“It’s pretty great at front-end development, for example, where I don’t need to always go to Opus or Fable for those things,” said Nix. He estimates that GLM 5.2 handles 10 to 20 percent of the work he sends to an LLM on a given day, and it’s now his first stop for some specific tasks, such as front-end design. </p><p>Others reported the same strength. Hasan said GLM 5.2 has “really good taste” in web design. <a href="https://www.linkedin.com/in/kacpermichalik/" rel="noopener noreferrer" target="_blank">Kacper Michalik</a>, a software engineer at Kraków, Poland–based <a href="https://screen.studio/" rel="noopener noreferrer" target="_blank">Screen Studio</a>, received good results while using GLM 5.2 to create forms for use on a website.</p><p>On the other hand, <a href="https://www.linkedin.com/in/sai-kiran-myadaram-027893242/" rel="noopener noreferrer" target="_blank">Sai Kiran Myadaram</a>, a software engineer at Bengaluru, India–based <a href="https://www.linkedin.com/company/indhic-ai/" rel="noopener noreferrer" target="_blank">Indhic AI</a>, reports less positive results with Z.ai. He signed up for Z.ai’s subscription plan the week GLM 5.2 launched and found the model burned through its token allotment quickly. “The weekly quota that Z.ai provides has been exhausted for me in less than two to three days,” he said. Michalik, who also accessed Z.ai directly, had no significant issues with the model’s quality but occasionally bumped into rate limits, though in his case he stuck to the free plan.</p><p>In addition to rate limits, Myadaram experienced problems with model hallucinations and overplanning when asked to tackle minor front-end fixes. “It’s messing up my code base,” said Myadaram. He’s since drifted back to OpenAI’s <a href="https://openai.com/codex/" rel="noopener noreferrer" target="_blank">Codex</a>. </p>]]></description><pubDate>Tue, 21 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-coding-assistant-china-anthropic</guid><category>Large-language-models</category><category>China</category><category>Ai-research</category><category>Ai-benchmarks</category><category>Anthropic</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-smartphone-running-a-chinese-ai-application-called-z-ai.jpg?id=67508250&amp;width=980"></media:content></item><item><title>How to Make an Invisible Drone</title><link>https://spectrum.ieee.org/invisible-spinning-drone</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/low-visibility-drone-flying-in-front-of-an-office-plant.jpg?id=67480624&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p><span>There are many words that I would never, ever use to describe a drone. Stealthy. Subtle. Whatever the opposite of obnoxious is. Much of this is because of the giant angry bee sound that drones tend to make, but it’s also the way that they look in flight: With uncannily linear movements and an even less canny ability to hover perfectly still, they tend to draw the eye as affronts to nature.</span></p><p>In a paper presented this week at <a href="https://roboticsconference.org/" target="_blank">Robotics Science and Systems 2026</a> in Sydney, roboticists from Northwestern University, Evanston, Ill., demonstrated a drone called Phantom Twist that is essentially invisible to humans, being an order of magnitude more difficult to see in flight than a typical quadrotor. They accomplished this with the aid of computational design, and while the resulting hardware is, I would argue, also an order of magnitude more of an affront to nature than a typical quadrotor represents, it’s pretty amazing how well it works.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="72b9c290bcfb7c1ef7a4e51a32cb7399" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/5KQ7dKs1dpQ?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">Phantom Twist spins so fast, it’s practically invisible.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p><span>The trick here is easy to see, even if the drone isn’t. By spinning in flight at between 15 and 25 hertz, Phantom Twist takes advantage of humans’ decidedly mediocre visual system to turn a solid spinning object into an opaque smear. Human eyes take some amount of time (typically about 100 milliseconds) to integrate what we see before sending the full scene off to our brains for processing. Moving objects can cause problems for this system, because if the movement is fast enough, our eyes are forced to average that motion across the scene, combining it with whatever is in the background and resulting in a transparent blur. This effect is called persistence of vision. For something that spins like Phantom Twist, that motion blur comes from the drone’s rapid rotation, and it works because most of the drone is cleverly designed to be empty space.</span></p><p>Drones that spin in flight are nothing new—we’ve covered a bunch of them in the past, including <a href="https://www.modlabupenn.org/piccolissimo/" target="_blank">Picolissimo</a> and any number of <a href="https://spectrum.ieee.org/spinning-drone" target="_self">samara</a> <a href="https://spectrum.ieee.org/foldable-monocopter-drone" target="_self">drones</a> inspired by the spinning flight of maple seeds. What makes Phantom Twist unique, and also very odd, is that the design was computationally optimized for low visibility. </p><h2>Controlling how drones like this fly</h2><p>Before we get into that, though, a quick note about how drones like this can even fly controllably, because it’s not at all obvious. With just a single motor and no control surfaces, the only possible control input is through the motor itself, and by pulsing the motor speed up or down at just the right time during each rotation, the drone can translate in any direction. Altitude control comes from changing overall motor thrust, and the drone‘s spinning nature makes it passively stable.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="A minimalist drone made of a few thin rods, wires and a miniature circuit board." class="rm-shortcode" data-rm-shortcode-id="0b6c02f1356af36a7437e22c62bbcbf9" data-rm-shortcode-name="rebelmouse-image" id="b2caf" loading="lazy" src="https://spectrum.ieee.org/media-library/a-minimalist-drone-made-of-a-few-thin-rods-wires-and-a-miniature-circuit-board.jpg?id=67480639&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Carbon fiber rods connect batteries, a controller, some counterweights, and a motor and propeller. The research robot also includes optical tracking tags.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p>The bits that you need for this kind of drone include the motor and propeller, a couple of batteries, a controller, some counterweights (which could be replaced with more batteries or payload), 0.8-mm carbon fiber rods to tie it all together, and a connector for the handheld launcher that gets the whole thing up to speed. The actual <em><em>arrangement</em></em> of these components is surprisingly flexible, and that’s where the invisibility comes in. </p><p>“The design space is high dimensional,” explains Northwestern’s <a href="https://www.mccormick.northwestern.edu/research-faculty/directory/profiles/rubenstein-michael.html" target="_blank">Michael Rubenstein</a>. “It’s very difficult for a human to reason through all the trade-offs between the physical constraints required for stable flight and the visual appearance of the spinning drone, and I don’t think we would have easily arrived at this low-visibility design ourselves.”</p><p>The visibility (or not) of Phantom Twist is primarily driven by the extent to which different components line up with each other from the perspective of someone looking at the drone. The more components that line up with each other as the drone flies, the less background you see through the spinning drone, and the more visible the drone becomes. Because you might be looking at the drone from a number of different angles, and also because the drone has to be stable enough for controlled flight, there are a bunch of different things that need to be optimized all at once, which is why computational design is effective here.</p><p>Phantom Twist’s final design was generated using an iterative optimizer which had a goal of minimizing a metric called <a href="https://eureka.patsnap.com/article/what-is-lpips-and-how-it-measures-perceptual-similarity" target="_blank">learned perceptual image patch similarity</a>, or LPIPS, while making sure that the design could still physically work. LPIPS is the difference between two images: a background image, and a background image with an overlay of the simulated spinning drone. The smaller that difference is, the more invisible that design is. It’s tricky for a human to consider all of the variables at once, but Rubenstein says that the final design does make intuitive sense, because “the automated pipeline prefers placements where components don’t visually overlap as it spins, or where the components are too close to the center of rotation.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Two variations of minimalist drones made from a few thin rods, wires and a miniature circuit board. Both are barely visible when in-flight." class="rm-shortcode" data-rm-shortcode-id="1869e49edd1b81064d2bfb41e0eedecb" data-rm-shortcode-name="rebelmouse-image" id="394e6" loading="lazy" src="https://spectrum.ieee.org/media-library/two-variations-of-minimalist-drones-made-from-a-few-thin-rods-wires-and-a-miniature-circuit-board-both-are-barely-visible-when.jpg?id=67480677&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Two iterations of Phantom Twist drones are shown with their handheld launching mechanisms. The better-optimized version [bottom row] relocates the launcher interface to remove components that are too close to the central axis, making them more visible.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p>Out of a starting set of around 20,000 feasible Phantom Twist configurations, the optimized design (the one that you see or don’t see in the pictures and videos) has a LPIPS score of 0.0104. A human-designed Phantom Twist is about twice as visible, with a LPIPS score of around 0.2, and a conventional quadrotor (of the same size) would be over 10 times more visible. And there’s still a bit more optimization that could be done with the electrical wiring as well as increasing the baseline transparency of the components themselves.</p><p>Phantom Twist is currently controlled using an optical tracking system, which means that it’s not yet capable of flying outside of a controlled environment. But Rubenstein has built <a href="https://www.roboticsproceedings.org/rss20/p102.pdf" target="_blank">other drones along similar principles</a> in the past, which have successfully flown outside, and he’s optimistic about using those techniques to break Phantom Twist out of the lab. The spinning behavior might even enable some useful sensing capabilities, he says. “An interesting possibility is mounting a camera on the spinning body. As the vehicle rotates, it could capture imagery in every direction, effectively creating a 360-degree view of its surroundings that could be used for onboard navigation and control.”</p><p>As for what a drone like Phantom Twist could be used for—assuming that the sound can be mitigated somewhat (and there are <a href="https://spectrum.ieee.org/pennysized-ionocraft-flies-with-no-moving-parts" target="_self">potential approaches to making that happen</a>), a stealthy microdrone could do all sorts of things with covert surveillance being the most obvious application. For his part, Rubenstein says that he’s personally excited about the potential for watching wildlife, “where a less-intrusive drone could observe animals while minimizing its impact on their natural behavior.” The elephants in particular <a href="https://spectrum.ieee.org/research-proves-drones-sound-like-bees-which-is-good-news-for-elephants" target="_self">would certainly appreciate that</a>.</p><p>For a deeper dive into all the particulars of this project, read the paper: <a href="https://arxiv.org/html/2605.11296v1" target="_blank"><em><em>Computational Design of a Low-Visibility UAV Using a Human-Aligned Perceptual Metric</em></em></a>, by Jingxian Wang, Chen Yu, David Matthews, Emma Alexander, Sam Kriegman, and Michael Rubenstein from Northwestern University, which is being presented this week at <a href="https://roboticsconference.org/" target="_blank">RSS 2026 in Sydney</a>.</p>]]></description><pubDate>Thu, 16 Jul 2026 16:09:21 +0000</pubDate><guid>https://spectrum.ieee.org/invisible-spinning-drone</guid><category>Robotics</category><category>Drones</category><category>Spinning-drones</category><category>Invisible-drones</category><category>Ai-design-optimization</category><category>Uav</category><dc:creator>Evan Ackerman</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/low-visibility-drone-flying-in-front-of-an-office-plant.jpg?id=67480624&amp;width=980"></media:content></item><item><title>Digital Surveillance Reshapes Fishery Enforcement in Indonesia</title><link>https://spectrum.ieee.org/fishery-satellite-surveillance</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/an-overhead-drone-photo-shows-an-array-of-closely-placed-fishing-boats-at-sea.jpg?id=67101579&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>In the eastern Indian Ocean, south of Java in the vast sea stretching toward Australia, a fishing vessel slightly alters its course while operating near the boundary of its authorized fishing ground. Nothing appears unusual on deck. Nets remain in the water. Engines maintain a steady speed. To the crew, it is an ordinary day at sea.</p><p>Yet hundreds of kilometers above, satellites continuously record the vessel’s position. At Indonesia’s Marine and Fisheries Resources Surveillance Station, in Cilacap, where I work, a monitoring platform receives the signal and automatically compares it against fishing permits, designated fishing grounds, vessel characteristics, and historical movement patterns. Within minutes, the system identifies a potential violation. Before any patrol vessel leaves port, before any inspector boards a vessel, and before any warning is issued, we have begun enforcement.</p><p>This transformation reflects a profound shift in maritime governance. The ocean has historically been opaque to regulators. States could only enforce laws where patrol vessels happened to be present. Today, however, integrated systems combining data from vessel monitoring systems (VMS), <a href="https://spectrum.ieee.org/earth-observation-satellites-small-constellations" target="_blank">satellite remote sensing</a>, geospatial analytics, and increasingly sophisticated data-processing tools are making marine activity visible at an <a href="https://www.science.org/doi/10.1126/science.aao5646" rel="noopener noreferrer" target="_blank">unprecedented scale</a>. Global Fishing Watch alone tracks <a href="https://globalfishingwatch.org/research/annual-report-2025/" rel="noopener noreferrer" target="_blank">hundreds of thousands of vessels</a> worldwide, generating a near real-time picture of fishing activity across the world’s oceans.</p><p>Indonesia has emerged as one of the most ambitious examples of this transition. As the world’s largest archipelagic state, managing more than 6 million square kilometers of maritime space, Indonesia faces a challenge familiar to many coastal nations: There are never enough patrol vessels. Digital surveillance is a practical necessity that makes my job possible, even as it creates new challenges.</p><h2>The Law of the Sea Meets Digital Reality</h2><p>The international legal framework governing the oceans was designed in an era when maritime enforcement depended almost entirely on physical presence. The <a href="https://www.un.org/depts/los/convention_agreements/texts/unclos/unclos_e.pdf" rel="noopener noreferrer" target="_blank">United Nations Convention on the Law of the Sea (UNCLOS), adopted in 1982,</a> assumes that states exercise authority through patrols, inspections, vessel boardings, and direct observation.</p><p>For countries with extensive coastlines and limited enforcement resources, this model has always faced practical constraints. Indonesia’s Fisheries Management Areas (WPP-NRI) span waters ranging from the Indian Ocean to the Pacific and from the Strait of Malacca to the maritime boundaries adjacent to Australia and Papua New Guinea. Monitoring such a vast domain solely through patrol operations is both expensive and operationally impossible.</p><p>Beginning in the late 2010s, Indonesia accelerated the integration of satellite-based monitoring into fisheries enforcement. Vessel monitoring systems became a cornerstone of this strategy. By early 2026, a total of <a href="https://ppid.kkp.go.id/upt/pelabuhan-perikanan-samudera-kendari/news/detail/9394-kapal-perikanan-sudah-pasang-vms/" rel="noopener noreferrer" target="_blank">9,394 Indonesian fishing vessels</a> were actively transmitting through the national VMS, representing an increase of 2,880 vessels during the 2021–2025 period. As part of Indonesia’s broader maritime surveillance architecture, VMS data are complemented by satellite remote sensing and other monitoring tools to help identify suspicious activities involving vessels operating without active transponders or outside the national VMS network.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A man in an Indonesian military uniform points to a large monitor showing an overview of thousands of fishing boats at sea near Indonesia." class="rm-shortcode" data-rm-shortcode-id="b9c4ba7e404292b97184b48952a729cc" data-rm-shortcode-name="rebelmouse-image" id="1d12e" loading="lazy" src="https://spectrum.ieee.org/media-library/a-man-in-an-indonesian-military-uniform-points-to-a-large-monitor-showing-an-overview-of-thousands-of-fishing-boats-at-sea-near.jpg?id=67101591&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Indonesian fisheries officials plan fishery patrols using data from tracking devices, satellites, and their understanding of the patterns of illegal fishing.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Indonesian Ministry of Marine Affairs and Fisheries</small></p><p>The implications extend far beyond vessel tracking. Continuous digital monitoring enables authorities to reconstruct vessel movements, identify suspicious behavioral patterns, detect unauthorized fishing activity, and verify compliance with licensing conditions. Rather than waiting to discover violations during patrol operations, regulators can increasingly prioritize inspections based on data-derived risk assessments.</p><p>Maritime governance is shifting from reactive enforcement toward predictive oversight.</p><h2>The Surprising Geography of Digital Enforcement</h2><p>The expansion of surveillance infrastructure has already generated measurable enforcement outcomes.</p><p>The Ministry of Marine and Fisheries Affairs Indonesia <a href="https://kkp.go.id/unit-kerja/djpsdkp/akuntabilitas-kinerja/pelaporan-kinerja/detail/laporan-kinerja-direktorat-jenderal-psdkp-tahun-20256a042458a936b.html" target="_blank">imposed 2,550 administrative sanctions during 2025</a>, many involving violations detected through the vessel monitoring system, including fishing outside authorized fishing grounds and deliberate deactivation of monitoring transmitters.</p><p>This statistic is significant because many of these violations would have been extremely difficult to detect under traditional patrol-based enforcement. A vessel that briefly crosses into a prohibited fishing zone may never encounter an enforcement vessel. Likewise, a captain who temporarily disables a transmitter may escape detection if oversight depends solely on physical inspections.</p><p>Digital monitoring fundamentally changes this equation. Every vessel movement creates a data trail. Authorities can reconstruct routes, identify anomalous behavior, and compare activities against permit conditions long after the event itself has occurred.</p><p>The first quarter of 2026 demonstrates the scale of this surveillance capability. During just three months, Indonesia’s fisheries monitoring system tracked 14,571 fishing vessels, 182 fishing gear units, and 208 registered home ports while identifying <a href="https://kkp.go.id/download-pdf-akuntabilitas-kinerja/akuntabilitas-kinerja-pelaporan-kinerja-laporan-kinerja-direktorat-jenderal-psdkp-twi-tahun-2026.pdf" target="_blank">491 suspected violations</a> across the country’s fisheries management areas. These violations included unauthorized fishing grounds, illegal high-seas operations, transshipment-related offenses, port-base discrepancies, licensing irregularities, and indications of poaching.</p><p>Such numbers reveal a fundamental transformation. Enforcement is no longer limited by the number of patrol vessels available at sea. Instead, surveillance capacity increasingly depends on the ability to collect, process, and interpret big data.</p><h2>Illegal Operators Are Learning Too</h2><p>Yet greater visibility does not eliminate illegal fishing. But it does change how poachers operate.</p><p>Indonesia’s expanding digital surveillance network, and a 2023 requirement that even small vessels use VMS when 12 nautical miles offshore, appears to have improved compliance among licensed fishing vessels. However, as enforcement capabilities become more sophisticated, some actors engaged in illegal fishing have also become more adept at exploiting technological and operational gaps.</p><p>Deliberately disabling VMS transmitters remains one of the most common enforcement concerns. While temporary signal losses, whether intentional or caused by technical failures—can complicate the reconstruction of vessel movements, they do not necessarily prevent authorities from detecting potentially illegal activity. Indonesia increasingly combines VMS with satellite-based observations, other maritime surveillance systems, intelligence-led analysis, and reports from <a href="https://repository.seafdec.org/bitstream/handle/20.500.12066/7649/Fish-surveillance-Indonesia.pdf?sequence=1&isAllowed=y" target="_blank">community-based surveillance groups</a> (<em>Pokmaswas</em>) to corroborate suspicious behavior and direct patrol resources where they are most needed. This layered approach—integrating digital technologies with local knowledge from coastal communities—helps reduce opportunities for illegal, unreported, and unregulated (IUU) fishing even when a single monitoring system is compromised.</p><p class="pull-quote">A compromised surveillance network could potentially disrupt enforcement operations just as effectively as a vessel evading patrol detection.</p><p>As digital surveillance expands, one lesson from Indonesia’s experience is that stronger monitoring does not eliminate illegal fishing—it changes how illegal operators behave. Improved compliance across much of the fishing fleet has been accompanied by increasingly sophisticated attempts by a smaller group of offenders to avoid detection. This reflects a broader reality of technology-enabled enforcement: As monitoring capabilities evolve, so do the strategies used to circumvent them.</p><p>The result is a technological arms race. Every improvement in surveillance capability encourages <a href="https://www.science.org/doi/10.1126/science.aad5686" target="_blank">new methods of avoidance</a>, whether through disabling tracking devices, manipulating vessel identities, or exploiting gaps between different monitoring systems. Enforcement agencies must therefore continuously refine their analytical methods, integrate multiple sources of maritime information, and adapt their operational strategies to keep pace with evolving behavior at sea. Effective digital fisheries governance is not defined by a single technology but by the ability to combine data, human expertise, and operational intelligence into a resilient and adaptive enforcement system.</p><h2>The Next Battle May Be Over Data Integrity</h2><p>The future of fisheries enforcement may ultimately depend less on detecting vessels and more on ensuring confidence in the digital systems that generate enforcement decisions.</p><p>As surveillance networks become increasingly integrated, questions surrounding cybersecurity, algorithmic accountability, and data integrity become more important. What happens if vessel tracking data are manipulated? How should authorities verify automated risk assessments? What safeguards exist when enforcement actions increasingly originate from algorithmic analysis rather than direct human observation?</p><p>These questions are no longer theoretical.</p><p>Modern fisheries governance increasingly depends on interconnected networks of satellites, communication systems, databases, cloud infrastructure, and analytical platforms. While these technologies dramatically improve visibility, they also create new vulnerabilities. A compromised surveillance network could potentially disrupt enforcement operations just as effectively as a vessel evading patrol detection.</p><p>For Indonesia, this means that investment in digital surveillance must be accompanied by investment in digital resilience. The effectiveness of a monitoring system ultimately depends not only on the volume of data collected but also on the credibility, security, and <a href="https://spectrum.ieee.org/data-integrity" target="_blank">reliability of the information produced</a>.</p><h2>Governing Oceans Through Data</h2><p>Indonesia’s experience illustrates a broader global transformation in maritime governance. The ocean is becoming increasingly transparent to regulators. Activities that once occurred beyond the reach of enforcement agencies can now be observed, analyzed, and investigated through interconnected digital systems.</p><p>The benefits are substantial. Expanded VMS adoption, improved monitoring coverage, and thousands of administrative enforcement actions demonstrate that digital surveillance can significantly enhance fisheries governance. Yet the transition also introduces new challenges involving data quality, cybersecurity, algorithmic accountability, and adaptive <a href="https://doi.org/10.3389/fmars.2018.00240" rel="noopener noreferrer" target="_blank">criminal behavior</a>.</p><p>The central question facing maritime regulators is how governments can ensure that increasingly powerful monitoring systems remain transparent, secure, and accountable while preserving public trust and legal legitimacy. The most important lesson may be that digital surveillance does not replace traditional enforcement. It changes where enforcement begins. For generations, maritime law enforcement started when a patrol vessel encountered a suspected violator. Today, it often starts when an algorithm detects a pattern.</p><p>That shift may prove as significant for ocean governance as the invention of radar was for maritime navigation.</p>]]></description><pubDate>Thu, 16 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/fishery-satellite-surveillance</guid><category>Fishing</category><category>Environmental-monitoring</category><category>Indonesia</category><category>Poaching</category><category>Surveillance</category><dc:creator>Yogi Putranto</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/an-overhead-drone-photo-shows-an-array-of-closely-placed-fishing-boats-at-sea.jpg?id=67101579&amp;width=980"></media:content></item><item><title>The First Chatbot’s Multiple Personalities</title><link>https://spectrum.ieee.org/eliza-chatbot-source-code</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/photo-collage-of-a-charismatic-elderly-man-raising-his-eyebrows-against-a-background-of-obscured-computer-programming-script.jpg?id=67155632&width=1245&height=700&coordinates=0%2C0%2C0%2C0"/><br/><br/><p class="intro-text"><em>ELIZA is remembered as the world’s first AI star, a kindly therapist in chatbot form that gently probed users’ worries. Even its creator, Joseph Weizenbaum, was surprised by the warm reception given to his experiment in human-machine interaction. For some, it heralded an age of automated psychotherapy, while others believed the program demonstrated sentience, a fallacy soon known as the “<a href="https://spectrum.ieee.org/why-people-demanded-privacy-to-confide-in-the-worlds-first-chatbot" target="_self">ELIZA effect</a>.” Based on published descriptions, ELIZA has been implemented on many different computers, but only recently has <a href="https://sites.google.com/view/elizaarchaeology/code" rel="noopener noreferrer" target="_blank">the actual source code</a> been unearthed from <a href="https://archivesspace.mit.edu/" rel="noopener noreferrer" target="_blank">MIT’s archives</a>. </em></p><p><em>In <a href="https://mitpress.mit.edu/9780262052481/inventing-eliza/" target="_blank">Inventing ELIZA: How the First Chatbot Shaped the Future of AI</a>, just published by <a href="https://mitpress.mit.edu/" rel="noopener noreferrer" target="_blank">MIT Press</a>, a squad of researchers analyze the code and reveal a complex program capable of much more than faking psychiatry. In fact, it could assume several different personas. The authors have also created <a href="https://sites.google.com/view/elizaarchaeology/try-eliza" rel="noopener noreferrer" target="_blank">a faithful emulation of the therapist persona that you can try yourself</a></em><span><em> after reading the book excerpt below.</em></span></p><p class="drop-caps">W<strong>hen it debuted in</strong> the mid-1960s, the ELIZA software program transformed the way people thought about interacting with computers. As the first chatbot, ELIZA demonstrated how a calculation machine might engage in conversation, ushering in a host of social and technical questions that still resonate today. Now we don’t think twice about interacting with a machine in real time, conversing over text, or even speaking into the air to ask about the weather. In many ways, ELIZA shaped not only the way we think about <em><em>interacting</em></em> with computers but also how we think <em><em>about</em></em> them. It began to give a reality to the <a href="https://spectrum.ieee.org/tag/science-fiction" target="_blank">science fiction</a> stories of how we expect computers to work.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Orange book cover titled \u201cInventing Eliza: How the First Chatbot Shaped the Future of AI\u201d" class="rm-shortcode" data-rm-shortcode-id="208d351be5eb0018c13ad5a993fa22b5" data-rm-shortcode-name="rebelmouse-image" id="03894" loading="lazy" src="https://spectrum.ieee.org/media-library/orange-book-cover-titled-u201cinventing-eliza-how-the-first-chatbot-shaped-the-future-of-ai-u201d.jpg?id=67155786&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">This article is adapted from the new book “Inventing ELIZA: <a href="https://mitpress.mit.edu/9780262052481/inventing-eliza/" target="_blank">How the First Chatbot Shaped the Future of AI</a>“ (MIT Press, 2026).</small></p><p>Although ELIZA was far from a faultless conversation partner, it astonished its users. The recent discovery and archaeology of the original ELIZA source code represents a significant intervention in the history of computing. By examining the actual implementation of ELIZA rather than relying on later reconstructions and reimplementations, we challenge taken-for-granted assumptions about this key software artifact.</p><p>For example, the source code reveals that ELIZA was not merely a simple pattern-matching chatbot but can be better understood as a sophisticated platform designed for multiple “personas,” or scripts, with a complex set of capabilities, including script editing and contextual memory. The script that most people conflate with the program ELIZA was actually called Doctor, which performed the role of a psychotherapist. Yet, like a modern chatbot prompted to behave with different personalities, ELIZA could take on many roles.</p><p class="pull-quote">“This code and script…reveal underlying assumptions about language, therapy, and human-computer interaction that continue to influence modern AI development.”</p><p>This unearthed material transforms our understanding of early AI development by demonstrating that Joseph Weizenbaum’s technical innovations were far more advanced than previously documented. Moreover, the discrepancies between his published descriptions and the actual implementation help to show the gap between theoretical computational models and their material instantiations in computer source code, a tension that continues to shape digital culture today.</p><p>Although many technical innovations have emerged in the decades since ELIZA, examining the ELIZA/Doctor code offers a rare glimpse into one of the earliest formalized attempts to model human conversation. What makes ELIZA particularly fascinating is not only its historical significance but also what it reveals about Weizenbaum’s views on both computing and human interaction. This code and script do not merely showcase programming techniques of the 1960s; they reveal underlying assumptions about language, therapy, and human-computer interaction that continue to influence modern AI development. By examining this code, we can start to uncover the sophisticated linguistic and programming techniques that allowed a rudimentary pattern-matching system to create a convincing simulation of understanding. But before we can read the lines of code, let us offer an overview of the system.</p><h2>How Did ELIZA Create Personas?</h2><p>The architectural distinction between ELIZA and Doctor represents an important design decision in AI history. Think of ELIZA as a system for interaction and Doctor as one set of rules that Weizenbaum devised, among others. This separation, manifested in ELIZA’s system-script dichotomy, presaged numerous contemporary software patterns, from configuration-as-data to plug-in architectures and domain-specific languages.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A 1960s chatbot program running on a 1980s IBM personal computer." class="rm-shortcode" data-rm-shortcode-id="25a3713ea68f943c1be636d8e3cdbc0d" data-rm-shortcode-name="rebelmouse-image" id="1b809" loading="lazy" src="https://spectrum.ieee.org/media-library/a-1960s-chatbot-program-running-on-a-1980s-ibm-personal-computer.jpg?id=67155917&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Based on published journal articles, ELIZA was re-created on many platforms, such as the IBM PC. However, the actual source code sat untouched in the MIT archives for many years. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VCF Museum at InfoAge</small></p><p>Without question, the historical context of 1960s computing fundamentally shaped ELIZA’s architecture as well. Decisions in computing that reflect material constraints create path dependencies and eventually become programming cultural norms. These constraints manifested in ELIZA’s single-pass processing, tape-based storage and stack-oriented implementation. Yet within these limitations, Weizenbaum crafted an elegant solution. These technical features, though invisible to the users, are crucial to creating the illusion of understanding that made ELIZA so compelling.</p><p>Weizenbaum explained many of ELIZA’s technical features in the 10-page paper published in the <a href="https://dl.acm.org/doi/10.1145/365153.365168" target="_blank">January 1966 edition of the journal <em><em>Communications of the Association for Computing Machinery</em></em></a> (<em><em>CACM</em></em>). But he chose to omit some essential details.</p><p>In that paper Weizenbaum published ELIZA’s best known dialogue, which begins,</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>Men are all alike.</p><p>IN WHAT WAY</p><p>They’re always bugging us about something or other.</p><p>CAN YOU THINK OF A SPECIFIC EXAMPLE</p><p>Well, my boyfriend made me come here.</p></div><p>This dialogue marked ELIZA’s public debut in 1966 as one of the examples produced by the Doctor script. By finding the source code for ELIZA and examining how it performs the Doctor script, we now better understand these two separate parts of a system and can explore the many other personas of ELIZA. In just some of the other scripts known to date, ELIZA was programmed to discuss math, <a href="https://spectrum.ieee.org/tag/poetry" target="_blank">poetry</a>, color, paradoxes, synchronization, relativity, France, and elevators.</p><p>These scripts work like templates. They are structured data that direct the ELIZA system to “play” a particular task or role. By comparing archival and published ELIZA dialogues from interactions with a variety of scripts, including Doctor, we can understand more about bot personas and how they function, paying close attention to how a bot evokes social dynamics between system and interactor.</p><p>Ultimately, studying the dialogues and scripts demonstrates the crucial role that collaboration plays in these exchanges, as bot and user cocreate the sense of their interaction. To understand the full range of ELIZA’s capabilities and conversational possibilities, let’s take a look at the variety of scripts that were created for the ELIZA system.</p><p>What distinguishes each ELIZA script is both its subject matter and the linguistic and stylistic choices used to deliver that content. These choices are not neutral; they can be said to construct a particular persona with characteristics that emerge through the script’s language patterns, vocabulary, and conversational approach. In short, it matters not just what you say but how you say it too.</p><p class="pull-quote">“The aim was less to create a functional automated therapist and more to find a suitably constrained role to match the limitations of the programming environment.”</p><p>For example, with the Doctor script Weizenbaum deliberately echoed the style of a Rogerian “talk” therapist. He chose this persona because the psychiatric mode is one of the few types of conversations in which one person can “assume the pose of knowing almost nothing of the real world. If, for example, one were to tell a psychiatrist ‘I went for a long boat ride’ and he responded, ‘Tell me about boats,’ one would not assume that he knew nothing about boats but that he had some purpose in so directing the subsequent conversation.”</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of paper loaded in a teletype machine with a few paragraphs of chatbot dialogue written on it. " class="rm-shortcode" data-rm-shortcode-id="86bbdc58c1e01df11061af97b1e9b45c" data-rm-shortcode-name="rebelmouse-image" id="97e5e" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-paper-loaded-in-a-teletype-machine-with-a-few-paragraphs-of-chatbot-dialogue-written-on-it.jpg?id=67161217&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The first users of ELIZA interacted with it via teletype terminals.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VCF Museum at InfoAge</small></p><p>Thus, the most famous persona created for ELIZA was a technical convenience. As human-computer interaction expert Lucy Suchman explains, “The Doctor program exploited the maxim that shared premises can remain unspoken: that the less we say in conversation, the more what is said is assumed to be self-evident.” In creating the original ELIZA effect, less was more.</p><p>The aim was less to create a functional automated therapist and more to find a suitably constrained role to match the limitations of the programming environment. Then Weizenbaum composed the script to match the role by choosing specific words that evoked rhetorical tone and characterization, for example, <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">LET’S DISCUSS FURTHER WHY YOU … WHAT DOES THAT SUGGEST TO YOU.</span> In Doctor, the machine side of the conversation needs to appear like a good listener who cares about what the user has mentioned before, so it often includes the user’s text in its replies and keeps its responses open-ended. Because a real doctor would be inquisitive, the script contains lots of<span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">WHAT</span> and<span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">WHY</span> questions. In other scripts and dialogues, the script keywords and assigned responses reveal the design decisions that help create different specific personas. This variation becomes increasingly apparent as we look at the wider range of ELIZA scripts.</p><h2>What Were Some of ELIZA’s Alternate Personas?</h2><p>Let’s start with a script called Neweng, short for New England. By the mid-1950s, scientists were already taking up Turing’s proposition for conversing with a computer as a tangible idea. In 1959, a “conversation machine” made by a small team at the University of Toronto Computation Centre could make small talk about the weather if preprompted with current data. Its script included a ranked series of time words, weather words, and emotion words, but the system had “a rather narrow outlook,” limited vocabulary, and no knowledge of syntax. According to its creators, “in spite of this crudity, the results in passable conversation are surprisingly good, which in a way is a commentary on the shallowness of ordinary conversation about the weather.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Typewritten list of inquisitive dialogue responses such as \u201cWhat does that suggest to you?\u201d" class="rm-shortcode" data-rm-shortcode-id="5fcf41cfad8ee21b240d837e7b66c660" data-rm-shortcode-name="rebelmouse-image" id="17794" loading="lazy" src="https://spectrum.ieee.org/media-library/typewritten-list-of-inquisitive-dialogue-responses-such-as-u201cwhat-does-that-suggest-to-you-u201d.jpg?id=67156058&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">This fragment of ELIZA’s source code shows how scripted responses were encoded, in this case to make ELIZA sound like a Rogerian therapist. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MIT Libraries, Distinctive Collections</small></p><p>One of ELIZA’s test scripts found in the archive had a similar function. Neweng does not soothe, therapize, or instruct, but only makes small talk. It chats as if at a boring dinner party.</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>WHAT NEW ENGLAND STATE DO YOU LIKE BEST OF ALL</p><p>i love rhode island</p><p>I ENJOY THE BEACHES ON LONG ISLAND SOUND</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like massachusetts</p><p>THAT IS MY HOME STATE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>you must like massachusetts too</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i enjoy skiing in new hampshire</p><p>PLEASANTLY COOL IN THE SUMMER UP THERE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like missouri</p><p>MISSOURI IS NOT A NEW ENGLAND STATE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like cape cod</p><p>THAT’S A BEAUTIFUL PART OF MASSACHUSETTS</p></div><p>What can we learn from such a basic conversation in which it seems like nothing interesting is shared? Despite the bland chat, a persona emerges from the bot script in phrases like <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">PLEASANTLY COOL</span> and <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">UP THERE</span> and in locations like <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">THE BEACHES ON LONG ISLAND SOUND</span>. From such comments, we can begin to imagine a persona who has not only a certain vocabulary but also certain hobbies and vacation spots. Perhaps implicit assumptions of race and class emerge as well.</p><p>A chatty persona might take on many forms, depending on where and when the conversation occurs. This one seems reserved, perhaps fitting its setting in 1960s New England. The system reminds the user that Missouri is not a New England state, but what if this conversation took place in Missouri, Texas, or Mexico? The machine persona would sound different in its cadence, tone, and references. What would we come to understand about a chat persona from Fire Island, from Brooklyn, from Berlin? What would they sound like, and what topics would they discuss?</p><p>These differences in subject matter do matter. They imply personas with entirely different backgrounds and experience, giving users wholly different interactions and affective relations. In this way, the Neweng script demonstrates how even simple algorithms making contextual responses about geography could generate a convincing sense of personhood and place. Whereas Neweng could be said to have created a casual, conversational persona focused on light social exchange, other scripts pushed ELIZA into more structured and educational roles. These scripts demonstrate how the system could be adapted not just for friendly chatter but for teaching.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Black-and-white portrait of a middle-aged balding man in aviator eyeglasses." class="rm-shortcode" data-rm-shortcode-id="a02ade38f8966ac94c795b313e1c3a6b" data-rm-shortcode-name="rebelmouse-image" id="fc521" loading="lazy" src="https://spectrum.ieee.org/media-library/black-and-white-portrait-of-a-middle-aged-balding-man-in-aviator-eyeglasses.jpg?id=67155997&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">Edwin Taylor, at MIT’s Education Research Center, developed alternate scripts for ELIZA, testing its ability to act as a teacher.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MIT Libraries, Distinctive Collections</small></p><p>Meet ELIZA the tutor, quite unlike ELIZA the therapist or the chatty neighbor. Intrvw, Canvec, FVP1, and Arithm are a set of ELIZA scripts created as teaching tools used in experiments by Edwin F. Taylor at MIT’s Education Research Center. These scripts run on later versions of ELIZA that incorporated an important technical innovation called conditional keyword matching.</p><p>Unlike the original ELIZA, which simply looked for keywords and generated responses based on their presence, these updated versions could track what had been discussed previously and branch into different conversational paths based on specific user answers. This development allowed ELIZA to simulate a kind of Socratic method, where a tutor guides learning through carefully sequenced questions that respond to student answers rather than simply presenting information.</p><p>These scripts construct the tutor persona through many subtle linguistic gestures that create characterization and rhetorical tone. This tone differs from that of Doctor, which asks open-ended questions and comes across as gentle and nonscientific. In the tutoring scripts, large blocks of informative text from the bot tend to dominate the conversation, and the tone is often more dry and unemotional in these explanations. The dialogues indicate structured scripts that include guidance to lead the student through narrow, Socratic learning paths.</p><p>In particular, the teaching scripts feature praise and critique. The dialogues for Intrvw, Canvec, and FVP1 are peppered with <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">EXCELLENT, VERY GOOD, RIGHT YOU ARE,</span> and <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">CONGRATULATIONS.</span> These create the sense of a supportive instructor cheering the student on. Such politeness has been taken up in contemporary bots like ChatGPT, which has been shown to perform better when people are polite back to it.</p><p>ELIZA could become a tutor more effectively as the system grew in its capabilities, another valuable reminder that ELIZA was not one program but a family of programs. After the publication of the 1966 <em><em>CACM</em></em> article, Weizenbaum continued to develop the systems for interaction and understanding. As an experiment, Weizenbaum wrote the Arithm script less as a tutor and more so to “to illustrate the power of the evaluator to which ELIZA has access.” It uses a friendly, plain language interface to let users do simple programming. The script can do calculations, assign variables to values, and perform operations on them. Math problems can be described in sentence form:</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>The radius of a globe is 10.</p><p>A globe is a sphere. A sphere is an object.</p><p>What is the area of the globe.</p><p>IT’S 1256.635916</p></div><p>The updated 1967 version of the ELIZA system can accumulate facts and store additional information. In this later version of ELIZA, when the system does not recognize information, it asks follow-up questions to gain data. As Weizenbaum explains, “The present script is designed to reveal, as opposed to conceal, lack of understanding and misunderstanding. Notice, for example, that when the program is asked to compute the area of the ball, it doesn’t yet know that a ball is a sphere and that when the diameter of the ball needs to be computed the fact that a ball is an object has also not yet been established.” Unlike Doctor, which asks questions to keep the conversation going, Arithm is building its store of, if not knowledge, then data and logic statements.</p><p>Although the variety of scripts helps us to see how a range of personas could be constructed through script programming ELIZA, they represent only half of the conversational process. A script can establish a foundation for a persona, but that persona only emerges fully through interaction with users who engage with it, interpret it, and respond to it in ways that may confirm, challenge, or transform the script’s implicit character. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Wed, 15 Jul 2026 15:35:53 +0000</pubDate><guid>https://spectrum.ieee.org/eliza-chatbot-source-code</guid><category>Chatbots</category><category>Ai</category><category>History</category><category>Eliza</category><category>Joseph-weizenbaum</category><dc:creator>Sarah Ciston</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/photo-collage-of-a-charismatic-elderly-man-raising-his-eyebrows-against-a-background-of-obscured-computer-programming-script.jpg?id=67155632&amp;width=980"></media:content></item><item><title>This AI Folds DNA Into Mini Masterpieces</title><link>https://spectrum.ieee.org/ai-dna-origami</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/3d-renderings-of-unique-dna-structures-resembling-a-flower-cone-and-corkscrew-shaped-loop.jpg?id=67155159&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p><span>Shaped like dogs, stars, and the Mona Lisa, you could mistake these DNA structures for fun-shaped macaroni if they weren’t only nanometers wide. South Korean scientists made the constructions using a technique called </span><a href="https://spectrum.ieee.org/nanoscale-interconnects-come-to-selfassembling-dna-origami" target="_self">DNA origami</a>,<span> which can bend genetic material into any form. Designing DNA strands so they’ll fold into a specific shape typically requires tedious manual work, but the researchers behind the playful fabrications have developed a shortcut using generative AI.</span></p><p>The AI model, called Generative SNUPI (short for Structured Nucleic Acids Programming Interface, and, yes, inspired by the dog), was created by research teams at <a href="https://en.snu.ac.kr/" target="_blank">Seoul National University</a> (SNU) and <a href="https://www.hanyang.ac.kr/web/eng" target="_blank">Hanyang University</a>. The work behind it, which was accepted for publication in <a href="https://www.nature.com/articles/s41467-026-73578-z" target="_blank"><em><em>Nature Communications</em></em></a><em><em>, </em></em>shows the model can conjure DNA origami designs that work in the real world for user-requested shapes. For a design like the Mona Lisa, that doesn’t mean simply tracing an outline; the model considers the chemical rules of DNA to tell researchers how unpaired DNA strands should be sequenced so that molecular forces will cause them to self-contort into the required shape.</p><p>DNA origami techniques have been around for <a href="https://www.nature.com/articles/nature04586" target="_blank">two decades now</a>, with potential applications ranging from nanoscale robots to therapeutic structures that interact with cells. But these innovations have been slowed by how time-consuming and expensive the DNA structure design process can be.</p><p>“Traditionally, we need some expertise, background knowledge, and know-how to design the proper nanostructures that we intend to make,” says <a href="https://www.linkedin.com/in/kyounghwa-jeon-7511041a3/" rel="noopener noreferrer" target="_blank">Kyounghwa Jeon</a>, a Ph.D. candidate at SNU. The work requires humans running algorithms and tweaking results until the desired shape is achieved and structurally stable. With Generative SNUPI, she says, users could, in theory, go straight from drawing a target shape to physically assembling the DNA. </p><p><a href="https://www.linkedin.com/in/rebecca-taylor-ph-d-022b854b/" rel="noopener noreferrer" target="_blank">Rebecca Taylor</a>, a professor of mechanical engineering at <a href="https://www.cmu.edu/" rel="noopener noreferrer" target="_blank">Carnegie Mellon University</a> who was not involved in the research, says the new generative platform is exciting for researchers. “The entire field is sort of enabled and held back by its tools. When you make a new tool that enables a new tech, a new capability, that’s just such a big advance for the field.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Traced outline of a Pug, surrounded by several dozen microscopic DNA structures bearing its exact shape. " class="rm-shortcode" data-rm-shortcode-id="62a7ef7bfdbc368bdc89ba8655839517" data-rm-shortcode-name="rebelmouse-image" id="e7d51" loading="lazy" src="https://spectrum.ieee.org/media-library/traced-outline-of-a-pug-surrounded-by-several-dozen-microscopic-dna-structures-bearing-its-exact-shape.jpg?id=67155167&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Generative SNUPI designs DNA sequences that, when synthesized, fold into nanoscale replicas of user-requested shapes.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Source images: <a href="https://www.nature.com/articles/s41467-026-73578-z" target="_blank">Chien Truong-Quoc, Kyounghwa Jeon, et al.</a></small></p><h2>How AI can design DNA origami</h2><p>Designing DNA origami using Generative SNUPI begins with a target shape. That could be something with complex curvature, like the outline of a dog’s face, or a more simple geometric pattern. Next, the new tech comes into play: Generative SNUPI applies a diffusion model, which adds and refines noise to the input shape to create the desired output in DNA form. Diffusion models are how platforms like <a href="https://openai.com/index/dall-e-3/" target="_blank">DALL-E</a> and <a href="https://www.midjourney.com/explore?tab=top" target="_blank">Midjourney</a> create <a href="https://spectrum.ieee.org/thermodynamic-computing-for-ai" target="_self">AI-generated imagery</a>.</p><p>“What it looks like is one of those kids crafts, where you decorate something with glue and then put glitter all over it,” says Taylor. When the noise is removed—or the glitter is shaken off—the design is revealed. “They’re basically just saying ‘populate this guide that I have with the DNA,’ but they also know how DNA comes together. … That’s the thing that it’s really been trained on.”</p><p>The arts-and-crafts metaphors only continue once Generative SNUPI returns the DNA sequences that form the target shape. Scientists chemically synthesize short DNA strands called staples and used biological methods to produce a long strand called a scaffold. The staples pull the scaffold into shape in a way that Jeon says is “very similar to stapling paper.” The staple-scaffold relationship exploits DNA’s imperative to bond guanine to cytosine and adenine to thymine; the exact positions of each of these molecules are dictated by Generative SNUPI during the design process. </p><p>Researchers were able to produce a variety of DNA origami structures, but some did not hold their shape at first, notes <a href="https://www.linkedin.com/in/do-nyun-kim-4b4830118/" target="_blank">Do-Nyun Kim</a>, an assistant professor of mechanical engineering at SNU. “This occurred not because Generative SNUPI had an error, but because the drawn shape was, in fact, structurally unstable,” he says. In response, they added a step before actually designing the DNA sequence to predict the structural integrity of the input shape. </p><p>To expand Generative SNUPI’s capacity for real-world applications, Kim says that DNA origami designs will need to be less rigid than what the model is currently able to produce. The technology reaching its full potential could mean life-saving uses like drug delivery and immunotherapy, but these uses often require flexibility.</p><p>“Most molecular structures are dynamic and reconfigure in response to external stimuli to perform their designated functions,” he says. “So, we plan to extend the current work to the design of dynamically reconfigurable structures in future research.”</p><p><em>This article appears in the September 2026 print issue as “</em><em>AI Model Folds DNA Into </em><em><em>Mini Mona Lisas</em>.”</em></p>]]></description><pubDate>Wed, 15 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-dna-origami</guid><category>Biotechnology</category><category>Dna-origami</category><category>Dna</category><category>Generative-ai</category><dc:creator>Alex Music</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/3d-renderings-of-unique-dna-structures-resembling-a-flower-cone-and-corkscrew-shaped-loop.jpg?id=67155159&amp;width=980"></media:content></item></channel></rss>