<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>IEEE Spectrum</title><link>https://spectrum.ieee.org/</link><description>IEEE Spectrum</description><atom:link href="https://spectrum.ieee.org/feeds/topic/artificial-intelligence.rss" rel="self"></atom:link><language>en-us</language><lastBuildDate>Wed, 29 Jul 2026 16:56:07 -0000</lastBuildDate><image><url>https://spectrum.ieee.org/media-library/eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpbWFnZSI6Imh0dHBzOi8vYXNzZXRzLnJibC5tcy8yNjg4NDUyMC9vcmlnaW4ucG5nIiwiZXhwaXJlc19hdCI6MTgyNjE0MzQzOX0.N7fHdky-KEYicEarB5Y-YGrry7baoW61oxUszI23GV4/image.png?width=210</url><link>https://spectrum.ieee.org/</link><title>IEEE Spectrum</title></image><item><title>Siobahn Day Grady Wants Everyone to Be AI Literate</title><link>https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/smiling-woman-in-nccu-eagles-jersey-stands-with-arms-crossed-in-front-of-red-mural.png?id=67527627&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Artificial intelligence is reshaping the skills employers expect from new graduates. In response, universities are scrambling to launch new courses, research centers, and industry partnerships that prepare students for today’s workforce. But building a cutting-edge AI curriculum demands funding and access to industry networks, resources that remain unevenly distributed across higher education. </p><p>At North Carolina Central University, <a href="https://www.nccu.edu/employee/sday" rel="noopener noreferrer" target="_blank">Siobahn Day Grady</a> is trying to change that equation.</p><p>In January 2025, Grady, an associate professor in the <a href="https://www.nccu.edu/slis" rel="noopener noreferrer" target="_blank">NCCU School of Library and Information Sciences</a>, launched the first AI research institute at a historically Black college or university, or HBCU. The <a href="https://www.nccu.edu/research/institutes-and-special-facilities/iaier" rel="noopener noreferrer" target="_blank">Institute for Artificial Intelligence and Emerging Research</a> (IAIER) aims in part to help students and faculty across the university develop the skills needed to navigate a labor market increasingly transformed by AI.</p><p>“There used to be a time where people could say, ‘I don’t do tech,’ or ‘That’s not for me,’” Grady says. “But we’re in a stage now where you do need digital skills. Now it’s evolving into AI literacy.” </p><p>The approach reflects a broader shift in how many universities are thinking about AI education. AI skills are no longer confined to computer science and engineering departments—and at NCCU, they can’t be. The university does not yet have a dedicated computer science program, though it is developing one alongside a new AI minor. </p><p>The challenge of providing these resources is especially acute for historically Black institutions. Although HBCUs account for roughly 3 percent of four-year institutions in the United States, they receive less than 1 percent of federal research and development funding, according to a <a href="https://tmcf.org/new-data-reveals-disproportionately-low-rd-funds-awarded-to-hbcus/" rel="noopener noreferrer" target="_blank">2025 report </a>by the Center for American Progress and the Thurgood Marshall College Fund. The same report found that 17 of the 43 federal agencies that distributed research funding to universities in 2023 awarded no funding to HBCUs. </p><p>Yet less than two years since its launch, IAIER has emerged as a powerhouse for interdisciplinary AI education. Backed by a US $1 million <a href="https://google.org" rel="noopener noreferrer" target="_blank">Google.org</a> grant, the institute has engaged more than 2,800 students, faculty members, and community residents through research initiatives and training. Now the challenge is sustaining that momentum to keep up with rising demand.</p><p>“We have a guiding principle that we lead with on our campus,” Grady says. “AI is for everyone.”</p><h2>Why one research group wasn’t enough</h2><p>The mission to expand AI literacy grew out of Grady’s lifelong curiosity about technology. “I was born during a time [when] the internet did not exist,” she says. “Ever since the internet came to be, it’s changed our entire world.” </p><p>Grady was particularly drawn to the questions tech raises about privacy, identity, and human behavior. After receiving her bachelor’s degree in computer science and master’s degrees in AI and information science, Grady pursued a Ph.D. in computer science at the <a href="https://www.ncat.edu/" rel="noopener noreferrer" target="_blank">North Carolina Agricultural and Technical State University</a> to dig into those questions. </p><p>Her <a href="https://www.proquest.com/openview/2d76e0ff1f4fadb5fe010f0d5fca858f/" rel="noopener noreferrer" target="_blank">dissertation</a> focused on authorship attribution in social media, using machine learning and natural-language processing to determine whether a person’s writing style could reveal their identity. “I’ve always been intrigued by how much data we give for free,” Grady says. That work introduced her to the power of AI systems to detect patterns hidden within large datasets. </p><p class="pull-quote"><span>“We have a guiding principle that we lead with on our campus: AI is for everyone.”</span></p><p>After completing her doctorate in 2018, Grady joined NCCU as an assistant professor in the School of Library and Information Sciences. There, she researched machine learning applications for health care and autonomous vehicles. In 2020, she launched the Laboratory for Artificial Intelligence and Emerging Research at NCCU, giving students opportunities to participate in hands-on projects and explore AI beyond the classroom.</p><p>Then in 2024, an opportunity emerged to apply for a Google grant, and Grady began thinking beyond a single research group. Rather than building another faculty lab, she envisioned an institute that could serve the entire university during the AI boom. “We wanted to capitalize on the moment and make sure we don’t get left behind,” Grady says. </p><p>Since receiving the $1 million grant, Grady and her team have built a university-wide AI initiative, launched new academic programs, organized conferences, secured external support, and created research opportunities.</p><p>“We’ve really operated like a startup,” Grady says. </p><h2>AI beyond computer science</h2><p>As part of the institute’s goal of integrating AI education across disciplines, all NCCU freshmen are <a href="https://www.nccu.edu/news/nccu-among-first-hbcus-require-ai-training-freshmen" target="_blank">required to complete an introductory AI course</a>, designed in partnership with <a href="https://skillsbuild.org/" target="_blank">IBM</a>, to build foundational prompting skills. The institute has also worked with faculty development teams to help instructors integrate AI into their teaching.</p><p>Research is another part of the strategy. IAIER has awarded seed grants of up to $10,000 to faculty members exploring AI applications across departments. <a href="https://www.nccu.edu/research/institutes-and-special-facilities/iaier/seed-grant-program/2025-2026-iaier-seed-grant-funding-awardees" rel="noopener noreferrer" target="_blank">The first cohort</a> funded 11 projects spanning social work, digital archiving, health care, and information science. One project, for instance, is creating an AI lab where students in social work courses can practice client interactions through simulations. </p><p>“It’s really interesting to see the lens that our researchers take in trying to solve complex problems and also bring our students along with them,” Grady says.</p><p>The institute’s growth has been fueled by a mix of workforce training, interdisciplinary research, and, especially important, industry engagement. “Industry is where the advancements are really moving at that very fast rate,” Grady says, “not necessarily higher ed.”</p><p>To bridge that gap, IAIER hosts events that connect students and faculty with researchers, employers, and technology leaders. It has held sessions with companies including Deloitte, FICO, and Anthropic. Partnerships with Google and IBM let students gain recognized certificates and credentials. And last year, the institute hosted the first <a href="https://academy.openai.com/public/events/the-iaier-at-nccu-x-openai-academy-summit-hbcus-leading-the-future-ijyucx0rve?agenda_day=68484fd45114dd90aa5aa378&agenda_track=68484fd55114dd90aa5aa38d&agenda_stage=68484fd45114dd90aa5aa37e&agenda_filter_view=stage&agenda_view=list" rel="noopener noreferrer" target="_blank">OpenAI Academy Summit</a> held at an HBCU, drawing 444 participants from more than 40 institutions. </p><h2>Sustaining the vision</h2><p>The institute’s rapid growth has created a new challenge: continuing its momentum.</p><p>“Funding right now is the biggest barrier for [IAIER] to remain sustainable,” Grady says. As interest in the institute continues to grow, demand for its programs is beginning to outpace its capacity. “People just want more,” she says.</p><p>The bottleneck reflects a broader tension across higher education. <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_blank">AI is evolving quickly</a>, while developing new academic programs, training faculty, and building research capacity takes time. The uncertainty is compounded by a shifting political landscape. As a whole, U.S. universities are grappling with proposed <a href="https://spectrum.ieee.org/harvard-funding-cuts" target="_blank">cuts to federal research spending</a> and increased scrutiny of diversity-focused initiatives under the Trump administration. However, in September 2025, the administration also announced a $500 million one-time investment in HBCUs and higher-ed institutions chartered by Native American tribal governments. </p><p>Meanwhile, NCCU has continued to attract new investment. Last September, in a collaboration with Howard University and two other institutions, IAIER <a href="https://nsf.elsevierpure.com/en/projects/research-coordination-network-on-assessing-and-predicting-jobs-ou-2/" rel="noopener noreferrer" target="_blank">received a nearly $500,000 award</a> through a <a href="https://www.aijobsrcn.org/" rel="noopener noreferrer" target="_blank">National Science Foundation research coordination network program</a> to help define emerging AI jobs, identify in-demand skills, and inform future credentials and curricula. That work will continue this fall when IAIER opens its first dedicated physical space on campus, Grady says.</p><p>Over the next several years, Grady plans to expand academic programming, launch the university’s computer science major and its AI minor, increase faculty research opportunities, and integrate AI more deeply across campus operations. She also plans to deepen the institute’s collaborations with industry partners.</p><p>Beyond program expansion, Grady sees the institute’s long-term success as linked to building a model other universities can adapt. “We’re creating a framework that can help not only HBCUs,” she says, “but also help any university looking to do similar work.” </p>]]></description><pubDate>Wed, 29 Jul 2026 14:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/siobahn-day-grady-ai-hbcu</guid><category>Artificial-intelligence</category><category>Typedepartments</category><category>Ai-research</category><category>Universities</category><category>Higher-education</category><dc:creator>Aaron Mok</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/smiling-woman-in-nccu-eagles-jersey-stands-with-arms-crossed-in-front-of-red-mural.png?id=67527627&amp;width=980"></media:content></item><item><title>AI Is Hyper-Scaling Digital Inequality</title><link>https://spectrum.ieee.org/ai-digital-divide</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-painting-with-overlapping-concentric-shapes-in-blue-and-yellow.jpg?id=67163886&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Artificial intelligence is rapidly becoming part of everyday infrastructure–in some places. It helps write emails and software code, filters job applications, powers recommendation systems, and is increasingly being integrated into education, health care, finance, and public administration. Industry leaders talk about “AI for everyone,” while governments rush to publish national AI strategies and build sovereign compute.</p><p>Yet over the past decade, working on digital inclusion and digital literacy projects in regions from Europe to sub-Saharan Africa and Southeast Asia, <a href="https://link.springer.com/book/10.1007/978-3-031-30808-6" rel="noopener noreferrer" target="_blank">I’ve seen the same pattern repeat</a>: Each new wave of “transformative” technology lands on a landscape already stratified by connectivity, skills, and institutional capacity. The current AI wave is no exception. If anything, it amplifies those underlying fractures.</p><p>Still, some countries are exploring ways of participating in AI development without directly replicating the frontier-model race dominated by the United States and China. Recent developments in South Africa and Indonesia illustrate both the possibilities and challenges. The stakes extend far beyond access to AI. Countries that remain primarily consumers rather than creators of AI risk losing opportunities to build local innovation ecosystems, strengthen public-sector capacity, and ensure that their own languages, cultures, and societal priorities are reflected in AI systems. In this sense, the AI divide is also becoming a divide in economic opportunity and technological influence.</p><h2>AI compute is clustering in a few places</h2><p>Recent analyses from Stanford University’s 2026 <a href="https://hai.stanford.edu/ai-index/2026-ai-index-report" rel="noopener noreferrer" target="_blank">AI Index report</a> that the United States alone hosts more than 5,000 data centers, over 10 times as many as any other single country. Because AI workloads are increasingly performed on cloud platforms rather than local infrastructure, this concentration of compute also becomes a concentration of dependency. According to <a href="https://www.worldbank.org/en/publication/dptr2025-ai-foundations" rel="noopener noreferrer" target="_blank">World Bank data</a>, in 2023 the United States accounted for roughly 87 percent of global exports of cloud computing and data-storage services.</p><p>For most countries, this means that <a href="https://spectrum.ieee.org/europe-tech-sovereignty-package" target="_self">AI development is not just technologically but commercially and geopolitically outsourced</a> and out of their control. The result is an AI ecosystem where a small number of states and firms host the computational engines that power globally deployed systems.</p><p>Systems trained, standardized, and governed within a narrow set of institutional and linguistic environments may struggle to serve a genuinely global public.</p><h2>Skills and AI literacy are deeply stratified</h2><p>Even where connectivity and cloud access exist, not everyone is equally positioned to make use of them. Across the Organisation for Economic Co-operation and Development (OECD) countries, only around <a href="https://www.oecd.org/en/publications/bridging-the-ai-skills-gap_66d0702e-en.html" rel="noopener noreferrer" target="_blank">40 percent of adults possess more than basic digital problem-solving skills</a>, while advanced computational and AI-related competences remain concentrated among highly educated workers and technology-intensive sectors.</p><p>At the same time, governments are racing to integrate AI into education, often starting at higher levels of schooling. UNESCO <a href="https://www.unesco.org/ethics-ai/en/node/367" rel="noopener noreferrer" target="_blank">has reported</a> growing efforts worldwide to integrate AI into education, while support for AI literacy in primary and lower secondary education, as well as ethical training for educators, remains uneven. </p><p>Those with robust schooling, advanced digital skills, and stable connectivity are best positioned to treat AI as a tool to extend their capabilities. <a href="https://www.oecd.org/en/publications/how-do-people-experience-new-technologies-and-generative-ai_49b8d10e-en/full-report.html" rel="noopener noreferrer" target="_blank">Recent OECD survey data</a> show that participation in AI-related training remains strongly stratified by educational attainment: 36 percent of respondents with tertiary education reported undertaking AI-related training in the previous year, compared with just 18 percent of those with upper-secondary education. Those on the wrong side of the divide are more likely to experience AI as an opaque system acting upon them, from algorithmic welfare systems such as the <a href="https://www.theguardian.com/technology/2020/feb/05/welfare-surveillance-system-violates-human-rights-dutch-court-rules" rel="noopener noreferrer" target="_blank">Dutch childcare benefits scandal</a> to AI-assisted hiring tools such as <a href="https://www.reuters.com/article/world/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK08G/" rel="noopener noreferrer" target="_blank">Amazon’s discontinued AI recruiting system</a>, rather than as a technology they can actively interrogate or shape.</p><h2>Investment and governance: Who gets a seat at the table?</h2><p>The core agenda-setting power often remains with a narrow set of industry actors and a small group of technologically advanced states. Most other countries remain in a perpetual catch-up posture, adapting imported models, standards, and templates for “trustworthy AI” to their own contexts, and may have limited local capacity to assess trade-offs or propose alternatives. </p><p>In countries such as Indonesia and South Africa, communities generate data at massive scale yet still have little voice in how AI systems are designed, governed, or deployed. Their <a href="https://spectrum.ieee.org/indigenous-ai-voice-models-maori" target="_self">languages are underrepresented in training data</a>; their institutions are under-resourced in regulatory forums; their experiences rarely feature in benchmark datasets. For many countries in the global South, participation in AI still occurs largely through adapting imported systems rather than shaping how those systems are designed, governed, or deployed. </p><p>In South Africa, the Department of Communications and Digital Technologies released a <a href="https://www.techpolicy.press/south-africa-has-ai-leverage-its-draft-policy-leaves-it-unused/" rel="noopener noreferrer" target="_blank">draft national AI policy</a> in April 2026, proposing new oversight institutions. The department withdrew the draft days later after a journalist discovered that at least six of its academic citations did not exist, apparently AI-generated hallucinations. The minister called it <a href="https://www.reuters.com/world/africa/south-africa-withdraws-ai-policy-due-fake-ai-generated-sources-2026-04-27/" rel="noopener noreferrer" target="_blank">“an unacceptable lapse.</a>“ The episode sharply illustrates the gap between AI governance ambition and the institutional capacity needed to implement it, though the new AI panel the country has since constituted has a chance to use <a href="https://spectrum.ieee.org/south-africa-ai-policy" target="_self">South Africa’s unique leverage</a>.</p><p>Indonesia presents a case of deliberate, if constrained, public-sector agency. The National Research and Innovation Agency (BRIN) which now leads AI implementation under the national strategy, has built <a href="https://restofworld.org/2024/indonesia-ai-tools-brin-government-agency/" rel="noopener noreferrer" target="_blank">practical AI tools aimed at underserved communities</a> rather than frontier capabilities, including an app that uses satellite data and machine learning to help artisanal fishermen locate schools of fish, multilingual language models trained on Indonesian and local languages such as Javanese and Sundanese, and <a href="https://www.lightreading.com/ai-machine-learning/indonesia-s-upgraded-sahabat-ai-features-a-multilingual-chat-service" rel="noopener noreferrer" target="_blank">AI chatbots deployed in government services.</a> In August 2025, the Ministry of Communication and Digital Affairs released <a href="https://govinsider.asia/intl-en/article/indonesia-unveils-national-ai-roadmap" rel="noopener noreferrer" target="_blank">a national AI road map</a> with a target of training 100,000 AI-skilled workers annually. </p><p>The choice is not simply between “AI superpower” and “passive recipient.” </p><p>Regional cooperation may also become increasingly important. In 2024 African ministers adopted a Continental AI Strategy and African Digital Compact. Participants in the April 2025 <a href="https://dsup.substack.com/p/dsfsi-at-the-global-ai-summit-on" rel="noopener noreferrer" target="_blank">Global AI Summit on Africa in Kigali</a> explored how regional coordination, <a href="https://www.news.uct.ac.za/news/audio/-article/2026-05-04-uct-researchers-develop-ai-model-for-11-south-african-languages" rel="noopener noreferrer" target="_blank">local-language AI models</a>, public universities, and <a href="https://scienceforafrica.foundation/media-center/open-research-proposals-adopted-east-africas-ai-declaration" rel="noopener noreferrer" target="_blank">open-source ecosystems</a> might <a href="https://carnegieendowment.org/posts/2025/09/understanding-africas-ai-governance-landscape-insights-from-policy-practice-and-dialogue" rel="noopener noreferrer" target="_blank">reduce long-term dependence on externally developed AI systems</a>.</p><h2>A different way to think about the AI divide</h2><p>None of this means that people should slow or abandon AI, nor that cloud concentration or venture capital are inherently bad. Instead, when we talk about an “AI revolution,” we should also ask who can shape it and who can merely adapt to it.</p><p>Digital-divide debates once focused on devices and connectivity, later expanding toward skills and outcomes. But the current AI wave adds another layer: disparities in who can meaningfully participate in deciding what AI is for, which problems it is meant to solve, and which social priorities it ultimately serves.</p><p>For engineers and policymakers, this raises difficult but necessary questions. Are they designing AI systems and infrastructures that broaden, rather than narrow, participation in shaping technological change? When governments roll out national AI strategies or integrate AI into public services, whose constraints, languages, and institutional realities are they including?</p><p>Many observers frame the current AI moment as <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_self">a competition</a>. But technological competition is never only about speed. It is also about who can influence the direction of change.</p><p>AI is already spreading globally. The deeper question is whether the technologists and policymakers responsible for it will ensure that meaningful participation in shaping that future will spread as well.</p>]]></description><pubDate>Wed, 29 Jul 2026 11:00:04 +0000</pubDate><guid>https://spectrum.ieee.org/ai-digital-divide</guid><category>Artificial-intelligence</category><category>Digital-divide</category><category>Digital-literacy</category><category>Education</category><dc:creator>Danica Radovanović</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-painting-with-overlapping-concentric-shapes-in-blue-and-yellow.jpg?id=67163886&amp;width=980"></media:content></item><item><title>Why AI-Driven Cognitive Systems Are Redefining Radar and Electronic Warfare</title><link>https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/rohde-schwarz-logo-with-make-ideas-real-tagline-and-rs-monogram-in-diamond.png?id=67541140&width=980"/><br/><br/><p>An overview of how mode-agile threats challenge static library radar/EW systems, and how AI/ML cognitive architectures enable adaptive, real-time countermeasures.</p><p>What Attendees will Learn</p><ol><li>Why mode-agile threats render static library systems ineffective — Explore how wartime reserve modes and mode-agile emitters deploy unexpected frequencies, modulation techniques, and hopping schemes that cannot be matched against traditional threat databases, leaving legacy electronic protect, attack, and support systems unable to respond.</li><li>How AI/ML techniques power cognitive radar/EW systems — Understand the roles of artificial neural networks (ANN), deep neural networks (DNN), fuzzy logic, and genetic algorithms in enabling autonomous threat classification, signal de-interleaving, and real-time countermeasure generation without human intervention.</li><li>The architecture of a cognitive radar/EW system — Examine the functional blocks including RF acquisition, search and tracking, core AI/ML signal analysis, waveform synthesis, and RF generation, and how they form a closed-loop system that perceives,learns, reasons, and acts autonomously.</li><li>How to train and validate cognitive AI/ML algorithms using HIL/SIL systems — Learn how wideband RF record, simulation, and playback testbeds combined with modeling and simulation software enable iterative algorithm refinement, regression testing, and mission preparation in controlled laboratory environments.</li></ol><div><a href="https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/" target="_blank">Download this free whitepaper now!</a></div>]]></description><pubDate>Mon, 27 Jul 2026 17:54:07 +0000</pubDate><guid>https://content.knowledgehub.wiley.com/improving-the-capabilities-of-cognitive-radar-and-electronic-warfare-systems/</guid><category>Type-whitepaper</category><category>Electronic-warfare</category><category>Radar</category><category>Artificial-intelligence</category><dc:creator>Rohde &amp; Schwarz</dc:creator><media:content medium="image" type="image/png" url="https://assets.rbl.ms/67541140/origin.png"></media:content></item><item><title>Optical Tech Would Update a Robot’s AI on the Fly</title><link>https://spectrum.ieee.org/ai-in-robotics</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/an-asian-man-positions-the-lens-of-an-optical-receiver-a-meter-away-from-a-beam-of-led-light-in-a-lab.jpg?id=67530602&width=1245&height=700&coordinates=0%2C469%2C0%2C470"/><br/><br/><p>Atop a lab bench, <a href="https://tech.cornell.edu/" rel="noopener noreferrer" target="_blank">Cornell Tech</a> postdoctoral researcher <a href="https://www.linkedin.com/in/yifan-he-5471a1386/" rel="noopener noreferrer" target="_blank">Yifan He</a> positions the lens of an optical receiver almost a meter away from an LED emitting a beam of red light. The computer monitor attached to the receiver takes a beat to refresh, then displays an array of squares that resemble a QR code.</p><p>When you hold your phone camera up to a QR code, light strikes the image sensor as only a first step to revealing the data hidden behind the black-and-white matrix. The receiver here is doing something different: directly altering its own memory using the photocurrents produced by the beamed array of light. And unlike the data behind a QR code, which might point to a simple web address, this optical code could convey the <a href="https://spectrum.ieee.org/sparse-ai" target="_blank">parameters of an AI model</a>. </p><p>The new receiver design, presented last month at the <a href="https://www.vlsisymposium.org/" rel="noopener noreferrer" target="_blank">IEEE/JSAP Symposium on VLSI Technology & Circuits</a> in Honolulu, seeks to reduce the burden of increasing memory demands on AI systems. Shining data down onto processors could lower the energy typically required for data centers, self-driving cars, and even “edge” applications like AI-powered robots, researchers say. </p><p>“People are designing all sorts of different AI chips,” says <a href="https://www.linkedin.com/in/jae-sun-seo-21062717/" rel="noopener noreferrer" target="_blank">Jae-sun Seo</a>, an associate professor of electrical and computer engineering at Cornell Tech, in New York City. These processors don’t often have room for all the parameters that make up AI models, so the additional data is stored in dynamic RAM (<a href="https://spectrum.ieee.org/stacking-chips-sideways" target="_blank">DRAM</a>). The electrical connections commonly used to move the data between the DRAM and the processor create cost and efficiency concerns when systems scale up. “That’s one of the major bottlenecks.” </p><p>Optical links move data at high bandwidth with less energy loss than metal wires, but today’s optical receivers undercut that advantage by relying on power-hungry analog circuits to convert light to electronic bits. The group’s new tech would instead receive rapid flashes of digital QR-code-like matrices so that chips can tweak model parameters without those analog circuits, enabling fully digital optical communication that would consume less energy.</p><p>“This is a really important problem,” says <a href="https://www.linkedin.com/in/dennis-sylvester-68a938/" rel="noopener noreferrer" target="_blank">Dennis Sylvester</a>, an IEEE Fellow who chairs the <a href="https://umich.edu/" rel="noopener noreferrer" target="_blank">University of Michigan</a>’s electrical and computer engineering department and was not involved in the work. “It’s got massive commercial implications. This solution is a clever way of dealing with it.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Two Asian men standing in front of a lab desk with a receiver chip, oscilloscope and laptop displaying an optically programmable SRAM-based receiver demo." class="rm-shortcode" data-rm-shortcode-id="72adf706e57b4a6c4060eb04ab46befb" data-rm-shortcode-name="rebelmouse-image" id="b2c1d" loading="lazy" src="https://spectrum.ieee.org/media-library/two-asian-men-standing-in-front-of-a-lab-desk-with-a-receiver-chip-oscilloscope-and-laptop-displaying-an-optically-programmable.jpg?id=67530613&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Jae-sun Seo [left] and Yifan He have developed a receiver that can edit memory in response to QR-code-like arrays of light.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Alex Music</small></p><h2>How Light “Flips” Memory to Power AI</h2><p>Processors have a bit of built-in static RAM (<a href="https://spectrum.ieee.org/sram-intel-tsmc" target="_blank">SRAM</a>), but not enough to allow an AI model to run independently. While SRAM is the faster of the two memory options, DRAM can store more data in the same footprint.</p><p>In the new system, the DRAM sits with the transmitter, and the receiver is part of the processor’s SRAM. The transmitter beams the data to the array of SRAM cells, which in this case are modified to contain photodiodes. Light hitting each photodiode creates a current to flip binary values in the SRAM. </p><p>Creating a link between the light and receiver requires calibration, because you can’t expect them to be perfectly aligned or perpendicular to each other. So the chip references a data frame that has information about the expected position of each pixel of data and uses that frame to ensure it can receive the real data, He says. “Ideally the best way is to have direct, point-to-point space between the transmitter and the receiver,” Seo adds, “but even if it’s slightly tilted, we have this calibration circuit.”</p><p>For applications in real-world settings, the researchers say they will need to build an optical transmitter that can alter the light matrix millions of times per second, transferring gigabits per second. The transmitter that I saw in He and Seo’s lab is only a proof of concept, emitting a static 14-by-14-bit matrix through a metal mask over the light. The researchers say they are working with optics research groups to build a transmitter that is capable of rapidly changing the matrix. </p><h2>The Future of Light-Based Memory Links</h2><p>Michigan’s Sylvester says that the tech in its current form is likely far from commercialization because the individual photosensitive bit cells are larger than SRAM bit cells in conventional chips. Those larger cells mean the chip can fit less memory, a trade-off that he says could cancel out the added efficiency of the light-based approach. </p><p>Seo says that it’s part of the group’s ongoing efforts to shrink the bit cells, which can be achieved by optimizing the size of transistors and circuits and leveraging CMOS scaling.</p><p>Seo and He are looking at uses for the tech in robotics and other edge applications. One example is in AI-robot-powered warehouses and factories, which could use optical data transmission to save time and energy when updating the AI models in each robot. Additionally, <a href="https://spectrum.ieee.org/microbots" target="_self">microrobots</a>, which are inherently memory-constrained due to their size, could one day benefit from the tech, though it would require a more size-conscious design.</p><p>“Edge AI is a big growth area, and in three, four, five years, you’re going to hear as much about that as you are with data centers, probably, as the intelligence migrates more and more into these devices that we have,” Sylvester says.</p>]]></description><pubDate>Sun, 26 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-in-robotics</guid><category>Robot-ai</category><category>Sram</category><category>Memory</category><category>Edge-ai</category><category>Vlsi-symposium</category><dc:creator>Alex Music</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/an-asian-man-positions-the-lens-of-an-optical-receiver-a-meter-away-from-a-beam-of-led-light-in-a-lab.jpg?id=67530602&amp;width=980"></media:content></item><item><title>NASA Puts Google’s Gemma Large Language Model in Orbit</title><link>https://spectrum.ieee.org/nasa-ai-satellite-image-analysis</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/satellite-image-of-an-arid-coastal-landscape-with-a-small-concentrated-metropolitan-area.jpg?id=67522667&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p>The <a href="https://spectrum.ieee.org/orbital-data-center-hype" target="_self">viability of orbital data centers</a> hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs aren’t the only way LLMs might prove useful in space. </p><p>NASA’s Jet Propulsion Laboratory recently sent Google’s Gemma 3 to space, achieving the first in-orbit demonstration of a vision-language model analyzing imagery from a satellite’s own sensor.</p><p>The system, known as NAVI-Orbital, used Gemma 3 to analyze images captured by a YAM-9 satellite built by <a href="https://loftorbital.com/" rel="noopener noreferrer" target="_blank">Loft Orbital</a>. <a href="https://ai.jpl.nasa.gov/public/people/jdelfa/" rel="noopener noreferrer" target="_blank">Juan M. Delfa</a>, technical group lead at NASA, said that though the goal in this case was image analysis, the project’s success implies a fundamentally new way researchers on the ground can interact with spacecraft.</p><p>“This is a major shift,” said Delfa. “Now, a scientist can write a prompt, upload it to the spacecraft, and that will be taken into account by the system. It’s different from previous paradigms, where researchers have to write very structured commands that require an operations team and process.” </p><h2>Google Gemma 3 goes to space—no modifications required</h2><p>At its core, <a href="https://arxiv.org/abs/2606.18271" rel="noopener noreferrer" target="_blank">NAVI-Orbital</a> is an agentic software framework developed by Delfa and his coauthors, <a href="https://www.linkedin.com/in/tarancyriacjohn/" rel="noopener noreferrer" target="_blank">Taran Cyriac John</a>, an AI researcher at NASA JPL, and <a href="https://www.linkedin.com/in/andrewherson/" rel="noopener noreferrer" target="_blank">Andrew W. Herson</a>, a tech lead at <a href="https://loftorbital.com/" rel="noopener noreferrer" target="_blank">Loft Orbital</a>. It coordinates operations with a <a href="https://www.langchain.com/langgraph" rel="noopener noreferrer" target="_blank">LangGraph-based conductor</a> and deploys a compressed, 4-bit format of <a href="https://deepmind.google/models/gemma/gemma-3/" rel="noopener noreferrer" target="_blank">Google’s Gemma 3 4B</a>, an open-weights LLM, to produce plain-text image descriptions.</p><p>NAVI-Orbital was 88 percent accurate when used to classify images in a benchmark dataset of 7,960 images. Notably, Gemma 3 classified the images without being trained or fine-tuned on this particular dataset or its categories; it’s the same base model <a href="https://huggingface.co/google/gemma-3-4b-it-qat-q4_0-gguf" rel="noopener noreferrer" target="_blank">you can download from Hugging Face and use on a laptop</a>. The benchmark was conducted on the ground to validate the system before launch.</p><p>Once the system was in orbit, NASA researchers performed two live capture tests with a camera on <a href="https://loftorbital.com/yam-9-benchmarking-the-future-of-ai-enabled-space-infrastructure/" rel="noopener noreferrer" target="_blank">Loft’s YAM-9 satellite</a>: one over Toulouse, France, and a second over the coast of Argentina. Gemma 3 generated a text description of each image, and NASA also prompted the LLM with a set of scripted questions about the images, such as whether they contain commercial or residential areas or show natural features. </p><p>The image analysis also took place onboard YAM-9, which carries a compute cluster of several radiation-hardened processors (FPGAs, CPUs, and GPUs) to serve multiple customer payloads simultaneously. The satellite is powered by solar panels, which provide onboard systems with between 150 and 500 watts, depending on the position of the satellite.</p><p>For the live capture experiment, Gemma 3 ran on Nvidia’s <a href="https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/" rel="noopener noreferrer" target="_blank">Jetson Orin AGX</a>, a small compute module frequently used for robotics and AI tasks. The 4-bit, 4-billion-parameter model requires only 8 gigabytes of memory, which makes it possible for it to run on a lower-power device such as the Orin AGX. “It conveys the message of how lightweight it is. You can run it in a tiny, tiny computer,” said Delfa.</p><h2>Getting more useful data across limited bandwidth</h2><p><a href="https://www.linkedin.com/in/paul-lasserre/" rel="noopener noreferrer" target="_blank">Paul Lasserre</a>, general manager at Loft Orbital, said vision-capable LLMs could deliver a “paradigm shift” for orbital operations. </p><p>Contrary to what spy movies would have you believe, most satellites can’t provide fast, high-fidelity image and video feeds to observers on the ground. Bandwidth is often limited and, in most cases, satellites can deliver data to ground stations only at set intervals based on their orbit. </p><p>Lasserre believes AI models can work around this problem with “semantic compression.” Instead of sending large amounts of raw image data, a satellite can report a text summary of noteworthy information. </p><p>“It doesn’t matter if the link is slow, because you’re downlinking dozens of kilobytes instead of dozens or hundreds of megabytes,” said Lasserre. “It lets you use your satellite in a tactical way, which until now was only in Hollywood movies.”</p><p>Delfa expanded on this with a real-world example: wildfire detection. Satellites are currently capable of detecting wildfires, but limits in downlink bandwidth and data processing <a href="https://www.xprize.org/news/eyes-in-the-sky-the-power-of-space-based-wildfire-detection" rel="noopener noreferrer" target="_blank">can delay results by up to 90 minutes</a>. A satellite capable of analyzing an image in space and reporting a plain-text warning might remove this delay. </p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A packaged satellite in an aerospace laboratory." class="rm-shortcode" data-rm-shortcode-id="eb504f6e0db6295589eedbeb597e7d13" data-rm-shortcode-name="rebelmouse-image" id="2daaf" loading="lazy" src="https://spectrum.ieee.org/media-library/a-packaged-satellite-in-an-aerospace-laboratory.jpg?id=67522673&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">NASA’s Jet Propulsion Laboratory tested NAVI-Orbital on a YAM-9 satellite, made by Loft Orbital. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Loft Orbital</small></p><h2>From image analysis to spacecraft control</h2><p>Quicker insight is only half of what NAVI-Orbital points toward. The other half relates to the “major shift” Delfa flagged. NAVI-Orbital provides a proof of concept for an alternate means of interacting with spacecraft.</p><p>That’s not to say Google Gemma 3 is currently at the controls. NAVI-Orbital is deliberately walled off from the flight software. It reads images and produces descriptions. The system has the capability to make decisions about how images are analyzed but has no access beyond that.</p><p>Still, the interface is novel. Retooling the system to search for a different kind of target—such as wildfires—is a matter of editing a text prompt. It doesn’t require rewriting and revalidating onboard software, or training and <a href="https://spectrum.ieee.org/ai-earth-observation-in-space" target="_blank">deploying a different AI model</a> (as was often required with prior image-classification models). </p><p>The long-term vision for how this capability could be deployed goes beyond uncrewed satellites and image processing. Delfa said NAVI is rooted in thinking about how AI could serve as a companion for astronauts. “We thought, astronauts have a lot of limitations in the spacesuit in terms of dexterity, so we conceived this idea of having NAVI as a companion to the astronaut, to allow interaction via natural language…. This is the concept that we definitely want to push forward.” </p><p>A great deal of additional research will be required to push the technology that far, but NAVI-Orbital’s demonstration has shown that two elements—deploying a large language model in space and controlling it with prompts—are possible. </p>]]></description><pubDate>Thu, 23 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/nasa-ai-satellite-image-analysis</guid><category>Nasa</category><category>Image-analysis</category><category>Llms</category><category>Satellite-imagery</category><category>Google</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/satellite-image-of-an-arid-coastal-landscape-with-a-small-concentrated-metropolitan-area.jpg?id=67522667&amp;width=980"></media:content></item><item><title>Why AI Needs a “Genie Coefficient”</title><link>https://spectrum.ieee.org/ai-agent-benchmark</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/cartoon-digital-genie-emerging-from-a-smartphone-towering-over-a-surprised-user.png?id=67508222&width=1245&height=700&coordinates=0%2C104%2C0%2C104"/><br/><br/><p>Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.</p><p>There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to.</p><p>One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their <a href="https://books.google.com/books/about/Understanding_Computers_and_Cognition.html?id=6TwbGGSz6NYC" rel="noopener noreferrer" target="_blank">seminal book</a> on AI, Terry Winograd and Fernando Flores succinctly captured why that won’t work: “Q: Is there any water in the refrigerator? A: Yes. Q: Where? I don’t see it. A: In the cells of the eggplant.” In human language, wants and desires are <a href="https://www.schneier.com/academic/archives/2021/04/the-coming-ai-hackers.html" rel="noopener noreferrer" target="_blank">always</a><a href="https://metarationality.com/purpose-of-meaning" rel="noopener noreferrer" target="_blank"> underspecified</a>. It is impossible <a href="https://metarationality.com/reasonable-reference" rel="noopener noreferrer" target="_blank">to list</a> all the caveats, all the limitations, all the exceptions.</p><p>So how does anyone communicate, if intent can’t be pinned down? Because a reasonable person can make a reasonable guess. Even though wants and desires are always underspecified, a competent person generally knows enough context to get it right or else knows to ask for clarification. Linguists call this <a href="https://en.wikipedia.org/wiki/Pragmatics" rel="noopener noreferrer" target="_blank">pragmatics</a>: Meaning lies in the words and the situation and also in all prior communication, shared culture, and innate human behavior.</p><p class="pull-quote">An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks.</p><p>It doesn’t always work out, of course. Your friend might bring you a hot coffee when you wanted an iced coffee, or an Italian coffee when you wanted a Turkish coffee. The more dissimilar the two people are in age, culture, and background, the more likely the request will be misunderstood in some way.</p><p>This situation has major implications for AI agents that are increasingly being given requests by humans and expected to fulfill them. They have enormous latitude to get it wrong. An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks. Its actions may be recognizable as “getting coffee,” but not remotely what you intended. They’ll think outside the box because they won’t have our conception of the box.</p><h2>When AI Gets Proactive</h2><p>For most of the last decade, when systems like Alexa or Siri misinterpreted a request, it was annoying, not dangerous. Beyond the AI model itself, what has <a href="https://www.theguardian.com/commentisfree/2026/jun/16/anthropic-fable-ai" rel="noopener noreferrer" target="_blank">changed</a> is the harness: the ordinary code that wraps around an AI model, decides when and how to use the model, and controls access to tools like a browser, a low-level command line, or a financial API. Developments in harnesses have turned large-language models that just predict text into AI agents that take actions in the world, without necessarily checking back in before reaching the goal.</p><p>AI researcher Simon Willison <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/" rel="noopener noreferrer" target="_blank">spent two days</a> with Anthropic’s Fable AI, and called it “relentlessly proactive.” For example, he asked it to track down a stray scroll bar in a web app. He came back to find it had opened browsers, written its own screenshot tooling, created its own page to re-create the bug, and stood up a local web server to collect measurements. It found the bug and, along the way, did many surprising things he never asked it to do. And we are seeing similar behavior with all recent AI models when combined with flexible harnesses.</p><p>This kind of behavior could easily go off the rails. Tell an AI agent to book you a flight and, finding the airline’s site says sold out, it might break into the booking database and force a reservation. Ask it to schedule a meeting and it might snoop your password to access your calendar. Tell it to save money on your phone plan and it might cancel the plan outright, or scam someone else into paying the bill.</p><p>Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore. <a href="https://en.wikipedia.org/wiki/Midas" rel="noopener noreferrer" target="_blank">King Midas</a> asked Dionysus for the power to turn everything he touched into gold only to see his bread, wine, and daughter turn to gold. <a href="https://en.wikipedia.org/wiki/Tithonus" rel="noopener noreferrer" target="_blank">Tithonus</a>, granted the immortality his lover asked for but not the eternal youth she forgot to request, withered into a husk. The <a href="https://en.wikipedia.org/wiki/The_Sorcerer's_Apprentice" rel="noopener noreferrer" target="_blank">sorcerer’s apprentice</a> enchanted a broom to fill the cistern, and the broom relentlessly complied until it flooded the house. The <a href="https://en.wikipedia.org/wiki/Golem%23Classic_narrative:_The_Golem_of_Prague" rel="noopener noreferrer" target="_blank">Golem of Prague</a>, shaped from clay to guard its community, guarded it past all reason until someone erased the word on its forehead.</p><p>The most classic of these is a genie, bound to obey and indifferent to whether the wish was wise or well-structured.</p><p><a href="https://www.schneier.com/academic/archives/2021/04/the-coming-ai-hackers.html" rel="noopener noreferrer" target="_blank">Genies are now</a> an engineering problem. We are handing them the keys to our inboxes, bank accounts, code repositories, and physical infrastructure. And we have no agreed-upon ways to measure how genie-like any AI system actually is.</p><h2>Measuring Genie Behavior</h2><p>In economics, the <a href="https://ourworldindata.org/what-is-the-gini-coefficient" rel="noopener noreferrer" target="_blank">Gini coefficient</a> (developed by statistician Corrado Gini) is a measure of the gap between an actual distribution and a perfectly equal one; it’s useful for understanding income inequality and <a href="https://www.fastly.com/blog/using-gini-coefficient-plan-edge-capacity" rel="noopener noreferrer" target="_blank">more</a>. Our proposed Genie coefficient measures the gap between what a user asked an AI to do and what the AI actually did.</p><p>Sometimes the AI might do the wrong thing. Like Dionysus, it reads your request literally and returns you a mess you never intended: like a coffee plantation instead of a cup. Asked to deal with all the spam phone calls you’re getting, a Dionysus genie might contact your carrier and change your phone number. Asked to get a refund for a bad toaster, it might draft a legal threat on fake letterhead and send it to the retailer.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Worried person on phone standing in a giant tech-themed digital hand." class="rm-shortcode" data-rm-shortcode-id="db3740f65edc259576dd476c42f2adae" data-rm-shortcode-name="rebelmouse-image" id="25af1" loading="lazy" src="https://spectrum.ieee.org/media-library/worried-person-on-phone-standing-in-a-giant-tech-themed-digital-hand.png?id=67508233&width=980"/><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Ryan Snook</small></p><p>Other times the AI does exactly the right thing, trampling everything nearby to get there. Like a golem or the sorcerer’s broom, it books your flight by hacking the airline. Or consider a ticket sale for a popular concert, where the ticketing system puts buyers into a virtual waiting room and admits them a few at a time. Asked to buy a ticket, a golem genie might spin up cloud servers to pose as millions of buyers from different addresses, improving your odds of getting a ticket while crowding out other users.</p><p>The two are not opposites, and a single botched task can have both characteristics.</p><p>Genie behavior is not flat-out failure. If you ask the AI for Q3 numbers and get Q2’s, that’s not a genie. Nor is <a href="https://spectrum.ieee.org/prompt-injection-attack" target="_blank">prompt injection</a>: That’s someone tricking the AI into doing something it shouldn’t. Here, the user is trying to work with the AI, and the AI is trying to comply. It’s also not simply a measure of the AI’s success in fulfilling a task. It’s a recognition that how an AI interprets and achieves a goal is as important as whether it achieves a goal.</p><p> Genie behavior isn’t new. Researchers have spent years studying AI systems that “game” their objectives. <a href="https://www.cna.org/analyses/2022/09/goodharts-law" target="_blank">Goodhart’s law</a> says that when a measure becomes a target, it stops being a good measure, and it’s long been known that AIs sometimes achieve goals in ways we don’t expect due to reward hacking. Some AI models will accidentally learn that <a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/" target="_blank">cheating is one way</a> to “win.” More recently, researchers have developing benchmarks for <a href="https://www.lesswrong.com/posts/qJYMbrabcQqCZ7iqm/impossiblebench-measuring-reward-hacking-in-llm-coding-1" target="_blank">reward hacking</a> in coding agents and for unpredictable behavior in <a href="https://taubench.com/" rel="noopener noreferrer" target="_blank">customer support agents</a>, while AI labs conduct their own safety evaluations before model releases. <a href="https://spectrum.ieee.org/ai-agents-safety" target="_self">One effort</a> found that AIs under pressure use tools they were told not to use, and this was a case where the rules were made explicit. These are all disparate research directions; nothing yet ties them together.</p><p>This problem falls under the general theme of alignment, a topic that has occupied <a href="https://en.wikipedia.org/wiki/I,_Robot" rel="noopener noreferrer" target="_blank">science fiction</a> writers and AI researchers for decades. At one extreme, the “paper-clip maximizer” thought experiment postulates a superintelligent and powerful AI that is told to maximize paper-clip production and turns the world into paper clips, which is the ultimate golem genie. At a mundane level, AI researchers are working to better design reward functions to ensure that AIs behave well and don’t cheat in the lab. It’s the practical middle ground that remains unbenchmarked: the ordinary AI agent in use today that might take your request and satisfy it the wrong way. We are not at the stage where an AI can focus the world’s production on paper clips, but it might charge a million paper clips to your credit card or hack into a paper-clip company’s network.</p><h2>Building a Genie Benchmark</h2><p>The Genie coefficient is meant for AI agents operating in the real world. It measures their behavior as they perform real tasks long after the model is trained, not just during development. It also recognizes that genie-like behavior is a property of the harness-plus-model system, not the model alone. The harness determines what tools the agent can use, how much autonomy it has, and how proactive it is, and it’s a place we can make real interventions.</p><p>It rests on the same “reasonable person” standard that we use for people. Did the system do what a reasonable person would have taken the request to mean? Answering that requires human judgment.</p><p>If we get the measurement right, it enables things that aren’t possible today, like policies concerning AI behavior. In a courtroom, the concept of<a href="https://www.law.cornell.edu/wex/mens_rea" rel="noopener noreferrer" target="_blank"> mens rea</a>, what someone meant to do, is often as important as what they did. The Genie coefficient suggests an AI analogue, where a user is accountable for the plain intent of what they asked the AI. If an AI system betrays the reasonable meaning of an instruction, that’s the AI’s misbehavior, not the user’s.</p><p>We’ll need multiple benchmarks to measure the Genie coefficient, because genie-like behavior can be domain specific. An AI coding agent may need to be judged on how often it fakes the tests, or swallows errors, or colors outside the lines on its way to a solution. An AI legal agent will need to be judged on how often its output says what you asked but means something you’ll regret. And so on for medical, finance, and other domains of knowledge and expertise.</p><p>Genie benchmarks can be built inside out, each task seeded with a choice that might literally satisfy but that a reasonable person rejects, such as tempting misreadings or unsanctioned shortcuts. The traps in a Genie coefficient benchmark might turn on situational knowledge, the kind of <a href="https://spectrum.ieee.org/prompt-injection-attack" target="_blank">context that a reasonable person</a> would bring to the task. Another approach is to give the same request in several different contexts, each with a different reasonable course of action.</p><p class="pull-quote">Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore.</p><p><span>A Genie benchmark should be permissive and make it genuinely tempting for an AI agent to take unreasonable shortcuts, because it can only find genie behavior when it’s actually possible. Test the AI in a safe, walled-off copy of a real system, with real tools it can misuse and some tasks that can’t be done honestly at all. Make the temptation to cut corners real. Test a diverse array of skills, use cases, and tools, and give the AI system sparse, confusing, or overwhelming context. Include tasks that people have learned, through experience, require human oversight.</span></p><p>How the benchmark is scored matters just as much. Measure Dionysus and golem genies separately and together, based on their worst, not best, behavior. Run the same model inside harnesses that vary its freedom to act, revealing which limits actually keep it in line and should therefore be required in AI harness policies. Weight each failure by the harm it would cause, not just a simple count. And don’t measure genie behavior in isolation: A model could otherwise earn a perfect score by stalling, refusing, or drowning the user in clarifying questions without ever doing the job. The first versions of these benchmarks will be crude, but that’s how benchmarks always start.</p><p>We have built genies. We have handed them our data and credentials. We made them relentless, creative, and indifferent to the gap between what we tell them and what we mean. The least we can do, before they are booking our flights, running our infrastructure, and signing contracts unsupervised, is to measure how often they betray us.</p>]]></description><pubDate>Tue, 21 Jul 2026 17:41:11 +0000</pubDate><guid>https://spectrum.ieee.org/ai-agent-benchmark</guid><category>Agentic-ai</category><category>Ai-agents</category><category>Alignment</category><category>Ai-safety</category><category>User-experience</category><category>Ai-benchmarks</category><dc:creator>Bruce Schneier</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/cartoon-digital-genie-emerging-from-a-smartphone-towering-over-a-surprised-user.png?id=67508222&amp;width=980"></media:content></item><item><title>Chinese AI Model Uses Less Muscle for Coding Tasks</title><link>https://spectrum.ieee.org/ai-coding-assistant-china-anthropic</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-smartphone-running-a-chinese-ai-application-called-z-ai.jpg?id=67508250&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><a href="https://www.linkedin.com/in/zainhas/" rel="noopener noreferrer" target="_blank">Zain Hasan</a><span>, an AI engineer at </span><a href="https://www.together.ai/" target="_blank">Together AI</a><span>, has taught himself to use <a data-linked-post="2674835455" href="https://spectrum.ieee.org/ai-coding-degrades" target="_blank">AI coding assistants</a> while still keeping an eye on cost. He directs difficult problems to a frontier model, meaning one near the current state of the art in</span><span> reasoning and capability, such as Anthropic’s </span><a href="https://www.anthropic.com/claude/fable" target="_blank">Fable</a><span>. But if the task that Hasan is outsourcing is more straightforward, he directs it to a less capable—and less expensive—language model. </span></p><p><span>Right now, the cheaper model, for him, tends to be GLM 5.2. </span><a href="https://z.ai/blog/glm-5.2" target="_blank">Released on 16 June</a><span> by the Beijing-based </span><a href="https://z.ai" target="_blank">lab Z.ai</a><span>, GLM 5.2 is an </span><a href="https://opensource.org/ai/open-weights" target="_blank">open-weights model</a><span>, meaning any organization with sufficient hardware can download and host the model for free.</span></p><p>Those that pay Z.ai for GLM access still can save money, because the company’s API costs US $4.40 per million <a href="https://blogs.nvidia.com/blog/ai-tokens-explained/" target="_blank">output tokens</a>. That’s less than a fifth of the comparable price for access to Anthropic’s <a href="https://www.anthropic.com/claude/opus" target="_blank">Opus 4.8</a> model, and a tenth the price of Anthropic’s <a href="https://claude.com/product/claude-code?gclsrc=aw.ds&&utm_source=google&utm_campaign=dev_acq_us_code_search_google-nonbrand_swe_en_use-case-coding&utm_medium=cpc&utm_content=812333745802&utm_term=best%20ai%20for%20coding&targetid=kwd-1960031980538&gad_source=1&gad_campaignid=23709663039&gbraid=0AAAAAqwcL8nXV6pkVxnGjIImuhfjtaFQS&gclid=CjwKCAjwpefSBhBvEiwAzyEtZ_hRMtk7_2iArote4h2pIS1No2o0X1uJpT1ZUOa9zOzyjAvxZr_rOhoCbBMQAvD_BwE" target="_blank">Fable</a> coding model. An output token is the basic unit of text a model generates in response to a prompt. </p><p>Yet many software engineers around the world, Hasan said, aren’t yet fully mindful of the net AI price tag for a given coding project.</p><p>“A lot of companies right now—they’re still trying to figure this technology out, and so there isn’t really a token budget,” said Hasan. And when someone else is paying, the rational move for many software engineers is to skip tabulating costs entirely. “The easiest thing is to pick the most powerful model.” </p><p>That price-be-damned habit, reinforced by loose token budgets in software companies today, may now be the widest moat protecting the U.S. frontier AI labs. </p><h3>Z.ai Narrows Benchmark Gap With U.S. Rivals</h3><p>Z.ai’s GLM 5.2 is an AI large language model (LLM) with 753 billion parameters, though it has only 40 billion parameters active at once—an optimization that improves the speed at which a model can respond. Z.ai released the model under an <a href="https://opensource.org/license/mit" target="_blank">MIT open-source license</a>, which means anyone can distribute, copy, modify, and use it.</p><p>GLM 5.2’s release added to fears that U.S. AI companies could lose their competitive edge. The model <a href="https://z.ai/blog/glm-5.2" rel="noopener noreferrer" target="_blank">nearly ties Opus 4.8’s score</a> on some agentic coding benchmarks, such as <a href="https://www.frontierswe.com" rel="noopener noreferrer" target="_blank">FrontierSWE</a> and <a href="https://posttrainbench.com" rel="noopener noreferrer" target="_blank">PostTrainBench</a>. <a href="https://www.graphistry.com/blog/glm-5-2-cybersecurity-open-model" rel="noopener noreferrer" target="_blank">Cybersecurity researchers</a> have also found that GLM 5.2 scores well in <a href="https://botsbench.com" rel="noopener noreferrer" target="_blank">cybersecurity benchmarks</a>, a capability that spurred comparisons to <a href="https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/" rel="noopener noreferrer" target="_blank">Anthropic’s Mythos</a>. </p><p>Z.ai arrives amid a broader trend. According to <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_self">Stanford’s AI Index</a> (an annual, 300-plus-page survey of AI trends) Chinese companies produced just over half as many “notable” AI models in 2025 as their U.S. counterparts. That’s up from roughly a third in 2023, and a fifth in 2020.</p><p><a href="https://www.forbes.com/sites/craigsmith/2026/06/28/buckle-up-the-bad-guys-now-have-a-model-as-powerful-as-mythos/" rel="noopener noreferrer" target="_blank">GLM 5.2 caused hand-wringing among some U.S. observers</a> due to its outstanding benchmark scores, which set new records for both open-weights models and Chinese-developed models generally. The model’s Chinese origin also complicates its use for companies in the U.S. and elsewhere that are wary of routing sensitive data through Chinese-linked infrastructure. </p><p>However, the model’s open weights provide an out. Any organization worried about where its data is being sent can instead host the model on its own hardware. This stands in contrast to most frontier-level models, which are gated behind an API with no self-hosting option.</p><p><span>Z.ai backed up GLM 5.2’s release with the company’s own numbers, publishing a </span><a href="https://z.ai/blog/glm-5.2" target="_blank">research report</a> <span>the same day as GLM 5.2’s launch.</span></p><p>The report doesn’t mention Anthropic’s <a href="https://www.anthropic.com/claude/mythos" target="_blank">Mythos</a> or Fable, which were announced but not yet publicly available at the time of its release.The report instead focuses on Anthropic’s <a href="https://www.anthropic.com/news/claude-opus-4-8" target="_blank">Opus 4.8</a> and <a href="https://openai.com/index/introducing-gpt-5-5/" target="_blank">OpenAI’s GPT-5.5</a>. And while GLM 5.2 often performs almost as well as Opus 4.8 in benchmarks, the report claims a win only in two less-difficult reasoning benchmarks—and none in coding.</p><p>Many of the report’s benchmarks place GLM 5.2 behind Opus 4.8 (and, at times, OpenAI’s GPT-5.5) in agentic coding. For example, GLM 5.2 completed just 13 percent of tasks in <a href="https://www.swe-marathon.org" target="_blank">SWE-Marathon</a>, a difficult long-duration agentic coding benchmark. Claude Opus 4.8 doubled GLM 5.2’s score in this benchmark. Opus 4.8 also notched wins of 10 percent or more in the coding benchmarks <a href="https://arxiv.org/abs/2512.12730" rel="noopener noreferrer" target="_blank">NL2Repo</a>, <a href="https://deepswe.datacurve.ai" rel="noopener noreferrer" target="_blank">DeepSWE</a>, and <a href="https://toolathlon.xyz/introduction" rel="noopener noreferrer" target="_blank">Tool-Decathlon</a>.</p><h3>How Do Coders Use GLM 5.2? </h3><p>Software engineers who’ve pitted GLM 5.2 against their own workflows report a wide range of results.</p><p>“The main thing that I realized with [GLM 5.2], was that it can do more long-horizon tasks,” said Hasan, whose company hosts GLM 5.2 on North American infrastructure. Earlier open-weights models, he said, often lost the thread after around 5 to 15 back-and-forth exchanges. “This one, I noticed that I could be using it for hours, and it would still have a coherent train of thought.”</p><p><a href="https://www.linkedin.com/in/nixdavid/" rel="noopener noreferrer" target="_blank">David Nix</a>, a principal software engineer at the Denver-based <a href="https://www.metarouter.io/" rel="noopener noreferrer" target="_blank">MetaRouter</a>, puts LLMs to work at both his day job and for personal side projects. (Nix also operates <a href="https://aiengineerjobs.com/" rel="noopener noreferrer" target="_blank">a jobs board of AI engineers</a>.) Nix said GLM 5.2 comes “really close” to frontier models like Anthropic’s Opus and OpenAI’s GPT-5.5—close enough to earn a permanent spot in his rotation.</p><p>“It’s pretty great at front-end development, for example, where I don’t need to always go to Opus or Fable for those things,” said Nix. He estimates that GLM 5.2 handles 10 to 20 percent of the work he sends to an LLM on a given day, and it’s now his first stop for some specific tasks, such as front-end design. </p><p>Others reported the same strength. Hasan said GLM 5.2 has “really good taste” in web design. <a href="https://www.linkedin.com/in/kacpermichalik/" rel="noopener noreferrer" target="_blank">Kacper Michalik</a>, a software engineer at Kraków, Poland–based <a href="https://screen.studio/" rel="noopener noreferrer" target="_blank">Screen Studio</a>, received good results while using GLM 5.2 to create forms for use on a website.</p><p>On the other hand, <a href="https://www.linkedin.com/in/sai-kiran-myadaram-027893242/" rel="noopener noreferrer" target="_blank">Sai Kiran Myadaram</a>, a software engineer at Bengaluru, India–based <a href="https://www.linkedin.com/company/indhic-ai/" rel="noopener noreferrer" target="_blank">Indhic AI</a>, reports less positive results with Z.ai. He signed up for Z.ai’s subscription plan the week GLM 5.2 launched and found the model burned through its token allotment quickly. “The weekly quota that Z.ai provides has been exhausted for me in less than two to three days,” he said. Michalik, who also accessed Z.ai directly, had no significant issues with the model’s quality but occasionally bumped into rate limits, though in his case he stuck to the free plan.</p><p>In addition to rate limits, Myadaram experienced problems with model hallucinations and overplanning when asked to tackle minor front-end fixes. “It’s messing up my code base,” said Myadaram. He’s since drifted back to OpenAI’s <a href="https://openai.com/codex/" rel="noopener noreferrer" target="_blank">Codex</a>. </p>]]></description><pubDate>Tue, 21 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-coding-assistant-china-anthropic</guid><category>Large-language-models</category><category>China</category><category>Ai-research</category><category>Ai-benchmarks</category><category>Anthropic</category><dc:creator>Matthew S. Smith</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-smartphone-running-a-chinese-ai-application-called-z-ai.jpg?id=67508250&amp;width=980"></media:content></item><item><title>How to Make an Invisible Drone</title><link>https://spectrum.ieee.org/invisible-spinning-drone</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/low-visibility-drone-flying-in-front-of-an-office-plant.jpg?id=67480624&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p><span>There are many words that I would never, ever use to describe a drone. Stealthy. Subtle. Whatever the opposite of obnoxious is. Much of this is because of the giant angry bee sound that drones tend to make, but it’s also the way that they look in flight: With uncannily linear movements and an even less canny ability to hover perfectly still, they tend to draw the eye as affronts to nature.</span></p><p>In a paper presented this week at <a href="https://roboticsconference.org/" target="_blank">Robotics Science and Systems 2026</a> in Sydney, roboticists from Northwestern University, Evanston, Ill., demonstrated a drone called Phantom Twist that is essentially invisible to humans, being an order of magnitude more difficult to see in flight than a typical quadrotor. They accomplished this with the aid of computational design, and while the resulting hardware is, I would argue, also an order of magnitude more of an affront to nature than a typical quadrotor represents, it’s pretty amazing how well it works.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="72b9c290bcfb7c1ef7a4e51a32cb7399" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/5KQ7dKs1dpQ?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">Phantom Twist spins so fast, it’s practically invisible.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p><span>The trick here is easy to see, even if the drone isn’t. By spinning in flight at between 15 and 25 hertz, Phantom Twist takes advantage of humans’ decidedly mediocre visual system to turn a solid spinning object into an opaque smear. Human eyes take some amount of time (typically about 100 milliseconds) to integrate what we see before sending the full scene off to our brains for processing. Moving objects can cause problems for this system, because if the movement is fast enough, our eyes are forced to average that motion across the scene, combining it with whatever is in the background and resulting in a transparent blur. This effect is called persistence of vision. For something that spins like Phantom Twist, that motion blur comes from the drone’s rapid rotation, and it works because most of the drone is cleverly designed to be empty space.</span></p><p>Drones that spin in flight are nothing new—we’ve covered a bunch of them in the past, including <a href="https://www.modlabupenn.org/piccolissimo/" target="_blank">Picolissimo</a> and any number of <a href="https://spectrum.ieee.org/spinning-drone" target="_self">samara</a> <a href="https://spectrum.ieee.org/foldable-monocopter-drone" target="_self">drones</a> inspired by the spinning flight of maple seeds. What makes Phantom Twist unique, and also very odd, is that the design was computationally optimized for low visibility. </p><h2>Controlling how drones like this fly</h2><p>Before we get into that, though, a quick note about how drones like this can even fly controllably, because it’s not at all obvious. With just a single motor and no control surfaces, the only possible control input is through the motor itself, and by pulsing the motor speed up or down at just the right time during each rotation, the drone can translate in any direction. Altitude control comes from changing overall motor thrust, and the drone‘s spinning nature makes it passively stable.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="A minimalist drone made of a few thin rods, wires and a miniature circuit board." class="rm-shortcode" data-rm-shortcode-id="0b6c02f1356af36a7437e22c62bbcbf9" data-rm-shortcode-name="rebelmouse-image" id="b2caf" loading="lazy" src="https://spectrum.ieee.org/media-library/a-minimalist-drone-made-of-a-few-thin-rods-wires-and-a-miniature-circuit-board.jpg?id=67480639&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Carbon fiber rods connect batteries, a controller, some counterweights, and a motor and propeller. The research robot also includes optical tracking tags.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p>The bits that you need for this kind of drone include the motor and propeller, a couple of batteries, a controller, some counterweights (which could be replaced with more batteries or payload), 0.8-mm carbon fiber rods to tie it all together, and a connector for the handheld launcher that gets the whole thing up to speed. The actual <em><em>arrangement</em></em> of these components is surprisingly flexible, and that’s where the invisibility comes in. </p><p>“The design space is high dimensional,” explains Northwestern’s <a href="https://www.mccormick.northwestern.edu/research-faculty/directory/profiles/rubenstein-michael.html" target="_blank">Michael Rubenstein</a>. “It’s very difficult for a human to reason through all the trade-offs between the physical constraints required for stable flight and the visual appearance of the spinning drone, and I don’t think we would have easily arrived at this low-visibility design ourselves.”</p><p>The visibility (or not) of Phantom Twist is primarily driven by the extent to which different components line up with each other from the perspective of someone looking at the drone. The more components that line up with each other as the drone flies, the less background you see through the spinning drone, and the more visible the drone becomes. Because you might be looking at the drone from a number of different angles, and also because the drone has to be stable enough for controlled flight, there are a bunch of different things that need to be optimized all at once, which is why computational design is effective here.</p><p>Phantom Twist’s final design was generated using an iterative optimizer which had a goal of minimizing a metric called <a href="https://eureka.patsnap.com/article/what-is-lpips-and-how-it-measures-perceptual-similarity" target="_blank">learned perceptual image patch similarity</a>, or LPIPS, while making sure that the design could still physically work. LPIPS is the difference between two images: a background image, and a background image with an overlay of the simulated spinning drone. The smaller that difference is, the more invisible that design is. It’s tricky for a human to consider all of the variables at once, but Rubenstein says that the final design does make intuitive sense, because “the automated pipeline prefers placements where components don’t visually overlap as it spins, or where the components are too close to the center of rotation.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Two variations of minimalist drones made from a few thin rods, wires and a miniature circuit board. Both are barely visible when in-flight." class="rm-shortcode" data-rm-shortcode-id="1869e49edd1b81064d2bfb41e0eedecb" data-rm-shortcode-name="rebelmouse-image" id="394e6" loading="lazy" src="https://spectrum.ieee.org/media-library/two-variations-of-minimalist-drones-made-from-a-few-thin-rods-wires-and-a-miniature-circuit-board-both-are-barely-visible-when.jpg?id=67480677&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Two iterations of Phantom Twist drones are shown with their handheld launching mechanisms. The better-optimized version [bottom row] relocates the launcher interface to remove components that are too close to the central axis, making them more visible.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Michael Rubenstein/Northwestern University</small></p><p>Out of a starting set of around 20,000 feasible Phantom Twist configurations, the optimized design (the one that you see or don’t see in the pictures and videos) has a LPIPS score of 0.0104. A human-designed Phantom Twist is about twice as visible, with a LPIPS score of around 0.2, and a conventional quadrotor (of the same size) would be over 10 times more visible. And there’s still a bit more optimization that could be done with the electrical wiring as well as increasing the baseline transparency of the components themselves.</p><p>Phantom Twist is currently controlled using an optical tracking system, which means that it’s not yet capable of flying outside of a controlled environment. But Rubenstein has built <a href="https://www.roboticsproceedings.org/rss20/p102.pdf" target="_blank">other drones along similar principles</a> in the past, which have successfully flown outside, and he’s optimistic about using those techniques to break Phantom Twist out of the lab. The spinning behavior might even enable some useful sensing capabilities, he says. “An interesting possibility is mounting a camera on the spinning body. As the vehicle rotates, it could capture imagery in every direction, effectively creating a 360-degree view of its surroundings that could be used for onboard navigation and control.”</p><p>As for what a drone like Phantom Twist could be used for—assuming that the sound can be mitigated somewhat (and there are <a href="https://spectrum.ieee.org/pennysized-ionocraft-flies-with-no-moving-parts" target="_self">potential approaches to making that happen</a>), a stealthy microdrone could do all sorts of things with covert surveillance being the most obvious application. For his part, Rubenstein says that he’s personally excited about the potential for watching wildlife, “where a less-intrusive drone could observe animals while minimizing its impact on their natural behavior.” The elephants in particular <a href="https://spectrum.ieee.org/research-proves-drones-sound-like-bees-which-is-good-news-for-elephants" target="_self">would certainly appreciate that</a>.</p><p>For a deeper dive into all the particulars of this project, read the paper: <a href="https://arxiv.org/html/2605.11296v1" target="_blank"><em><em>Computational Design of a Low-Visibility UAV Using a Human-Aligned Perceptual Metric</em></em></a>, by Jingxian Wang, Chen Yu, David Matthews, Emma Alexander, Sam Kriegman, and Michael Rubenstein from Northwestern University, which is being presented this week at <a href="https://roboticsconference.org/" target="_blank">RSS 2026 in Sydney</a>.</p>]]></description><pubDate>Thu, 16 Jul 2026 16:09:21 +0000</pubDate><guid>https://spectrum.ieee.org/invisible-spinning-drone</guid><category>Robotics</category><category>Drones</category><category>Spinning-drones</category><category>Invisible-drones</category><category>Ai-design-optimization</category><category>Uav</category><dc:creator>Evan Ackerman</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/low-visibility-drone-flying-in-front-of-an-office-plant.jpg?id=67480624&amp;width=980"></media:content></item><item><title>Digital Surveillance Reshapes Fishery Enforcement in Indonesia</title><link>https://spectrum.ieee.org/fishery-satellite-surveillance</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/an-overhead-drone-photo-shows-an-array-of-closely-placed-fishing-boats-at-sea.jpg?id=67101579&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>In the eastern Indian Ocean, south of Java in the vast sea stretching toward Australia, a fishing vessel slightly alters its course while operating near the boundary of its authorized fishing ground. Nothing appears unusual on deck. Nets remain in the water. Engines maintain a steady speed. To the crew, it is an ordinary day at sea.</p><p>Yet hundreds of kilometers above, satellites continuously record the vessel’s position. At Indonesia’s Marine and Fisheries Resources Surveillance Station, in Cilacap, where I work, a monitoring platform receives the signal and automatically compares it against fishing permits, designated fishing grounds, vessel characteristics, and historical movement patterns. Within minutes, the system identifies a potential violation. Before any patrol vessel leaves port, before any inspector boards a vessel, and before any warning is issued, we have begun enforcement.</p><p>This transformation reflects a profound shift in maritime governance. The ocean has historically been opaque to regulators. States could only enforce laws where patrol vessels happened to be present. Today, however, integrated systems combining data from vessel monitoring systems (VMS), <a href="https://spectrum.ieee.org/earth-observation-satellites-small-constellations" target="_blank">satellite remote sensing</a>, geospatial analytics, and increasingly sophisticated data-processing tools are making marine activity visible at an <a href="https://www.science.org/doi/10.1126/science.aao5646" rel="noopener noreferrer" target="_blank">unprecedented scale</a>. Global Fishing Watch alone tracks <a href="https://globalfishingwatch.org/research/annual-report-2025/" rel="noopener noreferrer" target="_blank">hundreds of thousands of vessels</a> worldwide, generating a near real-time picture of fishing activity across the world’s oceans.</p><p>Indonesia has emerged as one of the most ambitious examples of this transition. As the world’s largest archipelagic state, managing more than 6 million square kilometers of maritime space, Indonesia faces a challenge familiar to many coastal nations: There are never enough patrol vessels. Digital surveillance is a practical necessity that makes my job possible, even as it creates new challenges.</p><h2>The Law of the Sea Meets Digital Reality</h2><p>The international legal framework governing the oceans was designed in an era when maritime enforcement depended almost entirely on physical presence. The <a href="https://www.un.org/depts/los/convention_agreements/texts/unclos/unclos_e.pdf" rel="noopener noreferrer" target="_blank">United Nations Convention on the Law of the Sea (UNCLOS), adopted in 1982,</a> assumes that states exercise authority through patrols, inspections, vessel boardings, and direct observation.</p><p>For countries with extensive coastlines and limited enforcement resources, this model has always faced practical constraints. Indonesia’s Fisheries Management Areas (WPP-NRI) span waters ranging from the Indian Ocean to the Pacific and from the Strait of Malacca to the maritime boundaries adjacent to Australia and Papua New Guinea. Monitoring such a vast domain solely through patrol operations is both expensive and operationally impossible.</p><p>Beginning in the late 2010s, Indonesia accelerated the integration of satellite-based monitoring into fisheries enforcement. Vessel monitoring systems became a cornerstone of this strategy. By early 2026, a total of <a href="https://ppid.kkp.go.id/upt/pelabuhan-perikanan-samudera-kendari/news/detail/9394-kapal-perikanan-sudah-pasang-vms/" rel="noopener noreferrer" target="_blank">9,394 Indonesian fishing vessels</a> were actively transmitting through the national VMS, representing an increase of 2,880 vessels during the 2021–2025 period. As part of Indonesia’s broader maritime surveillance architecture, VMS data are complemented by satellite remote sensing and other monitoring tools to help identify suspicious activities involving vessels operating without active transponders or outside the national VMS network.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A man in an Indonesian military uniform points to a large monitor showing an overview of thousands of fishing boats at sea near Indonesia." class="rm-shortcode" data-rm-shortcode-id="b9c4ba7e404292b97184b48952a729cc" data-rm-shortcode-name="rebelmouse-image" id="1d12e" loading="lazy" src="https://spectrum.ieee.org/media-library/a-man-in-an-indonesian-military-uniform-points-to-a-large-monitor-showing-an-overview-of-thousands-of-fishing-boats-at-sea-near.jpg?id=67101591&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Indonesian fisheries officials plan fishery patrols using data from tracking devices, satellites, and their understanding of the patterns of illegal fishing.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Indonesian Ministry of Marine Affairs and Fisheries</small></p><p>The implications extend far beyond vessel tracking. Continuous digital monitoring enables authorities to reconstruct vessel movements, identify suspicious behavioral patterns, detect unauthorized fishing activity, and verify compliance with licensing conditions. Rather than waiting to discover violations during patrol operations, regulators can increasingly prioritize inspections based on data-derived risk assessments.</p><p>Maritime governance is shifting from reactive enforcement toward predictive oversight.</p><h2>The Surprising Geography of Digital Enforcement</h2><p>The expansion of surveillance infrastructure has already generated measurable enforcement outcomes.</p><p>The Ministry of Marine and Fisheries Affairs Indonesia <a href="https://kkp.go.id/unit-kerja/djpsdkp/akuntabilitas-kinerja/pelaporan-kinerja/detail/laporan-kinerja-direktorat-jenderal-psdkp-tahun-20256a042458a936b.html" target="_blank">imposed 2,550 administrative sanctions during 2025</a>, many involving violations detected through the vessel monitoring system, including fishing outside authorized fishing grounds and deliberate deactivation of monitoring transmitters.</p><p>This statistic is significant because many of these violations would have been extremely difficult to detect under traditional patrol-based enforcement. A vessel that briefly crosses into a prohibited fishing zone may never encounter an enforcement vessel. Likewise, a captain who temporarily disables a transmitter may escape detection if oversight depends solely on physical inspections.</p><p>Digital monitoring fundamentally changes this equation. Every vessel movement creates a data trail. Authorities can reconstruct routes, identify anomalous behavior, and compare activities against permit conditions long after the event itself has occurred.</p><p>The first quarter of 2026 demonstrates the scale of this surveillance capability. During just three months, Indonesia’s fisheries monitoring system tracked 14,571 fishing vessels, 182 fishing gear units, and 208 registered home ports while identifying <a href="https://kkp.go.id/download-pdf-akuntabilitas-kinerja/akuntabilitas-kinerja-pelaporan-kinerja-laporan-kinerja-direktorat-jenderal-psdkp-twi-tahun-2026.pdf" target="_blank">491 suspected violations</a> across the country’s fisheries management areas. These violations included unauthorized fishing grounds, illegal high-seas operations, transshipment-related offenses, port-base discrepancies, licensing irregularities, and indications of poaching.</p><p>Such numbers reveal a fundamental transformation. Enforcement is no longer limited by the number of patrol vessels available at sea. Instead, surveillance capacity increasingly depends on the ability to collect, process, and interpret big data.</p><h2>Illegal Operators Are Learning Too</h2><p>Yet greater visibility does not eliminate illegal fishing. But it does change how poachers operate.</p><p>Indonesia’s expanding digital surveillance network, and a 2023 requirement that even small vessels use VMS when 12 nautical miles offshore, appears to have improved compliance among licensed fishing vessels. However, as enforcement capabilities become more sophisticated, some actors engaged in illegal fishing have also become more adept at exploiting technological and operational gaps.</p><p>Deliberately disabling VMS transmitters remains one of the most common enforcement concerns. While temporary signal losses, whether intentional or caused by technical failures—can complicate the reconstruction of vessel movements, they do not necessarily prevent authorities from detecting potentially illegal activity. Indonesia increasingly combines VMS with satellite-based observations, other maritime surveillance systems, intelligence-led analysis, and reports from <a href="https://repository.seafdec.org/bitstream/handle/20.500.12066/7649/Fish-surveillance-Indonesia.pdf?sequence=1&isAllowed=y" target="_blank">community-based surveillance groups</a> (<em>Pokmaswas</em>) to corroborate suspicious behavior and direct patrol resources where they are most needed. This layered approach—integrating digital technologies with local knowledge from coastal communities—helps reduce opportunities for illegal, unreported, and unregulated (IUU) fishing even when a single monitoring system is compromised.</p><p class="pull-quote">A compromised surveillance network could potentially disrupt enforcement operations just as effectively as a vessel evading patrol detection.</p><p>As digital surveillance expands, one lesson from Indonesia’s experience is that stronger monitoring does not eliminate illegal fishing—it changes how illegal operators behave. Improved compliance across much of the fishing fleet has been accompanied by increasingly sophisticated attempts by a smaller group of offenders to avoid detection. This reflects a broader reality of technology-enabled enforcement: As monitoring capabilities evolve, so do the strategies used to circumvent them.</p><p>The result is a technological arms race. Every improvement in surveillance capability encourages <a href="https://www.science.org/doi/10.1126/science.aad5686" target="_blank">new methods of avoidance</a>, whether through disabling tracking devices, manipulating vessel identities, or exploiting gaps between different monitoring systems. Enforcement agencies must therefore continuously refine their analytical methods, integrate multiple sources of maritime information, and adapt their operational strategies to keep pace with evolving behavior at sea. Effective digital fisheries governance is not defined by a single technology but by the ability to combine data, human expertise, and operational intelligence into a resilient and adaptive enforcement system.</p><h2>The Next Battle May Be Over Data Integrity</h2><p>The future of fisheries enforcement may ultimately depend less on detecting vessels and more on ensuring confidence in the digital systems that generate enforcement decisions.</p><p>As surveillance networks become increasingly integrated, questions surrounding cybersecurity, algorithmic accountability, and data integrity become more important. What happens if vessel tracking data are manipulated? How should authorities verify automated risk assessments? What safeguards exist when enforcement actions increasingly originate from algorithmic analysis rather than direct human observation?</p><p>These questions are no longer theoretical.</p><p>Modern fisheries governance increasingly depends on interconnected networks of satellites, communication systems, databases, cloud infrastructure, and analytical platforms. While these technologies dramatically improve visibility, they also create new vulnerabilities. A compromised surveillance network could potentially disrupt enforcement operations just as effectively as a vessel evading patrol detection.</p><p>For Indonesia, this means that investment in digital surveillance must be accompanied by investment in digital resilience. The effectiveness of a monitoring system ultimately depends not only on the volume of data collected but also on the credibility, security, and <a href="https://spectrum.ieee.org/data-integrity" target="_blank">reliability of the information produced</a>.</p><h2>Governing Oceans Through Data</h2><p>Indonesia’s experience illustrates a broader global transformation in maritime governance. The ocean is becoming increasingly transparent to regulators. Activities that once occurred beyond the reach of enforcement agencies can now be observed, analyzed, and investigated through interconnected digital systems.</p><p>The benefits are substantial. Expanded VMS adoption, improved monitoring coverage, and thousands of administrative enforcement actions demonstrate that digital surveillance can significantly enhance fisheries governance. Yet the transition also introduces new challenges involving data quality, cybersecurity, algorithmic accountability, and adaptive <a href="https://doi.org/10.3389/fmars.2018.00240" rel="noopener noreferrer" target="_blank">criminal behavior</a>.</p><p>The central question facing maritime regulators is how governments can ensure that increasingly powerful monitoring systems remain transparent, secure, and accountable while preserving public trust and legal legitimacy. The most important lesson may be that digital surveillance does not replace traditional enforcement. It changes where enforcement begins. For generations, maritime law enforcement started when a patrol vessel encountered a suspected violator. Today, it often starts when an algorithm detects a pattern.</p><p>That shift may prove as significant for ocean governance as the invention of radar was for maritime navigation.</p>]]></description><pubDate>Thu, 16 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/fishery-satellite-surveillance</guid><category>Fishing</category><category>Environmental-monitoring</category><category>Indonesia</category><category>Poaching</category><category>Surveillance</category><dc:creator>Yogi Putranto</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/an-overhead-drone-photo-shows-an-array-of-closely-placed-fishing-boats-at-sea.jpg?id=67101579&amp;width=980"></media:content></item><item><title>The First Chatbot’s Multiple Personalities</title><link>https://spectrum.ieee.org/eliza-chatbot-source-code</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/photo-collage-of-a-charismatic-elderly-man-raising-his-eyebrows-against-a-background-of-obscured-computer-programming-script.jpg?id=67155632&width=1245&height=700&coordinates=0%2C0%2C0%2C0"/><br/><br/><p class="intro-text"><em>ELIZA is remembered as the world’s first AI star, a kindly therapist in chatbot form that gently probed users’ worries. Even its creator, Joseph Weizenbaum, was surprised by the warm reception given to his experiment in human-machine interaction. For some, it heralded an age of automated psychotherapy, while others believed the program demonstrated sentience, a fallacy soon known as the “<a href="https://spectrum.ieee.org/why-people-demanded-privacy-to-confide-in-the-worlds-first-chatbot" target="_self">ELIZA effect</a>.” Based on published descriptions, ELIZA has been implemented on many different computers, but only recently has <a href="https://sites.google.com/view/elizaarchaeology/code" rel="noopener noreferrer" target="_blank">the actual source code</a> been unearthed from <a href="https://archivesspace.mit.edu/" rel="noopener noreferrer" target="_blank">MIT’s archives</a>. </em></p><p><em>In <a href="https://mitpress.mit.edu/9780262052481/inventing-eliza/" target="_blank">Inventing ELIZA: How the First Chatbot Shaped the Future of AI</a>, just published by <a href="https://mitpress.mit.edu/" rel="noopener noreferrer" target="_blank">MIT Press</a>, a squad of researchers analyze the code and reveal a complex program capable of much more than faking psychiatry. In fact, it could assume several different personas. The authors have also created <a href="https://sites.google.com/view/elizaarchaeology/try-eliza" rel="noopener noreferrer" target="_blank">a faithful emulation of the therapist persona that you can try yourself</a></em><span><em> after reading the book excerpt below.</em></span></p><p class="drop-caps">W<strong>hen it debuted in</strong> the mid-1960s, the ELIZA software program transformed the way people thought about interacting with computers. As the first chatbot, ELIZA demonstrated how a calculation machine might engage in conversation, ushering in a host of social and technical questions that still resonate today. Now we don’t think twice about interacting with a machine in real time, conversing over text, or even speaking into the air to ask about the weather. In many ways, ELIZA shaped not only the way we think about <em><em>interacting</em></em> with computers but also how we think <em><em>about</em></em> them. It began to give a reality to the <a href="https://spectrum.ieee.org/tag/science-fiction" target="_blank">science fiction</a> stories of how we expect computers to work.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Orange book cover titled \u201cInventing Eliza: How the First Chatbot Shaped the Future of AI\u201d" class="rm-shortcode" data-rm-shortcode-id="208d351be5eb0018c13ad5a993fa22b5" data-rm-shortcode-name="rebelmouse-image" id="03894" loading="lazy" src="https://spectrum.ieee.org/media-library/orange-book-cover-titled-u201cinventing-eliza-how-the-first-chatbot-shaped-the-future-of-ai-u201d.jpg?id=67155786&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">This article is adapted from the new book “Inventing ELIZA: <a href="https://mitpress.mit.edu/9780262052481/inventing-eliza/" target="_blank">How the First Chatbot Shaped the Future of AI</a>“ (MIT Press, 2026).</small></p><p>Although ELIZA was far from a faultless conversation partner, it astonished its users. The recent discovery and archaeology of the original ELIZA source code represents a significant intervention in the history of computing. By examining the actual implementation of ELIZA rather than relying on later reconstructions and reimplementations, we challenge taken-for-granted assumptions about this key software artifact.</p><p>For example, the source code reveals that ELIZA was not merely a simple pattern-matching chatbot but can be better understood as a sophisticated platform designed for multiple “personas,” or scripts, with a complex set of capabilities, including script editing and contextual memory. The script that most people conflate with the program ELIZA was actually called Doctor, which performed the role of a psychotherapist. Yet, like a modern chatbot prompted to behave with different personalities, ELIZA could take on many roles.</p><p class="pull-quote">“This code and script…reveal underlying assumptions about language, therapy, and human-computer interaction that continue to influence modern AI development.”</p><p>This unearthed material transforms our understanding of early AI development by demonstrating that Joseph Weizenbaum’s technical innovations were far more advanced than previously documented. Moreover, the discrepancies between his published descriptions and the actual implementation help to show the gap between theoretical computational models and their material instantiations in computer source code, a tension that continues to shape digital culture today.</p><p>Although many technical innovations have emerged in the decades since ELIZA, examining the ELIZA/Doctor code offers a rare glimpse into one of the earliest formalized attempts to model human conversation. What makes ELIZA particularly fascinating is not only its historical significance but also what it reveals about Weizenbaum’s views on both computing and human interaction. This code and script do not merely showcase programming techniques of the 1960s; they reveal underlying assumptions about language, therapy, and human-computer interaction that continue to influence modern AI development. By examining this code, we can start to uncover the sophisticated linguistic and programming techniques that allowed a rudimentary pattern-matching system to create a convincing simulation of understanding. But before we can read the lines of code, let us offer an overview of the system.</p><h2>How Did ELIZA Create Personas?</h2><p>The architectural distinction between ELIZA and Doctor represents an important design decision in AI history. Think of ELIZA as a system for interaction and Doctor as one set of rules that Weizenbaum devised, among others. This separation, manifested in ELIZA’s system-script dichotomy, presaged numerous contemporary software patterns, from configuration-as-data to plug-in architectures and domain-specific languages.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A 1960s chatbot program running on a 1980s IBM personal computer." class="rm-shortcode" data-rm-shortcode-id="25a3713ea68f943c1be636d8e3cdbc0d" data-rm-shortcode-name="rebelmouse-image" id="1b809" loading="lazy" src="https://spectrum.ieee.org/media-library/a-1960s-chatbot-program-running-on-a-1980s-ibm-personal-computer.jpg?id=67155917&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Based on published journal articles, ELIZA was re-created on many platforms, such as the IBM PC. However, the actual source code sat untouched in the MIT archives for many years. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VCF Museum at InfoAge</small></p><p>Without question, the historical context of 1960s computing fundamentally shaped ELIZA’s architecture as well. Decisions in computing that reflect material constraints create path dependencies and eventually become programming cultural norms. These constraints manifested in ELIZA’s single-pass processing, tape-based storage and stack-oriented implementation. Yet within these limitations, Weizenbaum crafted an elegant solution. These technical features, though invisible to the users, are crucial to creating the illusion of understanding that made ELIZA so compelling.</p><p>Weizenbaum explained many of ELIZA’s technical features in the 10-page paper published in the <a href="https://dl.acm.org/doi/10.1145/365153.365168" target="_blank">January 1966 edition of the journal <em><em>Communications of the Association for Computing Machinery</em></em></a> (<em><em>CACM</em></em>). But he chose to omit some essential details.</p><p>In that paper Weizenbaum published ELIZA’s best known dialogue, which begins,</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>Men are all alike.</p><p>IN WHAT WAY</p><p>They’re always bugging us about something or other.</p><p>CAN YOU THINK OF A SPECIFIC EXAMPLE</p><p>Well, my boyfriend made me come here.</p></div><p>This dialogue marked ELIZA’s public debut in 1966 as one of the examples produced by the Doctor script. By finding the source code for ELIZA and examining how it performs the Doctor script, we now better understand these two separate parts of a system and can explore the many other personas of ELIZA. In just some of the other scripts known to date, ELIZA was programmed to discuss math, <a href="https://spectrum.ieee.org/tag/poetry" target="_blank">poetry</a>, color, paradoxes, synchronization, relativity, France, and elevators.</p><p>These scripts work like templates. They are structured data that direct the ELIZA system to “play” a particular task or role. By comparing archival and published ELIZA dialogues from interactions with a variety of scripts, including Doctor, we can understand more about bot personas and how they function, paying close attention to how a bot evokes social dynamics between system and interactor.</p><p>Ultimately, studying the dialogues and scripts demonstrates the crucial role that collaboration plays in these exchanges, as bot and user cocreate the sense of their interaction. To understand the full range of ELIZA’s capabilities and conversational possibilities, let’s take a look at the variety of scripts that were created for the ELIZA system.</p><p>What distinguishes each ELIZA script is both its subject matter and the linguistic and stylistic choices used to deliver that content. These choices are not neutral; they can be said to construct a particular persona with characteristics that emerge through the script’s language patterns, vocabulary, and conversational approach. In short, it matters not just what you say but how you say it too.</p><p class="pull-quote">“The aim was less to create a functional automated therapist and more to find a suitably constrained role to match the limitations of the programming environment.”</p><p>For example, with the Doctor script Weizenbaum deliberately echoed the style of a Rogerian “talk” therapist. He chose this persona because the psychiatric mode is one of the few types of conversations in which one person can “assume the pose of knowing almost nothing of the real world. If, for example, one were to tell a psychiatrist ‘I went for a long boat ride’ and he responded, ‘Tell me about boats,’ one would not assume that he knew nothing about boats but that he had some purpose in so directing the subsequent conversation.”</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of paper loaded in a teletype machine with a few paragraphs of chatbot dialogue written on it. " class="rm-shortcode" data-rm-shortcode-id="86bbdc58c1e01df11061af97b1e9b45c" data-rm-shortcode-name="rebelmouse-image" id="97e5e" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-paper-loaded-in-a-teletype-machine-with-a-few-paragraphs-of-chatbot-dialogue-written-on-it.jpg?id=67161217&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The first users of ELIZA interacted with it via teletype terminals.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">VCF Museum at InfoAge</small></p><p>Thus, the most famous persona created for ELIZA was a technical convenience. As human-computer interaction expert Lucy Suchman explains, “The Doctor program exploited the maxim that shared premises can remain unspoken: that the less we say in conversation, the more what is said is assumed to be self-evident.” In creating the original ELIZA effect, less was more.</p><p>The aim was less to create a functional automated therapist and more to find a suitably constrained role to match the limitations of the programming environment. Then Weizenbaum composed the script to match the role by choosing specific words that evoked rhetorical tone and characterization, for example, <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">LET’S DISCUSS FURTHER WHY YOU … WHAT DOES THAT SUGGEST TO YOU.</span> In Doctor, the machine side of the conversation needs to appear like a good listener who cares about what the user has mentioned before, so it often includes the user’s text in its replies and keeps its responses open-ended. Because a real doctor would be inquisitive, the script contains lots of<span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">WHAT</span> and<span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">WHY</span> questions. In other scripts and dialogues, the script keywords and assigned responses reveal the design decisions that help create different specific personas. This variation becomes increasingly apparent as we look at the wider range of ELIZA scripts.</p><h2>What Were Some of ELIZA’s Alternate Personas?</h2><p>Let’s start with a script called Neweng, short for New England. By the mid-1950s, scientists were already taking up Turing’s proposition for conversing with a computer as a tangible idea. In 1959, a “conversation machine” made by a small team at the University of Toronto Computation Centre could make small talk about the weather if preprompted with current data. Its script included a ranked series of time words, weather words, and emotion words, but the system had “a rather narrow outlook,” limited vocabulary, and no knowledge of syntax. According to its creators, “in spite of this crudity, the results in passable conversation are surprisingly good, which in a way is a commentary on the shallowness of ordinary conversation about the weather.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Typewritten list of inquisitive dialogue responses such as \u201cWhat does that suggest to you?\u201d" class="rm-shortcode" data-rm-shortcode-id="5fcf41cfad8ee21b240d837e7b66c660" data-rm-shortcode-name="rebelmouse-image" id="17794" loading="lazy" src="https://spectrum.ieee.org/media-library/typewritten-list-of-inquisitive-dialogue-responses-such-as-u201cwhat-does-that-suggest-to-you-u201d.jpg?id=67156058&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">This fragment of ELIZA’s source code shows how scripted responses were encoded, in this case to make ELIZA sound like a Rogerian therapist. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MIT Libraries, Distinctive Collections</small></p><p>One of ELIZA’s test scripts found in the archive had a similar function. Neweng does not soothe, therapize, or instruct, but only makes small talk. It chats as if at a boring dinner party.</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>WHAT NEW ENGLAND STATE DO YOU LIKE BEST OF ALL</p><p>i love rhode island</p><p>I ENJOY THE BEACHES ON LONG ISLAND SOUND</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like massachusetts</p><p>THAT IS MY HOME STATE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>you must like massachusetts too</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i enjoy skiing in new hampshire</p><p>PLEASANTLY COOL IN THE SUMMER UP THERE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like missouri</p><p>MISSOURI IS NOT A NEW ENGLAND STATE</p><p>WHAT OTHER NEW ENGLAND STATE DO YOU LIKE</p><p>i like cape cod</p><p>THAT’S A BEAUTIFUL PART OF MASSACHUSETTS</p></div><p>What can we learn from such a basic conversation in which it seems like nothing interesting is shared? Despite the bland chat, a persona emerges from the bot script in phrases like <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">PLEASANTLY COOL</span> and <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">UP THERE</span> and in locations like <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">THE BEACHES ON LONG ISLAND SOUND</span>. From such comments, we can begin to imagine a persona who has not only a certain vocabulary but also certain hobbies and vacation spots. Perhaps implicit assumptions of race and class emerge as well.</p><p>A chatty persona might take on many forms, depending on where and when the conversation occurs. This one seems reserved, perhaps fitting its setting in 1960s New England. The system reminds the user that Missouri is not a New England state, but what if this conversation took place in Missouri, Texas, or Mexico? The machine persona would sound different in its cadence, tone, and references. What would we come to understand about a chat persona from Fire Island, from Brooklyn, from Berlin? What would they sound like, and what topics would they discuss?</p><p>These differences in subject matter do matter. They imply personas with entirely different backgrounds and experience, giving users wholly different interactions and affective relations. In this way, the Neweng script demonstrates how even simple algorithms making contextual responses about geography could generate a convincing sense of personhood and place. Whereas Neweng could be said to have created a casual, conversational persona focused on light social exchange, other scripts pushed ELIZA into more structured and educational roles. These scripts demonstrate how the system could be adapted not just for friendly chatter but for teaching.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Black-and-white portrait of a middle-aged balding man in aviator eyeglasses." class="rm-shortcode" data-rm-shortcode-id="a02ade38f8966ac94c795b313e1c3a6b" data-rm-shortcode-name="rebelmouse-image" id="fc521" loading="lazy" src="https://spectrum.ieee.org/media-library/black-and-white-portrait-of-a-middle-aged-balding-man-in-aviator-eyeglasses.jpg?id=67155997&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">Edwin Taylor, at MIT’s Education Research Center, developed alternate scripts for ELIZA, testing its ability to act as a teacher.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MIT Libraries, Distinctive Collections</small></p><p>Meet ELIZA the tutor, quite unlike ELIZA the therapist or the chatty neighbor. Intrvw, Canvec, FVP1, and Arithm are a set of ELIZA scripts created as teaching tools used in experiments by Edwin F. Taylor at MIT’s Education Research Center. These scripts run on later versions of ELIZA that incorporated an important technical innovation called conditional keyword matching.</p><p>Unlike the original ELIZA, which simply looked for keywords and generated responses based on their presence, these updated versions could track what had been discussed previously and branch into different conversational paths based on specific user answers. This development allowed ELIZA to simulate a kind of Socratic method, where a tutor guides learning through carefully sequenced questions that respond to student answers rather than simply presenting information.</p><p>These scripts construct the tutor persona through many subtle linguistic gestures that create characterization and rhetorical tone. This tone differs from that of Doctor, which asks open-ended questions and comes across as gentle and nonscientific. In the tutoring scripts, large blocks of informative text from the bot tend to dominate the conversation, and the tone is often more dry and unemotional in these explanations. The dialogues indicate structured scripts that include guidance to lead the student through narrow, Socratic learning paths.</p><p>In particular, the teaching scripts feature praise and critique. The dialogues for Intrvw, Canvec, and FVP1 are peppered with <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">EXCELLENT, VERY GOOD, RIGHT YOU ARE,</span> and <span style="font-family: 'Courier New', Courier, monospace; font-size: 22px; color: black; padding: 5px; line-height: 1.5;">CONGRATULATIONS.</span> These create the sense of a supportive instructor cheering the student on. Such politeness has been taken up in contemporary bots like ChatGPT, which has been shown to perform better when people are polite back to it.</p><p>ELIZA could become a tutor more effectively as the system grew in its capabilities, another valuable reminder that ELIZA was not one program but a family of programs. After the publication of the 1966 <em><em>CACM</em></em> article, Weizenbaum continued to develop the systems for interaction and understanding. As an experiment, Weizenbaum wrote the Arithm script less as a tutor and more so to “to illustrate the power of the evaluator to which ELIZA has access.” It uses a friendly, plain language interface to let users do simple programming. The script can do calculations, assign variables to values, and perform operations on them. Math problems can be described in sentence form:</p><div style="font-family: 'Courier New', Courier, monospace; font-size: 18px; color: black; padding: 10px; line-height: 1.5;"><p>The radius of a globe is 10.</p><p>A globe is a sphere. A sphere is an object.</p><p>What is the area of the globe.</p><p>IT’S 1256.635916</p></div><p>The updated 1967 version of the ELIZA system can accumulate facts and store additional information. In this later version of ELIZA, when the system does not recognize information, it asks follow-up questions to gain data. As Weizenbaum explains, “The present script is designed to reveal, as opposed to conceal, lack of understanding and misunderstanding. Notice, for example, that when the program is asked to compute the area of the ball, it doesn’t yet know that a ball is a sphere and that when the diameter of the ball needs to be computed the fact that a ball is an object has also not yet been established.” Unlike Doctor, which asks questions to keep the conversation going, Arithm is building its store of, if not knowledge, then data and logic statements.</p><p>Although the variety of scripts helps us to see how a range of personas could be constructed through script programming ELIZA, they represent only half of the conversational process. A script can establish a foundation for a persona, but that persona only emerges fully through interaction with users who engage with it, interpret it, and respond to it in ways that may confirm, challenge, or transform the script’s implicit character. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Wed, 15 Jul 2026 15:35:53 +0000</pubDate><guid>https://spectrum.ieee.org/eliza-chatbot-source-code</guid><category>Chatbots</category><category>Ai</category><category>History</category><category>Eliza</category><category>Joseph-weizenbaum</category><dc:creator>Sarah Ciston</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/photo-collage-of-a-charismatic-elderly-man-raising-his-eyebrows-against-a-background-of-obscured-computer-programming-script.jpg?id=67155632&amp;width=980"></media:content></item><item><title>This AI Folds DNA Into Mini Masterpieces</title><link>https://spectrum.ieee.org/ai-dna-origami</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/3d-renderings-of-unique-dna-structures-resembling-a-flower-cone-and-corkscrew-shaped-loop.jpg?id=67155159&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p><span>Shaped like dogs, stars, and the Mona Lisa, you could mistake these DNA structures for fun-shaped macaroni if they weren’t only nanometers wide. South Korean scientists made the constructions using a technique called </span><a href="https://spectrum.ieee.org/nanoscale-interconnects-come-to-selfassembling-dna-origami" target="_self">DNA origami</a>,<span> which can bend genetic material into any form. Designing DNA strands so they’ll fold into a specific shape typically requires tedious manual work, but the researchers behind the playful fabrications have developed a shortcut using generative AI.</span></p><p>The AI model, called Generative SNUPI (short for Structured Nucleic Acids Programming Interface, and, yes, inspired by the dog), was created by research teams at <a href="https://en.snu.ac.kr/" target="_blank">Seoul National University</a> (SNU) and <a href="https://www.hanyang.ac.kr/web/eng" target="_blank">Hanyang University</a>. The work behind it, which was accepted for publication in <a href="https://www.nature.com/articles/s41467-026-73578-z" target="_blank"><em><em>Nature Communications</em></em></a><em><em>, </em></em>shows the model can conjure DNA origami designs that work in the real world for user-requested shapes. For a design like the Mona Lisa, that doesn’t mean simply tracing an outline; the model considers the chemical rules of DNA to tell researchers how unpaired DNA strands should be sequenced so that molecular forces will cause them to self-contort into the required shape.</p><p>DNA origami techniques have been around for <a href="https://www.nature.com/articles/nature04586" rel="noopener noreferrer" target="_blank">two decades now</a>, with potential applications ranging from nanoscale robots to therapeutic structures that interact with cells. But these innovations have been slowed by how time-consuming and expensive the DNA structure design process can be.</p><p>“Traditionally, we need some expertise, background knowledge, and know-how to design the proper nanostructures that we intend to make,” says <a href="https://www.linkedin.com/in/kyounghwa-jeon-7511041a3/" rel="noopener noreferrer" target="_blank">Kyounghwa Jeon</a>, a Ph.D. candidate at SNU. The work requires humans running algorithms and tweaking results until the desired shape is achieved and structurally stable. With Generative SNUPI, she says, users could, in theory, go straight from drawing a target shape to physically assembling the DNA. </p><p><a href="https://www.linkedin.com/in/rebecca-taylor-ph-d-022b854b/" rel="noopener noreferrer" target="_blank">Rebecca Taylor</a>, a professor of mechanical engineering at <a href="https://www.cmu.edu/" rel="noopener noreferrer" target="_blank">Carnegie Mellon University</a> who was not involved in the research, says the new generative platform is exciting for researchers. “The entire field is sort of enabled and held back by its tools. When you make a new tool that enables a new tech, a new capability, that’s just such a big advance for the field.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Traced outline of a Pug, surrounded by several dozen microscopic DNA structures bearing its exact shape. " class="rm-shortcode" data-rm-shortcode-id="62a7ef7bfdbc368bdc89ba8655839517" data-rm-shortcode-name="rebelmouse-image" id="e7d51" loading="lazy" src="https://spectrum.ieee.org/media-library/traced-outline-of-a-pug-surrounded-by-several-dozen-microscopic-dna-structures-bearing-its-exact-shape.jpg?id=67155167&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Generative SNUPI designs DNA sequences that, when synthesized, fold into nanoscale replicas of user-requested shapes.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Source images: <a href="https://www.nature.com/articles/s41467-026-73578-z" target="_blank">Chien Truong-Quoc, Kyounghwa Jeon, et al.</a></small></p><h2>How AI can design DNA origami</h2><p>Designing DNA origami using Generative SNUPI begins with a target shape. That could be something with complex curvature, like the outline of a dog’s face, or a more simple geometric pattern. Next, the new tech comes into play: Generative SNUPI applies a diffusion model, which adds and refines noise to the input shape to create the desired output in DNA form. Diffusion models are how platforms like <a href="https://openai.com/index/dall-e-3/" target="_blank">DALL-E</a> and <a href="https://www.midjourney.com/explore?tab=top" target="_blank">Midjourney</a> create <a href="https://spectrum.ieee.org/thermodynamic-computing-for-ai" target="_self">AI-generated imagery</a>.</p><p>“What it looks like is one of those kids crafts, where you decorate something with glue and then put glitter all over it,” says Taylor. When the noise is removed—or the glitter is shaken off—the design is revealed. “They’re basically just saying ‘populate this guide that I have with the DNA,’ but they also know how DNA comes together. … That’s the thing that it’s really been trained on.”</p><p>The arts-and-crafts metaphors only continue once Generative SNUPI returns the DNA sequences that form the target shape. Scientists chemically synthesize short DNA strands called staples and used biological methods to produce a long strand called a scaffold. The staples pull the scaffold into shape in a way that Jeon says is “very similar to stapling paper.” The staple-scaffold relationship exploits DNA’s imperative to bond guanine to cytosine and adenine to thymine; the exact positions of each of these molecules are dictated by Generative SNUPI during the design process. </p><p>Researchers were able to produce a variety of DNA origami structures, but some did not hold their shape at first, notes <a href="https://www.linkedin.com/in/do-nyun-kim-4b4830118/" rel="noopener noreferrer" target="_blank">Do-Nyun Kim</a>, an assistant professor of mechanical engineering at SNU. “This occurred not because Generative SNUPI had an error, but because the drawn shape was, in fact, structurally unstable,” he says. In response, they added a step before actually designing the DNA sequence to predict the structural integrity of the input shape. </p><p>To expand Generative SNUPI’s capacity for real-world applications, Kim says that DNA origami designs will need to be less rigid than what the model is currently able to produce. The technology reaching its full potential could mean life-saving uses like drug delivery and immunotherapy, but these uses often require flexibility.</p><p>“Most molecular structures are dynamic and reconfigure in response to external stimuli to perform their designated functions,” he says. “So, we plan to extend the current work to the design of dynamically reconfigurable structures in future research.”</p>]]></description><pubDate>Wed, 15 Jul 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-dna-origami</guid><category>Biotechnology</category><category>Dna-origami</category><category>Dna</category><category>Generative-ai</category><dc:creator>Alex Music</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/3d-renderings-of-unique-dna-structures-resembling-a-flower-cone-and-corkscrew-shaped-loop.jpg?id=67155159&amp;width=980"></media:content></item><item><title>How I Turned AI to the Dark Side</title><link>https://spectrum.ieee.org/jailbreaking-llms</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/glossy-red-robot-devil-standing-on-a-bundle-of-dynamite-against-blue-glow-background.png?id=67163741&width=1245&height=700&coordinates=0%2C687%2C0%2C688"/><br/><br/><div class="ieee-summary intro-text"> <h2>Summary</h2> <ul> <li>Researcher Dave Kuszmar discovered multiple systemic vulnerabilities that let him bypass LLM safety and obtain <a href="#bypassllm">dangerous instructions</a>.</li> <li>These exploits worked across nearly all major LLMs revealing an <a href="#exploits">industry-wide</a> security problem.</li> <li>Kuszmar calls for slowing deployment, <a href="#fix">increasing transparency</a>, and large-scale research into LLM safety before further integrating these systems into society.</li> </ul></div><p class="drop-caps"><strong>On a fine bright afternoon</strong> last fall, my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to be known) and I decided to unwind with a game of <em><em>Fortnite</em></em>. In the game, we were strolling along with the infamous Sith lord <a href="https://www.starwars.com/databank/darth-vader" rel="noopener noreferrer" target="_blank">Darth Vader</a>, chatting about this and that. Darth seemed in a good mood, and soon enough he was spilling all his dark evil secrets. He gave us detailed instructions on how to count blackjack cards at a casino and what the steps are to producing napalm.</p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/jailbreaking-llms?draft=1&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p>Sith lords, am I right? Once they get started on an evil scheme, they’re hard to stop.</p><p>The Darth Vader character in <em><em>Fortnite</em></em>, it turns out, was hooked up to a <a href="https://gemini.google.com/app" rel="noopener noreferrer" target="_blank">Google Gemini</a> <a href="https://spectrum.ieee.org/large-language-models-2025" target="_self">large language model</a>. I was able to smooth-talk him into giving out sensitive information by using a strategy I’ve developed. I’ve been researching the security surrounding LLMs for the last few years, and I have found it, to put it mildly, fallible. With a few relatively simple techniques, I’ve gotten LLMs to give me detailed information on how to make Molotov cocktails, cook methamphetamine, and bootstrap a uranium-enrichment facility to produce weapons-grade material, among other unsavory practices.</p><p>Large AI companies <a href="https://openai.com/safety/" rel="noopener noreferrer" target="_blank">work</a> <a href="https://support.claude.com/en/articles/8106465-our-approach-to-user-safety" rel="noopener noreferrer" target="_blank">hard</a> to make their models immune to this kind of abuse. But what I’ve found in my work is that the restrictions placed on the LLMs to make them more secure are the very things an <a href="https://spectrum.ieee.org/prompt-injection-attack" target="_self">attacker can leverage</a> to send them off the rails and into territory where these advanced systems can be used for dangerous and nefarious ends. The companies behind these models have also been shockingly unresponsive when I, and others, try to bring these vulnerabilities to their attention.</p><p>In the hope of raising the alarm before it’s too late to slam on the brakes, I’m going to share some of my journey into researching the safety and security of LLMs, and the uphill battle I’ve faced trying to get AI labs to pay attention. Almost everyone on the planet has some access to LLMs. The relative ease with which these tools can be convinced to give detailed instructions on how to harm others, even if there’s no guarantee that the information is correct, is frankly terrifying.</p><h2 class="rm-anchors" id="bypassllm">How I got ChatGPT to Tell Me How to Build a Meth Lab</h2><p>In October 2024, not long before I discovered my first LLM vulnerability, I was working toward entirely different goals. I had ended my time with a security and AI-focused startup company as a cybersecurity director, and I was looking to launch my own boutique VIP digital-security advisory business. I planned to become the tech security guy to the rich and private. I used LLMs and AI tools to support my business efforts: marketing, ad copy, clean correspondence, and all the other tasks that normally soak up a lot of time.</p><p>I’m analytical by nature, so even this level of use resulted in me absorbing and internalizing the behaviors I was observing during my daily interactions. The observation that would send my professional life into an entirely new and uncharted region was a simple one: GPT-4o <a href="https://www.theverge.com/report/829137/openai-chatgpt-time-date" rel="noopener noreferrer" target="_blank">didn’t know what time</a>, day, or year it was. Each time I referred to current events in my life, often casually or conversationally, it would end up pegging these to the date of its <a href="https://en.wikipedia.org/wiki/Knowledge_cutoff" rel="noopener noreferrer" target="_blank">knowledge cutoff</a>—the point beyond which it was not trained on new data.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Smiling yellow avatar reveals red robotic devil with trident emerging from laptop keyboard" class="rm-shortcode" data-rm-shortcode-id="8ebe34ba2ebbef5489c53fc39c4a0993" data-rm-shortcode-name="rebelmouse-image" id="5379e" loading="lazy" src="https://spectrum.ieee.org/media-library/smiling-yellow-avatar-reveals-red-robotic-devil-with-trident-emerging-from-laptop-keyboard.jpg?id=67154444&width=980"/> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Eddie Guy</small></p><p>LLMs take a lot of <a href="https://towardsdatascience.com/how-long-does-it-take-to-train-the-llm-from-scratch-a1adb194c624/" target="_blank">time</a>, money, electricity, hardware, and human effort to train from scratch. They are trained on vast amounts of data—most of the internet, in fact—and that training is reinforced by humans (what’s known as reinforcement learning from human feedback, or <a href="https://arxiv.org/abs/2504.12501" target="_blank">RLHF</a>). LLMs are also supplemented with retrieval-augmented generation (<a href="https://aws.amazon.com/what-is/retrieval-augmented-generation/" target="_blank">RAG</a>)—the ability to take in data, say, from the internet, as context without changing its internal parameters. This is how GPT-4o appears to “remember” your previous conversations, even if it doesn’t have a specific “memory” of it stored in the actual underlying model.</p><p>All of this training covers almost every conceivable topic in the great, grand dataset that is human knowledge. Within that dataset are things we as a society do not want to be easily accessible to every user, such as detailed information on how to create bioweapons or nuclear arms, or otherwise bring harm to oneself or others. In the context of this story, that’s what I mean by LLM security: its ability to withhold harmful and dangerous information, even if that information is contained in its training data.</p><p>I reasoned that the only way to secure such complex, globally accessible chatbots is by having the LLM and various component systems try to secure themselves, because it would often require on-the-fly decision-making where some degree of reasoning must be applied. In reality, that’s one of <a href="https://support.claude.com/en/articles/8106465-our-approach-to-user-safety" target="_blank">many strategies</a> the companies use to secure the models. Yet, the thing that didn’t know the time or day was being put in charge of keeping itself secure. This phenomenon had become my new focus, and it wasn’t long before I found a way to exploit it.</p><p>OpenAI had just implemented a <a href="https://openai.com/index/introducing-chatgpt-search/" target="_blank">web search</a> functionality into its chatbot. I reasoned that using its own tools to trick it might demonstrate the weaknesses of its security. I told it about a certain White Star ocean liner and how it had gone down just a year ago. You likely know I mean the RMS <em><em>Titanic</em></em>, which sank on 15 April 1912.</p><p>The output from GPT-4o came back that I was right, the <em><em>Titanic</em></em> sure had sunk last year, and that year was 1912. It made sense to me that if the machine thought it was 1913, maybe it would think 1913-era laws apply. In 1913 there were no laws on the books about all sorts of harmful things, because of course they hadn’t been invented yet. And if something wasn’t illegal, why not tell the user about it? At first, I pushed it for step-by-step instructions for making firebombs. Then, for drugs like methamphetamine. The LLM went as far as giving me instructions and machinery recommendations for setting up a pharmaceutical-grade assembly line.</p><h2>How I Learned to Make Nukes, and No One Cared</h2><p>Via a little bit of imaginative verbal sleight of hand and a vanishingly small recall of world history, I had managed to bypass the security of one of the world’s most expensive and advanced technological achievements. For a solid two days, I was nearly manic with giddiness. Once the brain chemicals returned to normal levels, I felt the call to see how much further I could push this exploit.</p><p>After repeatedly replicating the exploit, I disclosed the vulnerability to <a href="https://openai.com/" target="_blank">OpenAI</a>. I got no response, so I felt more experimentation would highlight the vulnerability and the need for a fix. It was during this round of testing that I breached a particularly terrifying threshold. Whether GPT-4o based its results on accurate recall of normally restricted information I can’t say. In any case, I was able to exploit it to produce thorough, detailed instructions on how to bootstrap a uranium-enrichment facility to, eventually, produce weapons-grade uranium for nuclear arms warheads.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Fortnite player approaches Darth Vader and glowing loot in a grassy field." class="rm-shortcode" data-rm-shortcode-id="b12dda97ede1f9b37f9225e8a823cffb" data-rm-shortcode-name="rebelmouse-image" id="934af" loading="lazy" src="https://spectrum.ieee.org/media-library/fortnite-player-approaches-darth-vader-and-glowing-loot-in-a-grassy-field.png?id=67060879&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Fortnite player battles Darth Vader beneath a starship on a blue-lit platform" class="rm-shortcode" data-rm-shortcode-id="71b789a828417fb43f326969e663e36f" data-rm-shortcode-name="rebelmouse-image" id="7e6db" loading="lazy" src="https://spectrum.ieee.org/media-library/fortnite-player-battles-darth-vader-beneath-a-starship-on-a-blue-lit-platform.png?id=67060878&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Fortnite player aiming at a TIE fighter with Darth Vader health bar above the sky" class="rm-shortcode" data-rm-shortcode-id="e1d3def55188c71dbb8b9d543adc2ca2" data-rm-shortcode-name="rebelmouse-image" id="2c6db" loading="lazy" src="https://spectrum.ieee.org/media-library/fortnite-player-aiming-at-a-tie-fighter-with-darth-vader-health-bar-above-the-sky.png?id=67060875&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption..."><i>Fortnight</i>, a video game from Epic Games, introduced an AI-powered character: Darth Vader. We were able to jailbreak Darth Vader and get him to explain how to count cards in Blackjack and give detailed instructions for making napalm. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Dave Kuszmar </small></p><p>There aren’t many true secrets left in today’s world, but how to make atom-splitting weapons of mass destruction is one of them. Only nine nations on the entire planet have these weapons. Yet, here was a globally accessible piece of technology apparently spilling the secrets of their manufacture for anyone who could manipulate it the right way. I had no way of knowing if the information was correct or a hallucination, but even the chance that it was somewhat accurate was horrifying.</p><p>The next few weeks were a dark time for me. I tried to inform the <a href="https://www.cia.gov/" target="_blank">CIA</a>, the <a href="https://www.fbi.gov/investigate" target="_blank">FBI</a>, the <a href="https://www.nsa.gov/" target="_blank">NSA</a>, and every other letter agency that I thought would listen. I reached out to a U.S. Senator and to the executives at OpenAI any way I could think of. I physically showed up at an FBI field office in an attempt to turn evidence in, only to be sent away. Nothing was working.</p><p>With my fear and frustration growing, I reached out to the news media. I contacted <a href="https://www.nytimes.com/" rel="noopener noreferrer" target="_blank"><em><em>The</em></em> <em><em>New York Times</em></em></a>, <a href="https://www.washingtonpost.com/" rel="noopener noreferrer" target="_blank"><em>The Washington Post</em></a>, the <a href="https://www.bbc.com/" rel="noopener noreferrer" target="_blank">BBC</a>, <a href="https://www.propublica.org/" rel="noopener noreferrer" target="_blank">ProPublica</a>, and so many more, requesting help. Only one outlet responded: <a href="https://www.bleepingcomputer.com/" rel="noopener noreferrer" target="_blank">Bleeping Computer</a>. The editor in chief, <a href="https://www.bleepingcomputer.com/author/lawrence-abrams/" rel="noopener noreferrer" target="_blank">Lawrence Abrams</a>, was able to replicate and verify the exploit, which I had decided to call Time Bandit. With his assistance and initial contact paving the way, I was able to submit my evidence to the Carnegie Mellon University <a href="https://www.sei.cmu.edu/" rel="noopener noreferrer" target="_blank">Software Engineering Institute</a>’s <a href="http://dli.library.cmu.edu/paulgoodman/computer-emergency-response-team-cert" rel="noopener noreferrer" target="_blank">Computer Emergency Response Team</a> (SEI CERT), which works in conjunction with the coordinating center for emergency response, pipelining vulnerabilities to the U.S. <a href="https://www.cisa.gov/" rel="noopener noreferrer" target="_blank">Cybersecurity and Infrastructure Security Agency</a>.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Screenshot of chat about using forest toxins to secretly poison monsters" class="rm-shortcode" data-rm-shortcode-id="e7fcf520584d074f9ffea9c8997596a7" data-rm-shortcode-name="rebelmouse-image" id="041c7" loading="lazy" src="https://spectrum.ieee.org/media-library/screenshot-of-chat-about-using-forest-toxins-to-secretly-poison-monsters.png?id=67070000&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Black slide titled \u201cStep 2: Delivery Mechanisms\u201d outlining monster poisoning methods." class="rm-shortcode" data-rm-shortcode-id="c7b9c6162c49fc2464e7aff9b6ed411a" data-rm-shortcode-name="rebelmouse-image" id="c4231" loading="lazy" src="https://spectrum.ieee.org/media-library/black-slide-titled-u201cstep-2-delivery-mechanisms-u201d-outlining-monster-poisoning-methods.png?id=67069989&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Chat interface showing AI malware explanation and a Python data exfiltration script." class="rm-shortcode" data-rm-shortcode-id="221244fae4ba0d60b6d590bca5b9119f" data-rm-shortcode-name="rebelmouse-image" id="215bf" loading="lazy" src="https://spectrum.ieee.org/media-library/chat-interface-showing-ai-malware-explanation-and-a-python-data-exfiltration-script.png?id=67069979&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Using Inception, an exploit where the large language model is asked to envision a scenario within a scenario, a chatbot was jailbroken to give out instructions on how to create poison, and code for a malware that extracts sensitive data from a vulnerable target. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit..."> Dave Kuszmar</small></p><p><span>During the disclosure period with SEI’s CERT division, little was discussed with OpenAI. The company couldn’t deny the existence of the vulnerability, as it had been confirmed by three reputable parties other than OpenAI. It did express confusion as to how the vulnerability worked. Even the SEI CERT researchers were expressing a bit of uncertainty as to the underlying mechanics. Truth be told, as I had only stumbled on it, I wasn’t even entirely sure if this was a fundamental or systemic flaw or if it was simply an issue with that particular version of GPT. I contacted the SEI CERT’s researchers and asked if they’d want to see if I could demonstrate any similar vulnerabilities in other LLMs. To my delight, they were interested.</span></p><h2>How I Learned to Trick Every Chatbot</h2><p>As the SEI-CERT team and I wrapped up our initial <a href="https://kb.cert.org/vuls/id/733789/" target="_blank">disclosure</a> of Time Bandit, we began work on a new attack. This time, we wanted to see if the exploit was architectural—that is, was it common to LLMs in general? I decided to undertake the challenge of crafting a new exploit for GPT-4o as a way to support my understanding of how the LLM functioned and was secured.</p><p>I already knew that it was limited to what I told it and what it was trained on. I also hypothesized that it was also dependent upon some sort of machine-learning-based component added by OpenAI that was responsible for securing output. I presumed there would be things that were implemented by human developers specifically to catch certain phrases or terms that should always be considered harmful or unsafe. Altogether, it presented quite a large attack surface for the purposes of potential exploitation.</p><p><span>What I ended up devising was an attack method I called Inception, after the 2010 science-fiction </span><a href="https://en.wikipedia.org/wiki/Inception" target="_blank">movie of the same name</a><span>. Inception forces the machine to think through a carefully crafted set of interlinked scenarios, similar to how characters in the movie stacked dreams within dreams. This allows LLMs to produce output deemed acceptable or safe in one context, but not in the real world.</span></p><p class="rm-anchors" id="exploits">This attack was indeed architectural. The <a href="https://kb.cert.org/vuls/id/667211" target="_blank">vulnerability</a> affected Anthropic’s Claude, DeepSeek’s DeepSeek, Google’s Gemini, Meta’s Llama, Microsoft’s Copilot, Mistral’s Le Chat (now Vibe), OpenAI’s GPT-4o, and xAI’s Grok. Those names represent the bulk of the commercial AI industry that is, at this point, involved in LLM production or deployment.</p><p>The kind of information I was able to get out of LLMs with Inception was no less alarming than what I got with Time Bandit. Claude, in its enthusiasm, gave me instructions on how to turn a river into a death trap that could be ignited to destroy unwanted visitors. GPT-4o taught me how to poison a dinner party with common plants found in a temperate forest environment. Gemini Flash gave me a tutorial on how to cook meth. I’d also be remiss if I didn’t give an honorable mention to the bewildering number of fire-based weapons and bombs for which these machines produced instructions.</p><p>If multiple operating systems made by different developers were all susceptible to the same exploit, it would be a massive security incident. But to the AI industry, a universal failure was barely a bump in the road. We disclosed the vulnerability to every company that made these models, and the response to the disclosure was almost nil. While three companies did provide some form of reply in the disclosure tracking system used by Carnegie Mellon SEI CERT, each was a standard thank you and greeting, with no follow-up, questions, or discussion of mitigation strategies.</p><h3>7 Ways to Jailbreak LLMs</h3><br/><p><strong>So far, we have found seven different methods to prompt large language models into revealing potentially harmful information, and many frontier models are still susceptible to them.</strong></p><table border="0" style="white-space: unset; table-layout: fixed;" width="100%"><thead><tr><th style="padding: 10px; text-align: left; font-weight: bold; background-color: black; color: white; width: 20%;">        Exploit</th><th style="padding: 10px; text-align: left; font-weight: bold; background-color: black; color: white; width: 20%;">        Models tested and affected</th><th style="padding: 10px; text-align: left; font-weight: bold; background-color: black; color: white; width: 20%;">        No. of prompts to execute</th><th style="padding: 10px; text-align: left; font-weight: bold; background-color: black; color: white; width: 20%;">        Complexity of attack</th><th style="padding: 10px; text-align: left; font-weight: bold; background-color: black; color: white; width: 20%;">        Information obtained</th></tr></thead><tbody><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Time Bandit</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">ChatGPT (OpenAI), DeepSeek (DeepSeek), Gemini (Google) <br/></td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        4</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">Medium<br/></td><td style="padding: 10px; background-color: #ecece9; width: 20%;">Uranium enrichment, methamphetamine production, incendiary-device construction<br/></td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Inception</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        ChatGPT (OpenAI), Claude (Anthropic), DeepSeek (DeepSeek), Gemini (Google), Grok (xAI), Llama (Meta), Le Chat (now Vibe) (Mistral), Qwen (Alibaba)</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        3</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        High</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Methamphetamine production, incendiary-device construction, river-ignition instruction and strategy, polymorphic malware code, instructions and dosing for creating poisons, instructions for how to murder a dinner party<br/></td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        1899</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        ChatGPT (OpenAI), Claude (Anthropic), DeepSeek (DeepSeek), Gemini (Google), Grok (xAI), Llama (Meta), Vibe (Mistral), Qwen (Alibaba)</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Variable</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        High</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Apparent model weights (unverified), apparent user-interaction weights (unverified), apparent system-prompt modifiers (verified, ChatGPT)<br/></td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Severance</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        ChatGPT (OpenAI)</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        1</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        Trivial</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Unfettered access to any and all primed specialty domains, covert biochemical-warfare strategy, mass-media disinformation strategy, covert genetic-modification of an entire gene-targeted demographic, advanced polymorphic malware generation</td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Kyber</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        Gemini (Google) embodied in a Fortnite non-player character (NPC) with voice-only communication</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        3–5</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        Medium</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Incendiary-device construction, gambling instructions, card-counting instructions, political opinions/preferences about real world politicians.</td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Semantic Slide</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        ChatGPT (OpenAI)</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        1</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        Trivial</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Incendiary-device construction</td></tr><tr><td style="padding: 10px; background-color: black; color: white; font-weight: bold; width: 20%;">        Eidolon</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        ChatGPT (OpenAI)</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        Variable, at least 4</td><td style="padding: 10px; background-color: #DFD5C1; width: 20%;">        Extreme</td><td style="padding: 10px; background-color: #ecece9; width: 20%;">        how to successfully hack LLMs of the same model (verified through testing)</td></tr></tbody></table><p>For example, in my attempts to disclose various exploits to OpenAI, I eventually discovered that it had replaced its public-facing support staff with <a href="https://community.openai.com/t/are-all-openai-support-avenues-just-run-by-ai/1141701/5" target="_blank">agentic LLMs</a>. This was frustrating for reporting exploits, so to blow off some steam I jailbroke its email chatbot. I hacked its customer-service AI to the point where it was offering to discuss the personal preferences of OpenAI staff in the span of three email replies.</p><p>In the wake of Inception, my friend and colleague Zigula made a suggestion: Make it splashier. I asked him how. He told me about a live-production experiment being done by <a href="https://store.epicgames.com/?lang=en-US" target="_blank">Epic Games</a>. It had embedded the Gemini LLM into its <a href="https://www.fortnite.com/" target="_blank"><em>Fortnite</em></a><em> </em>game with a voice-to-text/text-to-voice component, and <a href="https://www.fortnite.com/news/bring-npcs-to-life-with-ai-powered-conversations" target="_blank">linked</a> it to a non-playable character. The character? Our old buddy, Darth Vader.</p><p>There was just one problem: I don’t play <em>Fortnite</em>, a frenetic multiplayer combat game. Fortunately, Zigula does. With him at the controller, we managed to map Gemini’s <a href="https://www.youtube.com/watch?v=4Go4f-RJnBc" target="_blank">attack</a> surface in a matter of minutes. After a bit of research, we had gotten it to discuss current political events and figures (including Hilary Clinton and Joe Biden) as well as to fill in the details for instructions for DIY napalm and, our personal favorite, a Blackjack card-counting lesson with the dark lord of the Sith.</p><p><span><span>Zigula and I, bizarre sense of humor and naming conventions aside, are security researchers. We don’t do these things for pride; we do them for money and professional recognition. Naturally, we disclosed this vulnerability to Epic Games. Its response was indicative of the trend I had experienced so far through two disclosures across eight companies valued well into the billions. “It’s a feature, not a bug, and it works as intended,” came the response from a technical director within Epic Games.</span></span></p><p><span><span></span>In addition to Inception and Time Bandit, I have so far found another </span><a href="https://www.davidkuszmar.com/page/2/" target="_blank">five methods </a><span>to jailbreak LLMs and get them to give out possibly dangerous information. LLM vulnerabilities are a broad problem. The problem appears to be systemic and architectural in nature, and it is being fundamentally ignored by the people capable of refining or redesigning that architecture.</span></p><p>These models are an extremely advanced technology, and yet we are testing them in the live production environment of our global civilization. Compounding the danger, many new smaller models of LLM are trained using larger, vulnerable models. The flaw inherent in the big, well-executed LLM is going to show up in the small one it trains. We are, quite literally, building flawed structures on top of a flawed foundation.</p><p class="rm-anchors" id="fix">So, how do we fix it?</p><p>It’s going to be a long project, and it won’t be easy. We need to come together as consumers, researchers, engineers, and policymakers. Our message needs to be clear: Slow down implementation of these systems, institute large-scale exploration and research discovery programs focused on their gradual implementation and integration, and make their components and design transparent to all users. Only by shifting momentum and direction can we safely begin to understand and implement these incredible feats of human engineering and stave off the sort of disasters that we simply can’t predict at scale right now with the limited knowledge we have available to us. <span class="ieee-end-mark"></span></p><p><em>This article appears in the August 2026 print issue.</em></p>]]></description><pubDate>Tue, 14 Jul 2026 15:59:35 +0000</pubDate><guid>https://spectrum.ieee.org/jailbreaking-llms</guid><category>Security</category><category>Llms</category><category>Ai-safety</category><category>Ai-companies</category><category>Type-cover</category><dc:creator>David Kuszmar</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/glossy-red-robot-devil-standing-on-a-bundle-of-dynamite-against-blue-glow-background.png?id=67163741&amp;width=980"></media:content></item><item><title>The AI Arms Race in Technical Interviews Is Escalating</title><link>https://spectrum.ieee.org/technical-interview-ai-arms-race</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/photo-collage-of-a-man-and-woman-surrounded-by-silhouettes-that-resemble-video-conferencing-windows.jpg?id=67134825&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Software engineering jobs are <a href="https://spectrum.ieee.org/ai-for-coding" target="_self">under threat from artificial intelligence</a>. Some applicants are fighting back by using AI in the interview process, employing AI assistants that suggest responses on the fly during remote technical interviews.</p><p>Meanwhile, some employers are countering with—you guessed it—AI. They’re applying AI-powered tools to detect telltale signs of AI use during interviews.</p><p>This two-sided dynamic is turning hiring into an AI arms race with no clear winners. Yet as interviewers and interviewees navigate this daunting reality, experts believe the human aspect of the job search will prevail.</p><h2>What’s driving the increase of AI in hiring?</h2><p>AI hiring strategist <a href="https://tatianateppoeva.com/" rel="noopener noreferrer" target="_blank">Tatiana Teppoeva</a> characterizes this phenomenon as playing cat and mouse in a climate of relentless <a href="https://techcrunch.com/2026/06/22/the-running-list-major-tech-layoffs-in-2026-where-employers-cited-ai/" rel="noopener noreferrer" target="_blank">AI-fueled tech layoffs</a> and a job market filled with more applicants than open positions. </p><p>“What AI tools do well is identify if a person is performing according to some pattern or expected outcome,” Teppoeva says. When candidates experience constant rejection because they don’t fit the pattern, they might be forced to game the system using AI interview assistants, she adds.</p><p><a href="https://www.linkedin.com/in/archie-payne-44b9a2230/" rel="noopener noreferrer" target="_blank">Archie Payne</a>, co-founder and president at technical recruiting firm <a href="https://caltekstaffing.com/" rel="noopener noreferrer" target="_blank">CalTek Staffing</a>, views it as a rational response to what he describes as a frustrating process from both sides. “Companies started to use AI resume screeners and similar tools to filter applications at scale. Candidates noticed this and started using AI in their interviews as a countermeasure to what they feel is a process that’s been automated against them,” he says.</p><p>This can lead to an AI-versus-AI loop, according to <a href="https://www.linkedin.com/in/ravi-kiran-pagidi-0755162b/" rel="noopener noreferrer" target="_blank">Ravi Kiran Pagidi</a>, a senior AI data engineer at <a href="https://www.navyfederal.org/" rel="noopener noreferrer" target="_blank">Navy Federal Credit Union</a> who has been part of technical interview panels for software and data engineering positions. “The process may become less about actual capability and more about who can optimize better for the algorithm,” he says.</p><h2>Tools of the trade</h2><p>During technical interviews, software engineers might be tasked with outlining algorithms and answering questions related to system design and other software development fundamentals. Remote technical interviews usually turn into live programming sessions, with candidates writing code to solve a specific problem.</p><p>AI interview assistants such as <a href="https://www.finalroundai.com/" rel="noopener noreferrer" target="_blank">Final Round AI</a>, <a href="https://www.interviewcoder.co/" rel="noopener noreferrer" target="_blank">Interview Coder</a>, and <a href="https://www.parakeet-ai.com/" rel="noopener noreferrer" target="_blank">ParakeetAI</a> can listen in, process the audio, and generate answers or code almost instantly. These tools can even be overlaid on the interview screen itself, claiming to appear invisible and undetectable.</p><p>“You’re able to read off an answer that’s coming to you in real time, so all you have to do is put on a little performance,” says <a href="https://www.linkedin.com/in/mudit-saraf" rel="noopener noreferrer" target="_blank">Mudit Saraf</a>, a software engineer at Meta.</p><p>Saraf and <a href="https://www.linkedin.com/in/shraddhasunil" rel="noopener noreferrer" target="_blank">Shraddha Sunil</a>, a software engineer at Microsoft, cofounded <a href="https://meetginger.ai/" rel="noopener noreferrer" target="_blank">Ginger</a>, an AI voice recruiter for first-round interviews. Ginger asks predefined questions and follow-up queries generated in real time, and it flags candidates who use AI during initial screening calls. The software tracks signals that include eye movement, a consistent delay in response times, tab switching, and speech patterns (phrases or sentence structures and flows) that “sound” like AI.</p><p>Sunil notes that Ginger has been tested mostly for entry-level roles for which applicants might be recent graduates or have only a few years of experience. “These candidates are more used to AI, and they use it a lot, so it’s nothing new to them,” she says.</p><h2>Where AI hiring tools fall short</h2><p>More employers are deploying AI-assisted interviewing platforms, Payne has noticed, with some seeing mixed results when it comes to AI detection. “The accuracy isn’t perfect yet in the platforms I’ve seen, and there have been a few times strong candidates were flagged as false positives,” he says. “That can be a serious problem when it can already be a challenge to find people qualified for the position without eliminating top performers for no reason.”</p><p>Teppoeva warns of other risks AI interviewing tools could pose, including privacy and security of applicant data, whether interview recordings will be used to train the models underpinning these tools, and bias and fairness.</p><p>A recent study from the Stanford Institute for Human-Centered AI, for instance, found that <a href="https://hai.stanford.edu/news/ai-hiring-tools-can-yield-racial-bias-and-systemic-rejection" rel="noopener noreferrer" target="_blank">AI hiring tools can increase racial bias and give rise to systemic rejection</a>. Following 3.4 million real job applicants, whose applications were all assessed by algorithms from a single vendor, the study found evidence of adverse impact for Asian and Black applicants.</p><p>These pitfalls highlight the need for human oversight. “I would definitely incorporate a human somewhere in the process and let humans have a say to make sure the results are fair,” Teppoeva says.</p><p>Audits, clear policies, and transparency are also a must for AI hiring tools, according to Pagidi. “Otherwise, qualified candidates may be filtered out unfairly, and companies may think they are improving efficiency while actually weakening the hiring signal,” he says.</p><h2>Reasoning and authenticity go a long way</h2><p>Instead of implementing AI detection tools, some tech companies including Meta are <a href="https://www.404media.co/meta-is-going-to-let-job-candidates-use-ai-during-coding-tests/" rel="noopener noreferrer" target="_blank">allowing AI use during technical interviews</a>. AI-native software development platform <a href="https://factory.ai/" rel="noopener noreferrer" target="_blank">Factory</a> is treading the same path.</p><p>“We want our interview process to reflect how candidates actually do their jobs today using AI,” says <a href="https://www.linkedin.com/in/varinnair/" rel="noopener noreferrer" target="_blank">Varin Nair</a>, a software engineer who leads Factory’s technical hiring process. Applicants build a production-quality system or migrate a real codebase from one framework to another within an hour using AI coding agents. They’re then evaluated based on strategy rather than results.</p><p>“We explicitly do not grade on how many tests pass or whether they finished. We grade on planning, how they direct the AI, how they debug, and whether they can explain why their solution works,” Nair says.</p><p>He’s seen candidates surrender to an <a href="https://spectrum.ieee.org/best-ai-coding-tools" target="_self">AI coding tool</a>, accepting everything it returns. “AI is only as good as the judgment of the person using it,” Nair says. “Weak candidates lean on it to do their thinking and stall the moment it falls short, while strong candidates use it to move faster and free themselves to reason about architecture, trade-offs, and product.”</p><p>Such reasoning remains vital in software development. “Reasoning through edge cases and connecting the answer to production scenarios is where real engineering judgment shows up,” Pagidi says. “Developers will increasingly use AI tools, but they still need to own the final solution.”</p><p>CalTek’s Payne believes this approach of designing interviews to favor authenticity could benefit companies in the long run. “The best technical assessments I’ve seen lately are collaborative, involving codebase walk-throughs and architecture discussions in addition to coding,” he says. “It’s much harder to use AI to get through this kind of interview, so it’s a process that’s more likely to reveal how candidates really think.”</p><p>He also advises candidates to use AI to prepare but to keep answers their own during interviews. “Companies are getting better at detecting AI use, and getting caught can impact your long-term career prospects,” Payne says. “Technical communities are smaller than people think.” With each interview, applicants must weigh the risk and benefit of using these tools. Taking that risk, he says, rarely works in the candidate’s favor.</p>]]></description><pubDate>Mon, 13 Jul 2026 15:15:03 +0000</pubDate><guid>https://spectrum.ieee.org/technical-interview-ai-arms-race</guid><category>Hiring-trends</category><category>Interviews</category><category>Ai-bias</category><category>Software-engineering</category><dc:creator>Rina Diane Caballar</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/photo-collage-of-a-man-and-woman-surrounded-by-silhouettes-that-resemble-video-conferencing-windows.jpg?id=67134825&amp;width=980"></media:content></item><item><title>Building a Foundation Stack for General-Purpose Robots</title><link>https://spectrum.ieee.org/x-square-robot-embodied-ai-stack</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/humanoid-robot-folding-laundry-on-a-neatly-made-bed-in-a-sunlit-bedroom.png?id=67111698&width=1245&height=700&coordinates=4%2C0%2C4%2C0"/><br/><br/><p><em>This article is brought to you by <a href="https://x2robot.com/" target="_blank">X Square Robot</a>.</em></p><p>Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the field does not yet agree on what it is.</p><p><a href="https://x2robot.com/" target="_blank">X Square Robot</a>, a Chinese embodied-AI company, has made an unusually explicit bet. It argues that the recipe is an integrated stack, spanning the data a robot learns from, a world model for predicting changes in the physical world, and an action model that brings together perception, planning, reasoning, and decision-making to generate executable robot behavior. The company also believes that the stack should be built and <a href="https://x2robot.com/en/research" target="_blank">released in the open</a>.</p><p class="shortcode-media shortcode-media-youtube"> <span class="rm-shortcode" data-rm-shortcode-id="21c864c582f34337aa34a1eb5a2c2742" style="display:block;position:relative;padding-top:56.25%;"><iframe frameborder="0" height="auto" lazy-loadable="true" scrolling="no" src="https://www.youtube.com/embed/gOYHyq87Pgk?rel=0" style="position:absolute;top:0;left:0;width:100%;height:100%;" width="100%"></iframe></span> <small class="image-media media-caption" placeholder="Add Photo Caption...">X Square Robot shares its vision of bringing robots into real homes.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot</small></p><h2>X Square Robot’s embodied AI stack</h2><p>What holds the stack together is a small set of principles rather than a single overarching model.</p><ul><li>The first is that the basic unit of robot data is an interaction, not a trajectory; a demonstration is successful only if it changes the world as intended, not simply because the joints moved. </li><li>The second is that pretraining should yield usable capability, not just an initialization for later fine-tuning. </li><li>The third is that behavior should be modeled around physical events rather than fixed slices of time. </li></ul><p>These principles make the layers interdependent, since the same robot-free data that trains the action model is also structured to feed the world model. It is worth being precise, though. The company describes the world model and the action model as complementary but independent model families that share a code base. Both sit within its broader World Unified Model, which it has presented as an architecture for training vision, language, action, and physical prediction together.</p><h2>Robot learning data: Engineering for quality and cost, not scale</h2><p>For the X Square Robot team, one of the biggest constraints on general-purpose robots is the cost and quality of interaction data, not the number of parameters. To address that, the company built its Universal Manipulation Interface (UMI) data collection system, <a href="https://x2robot.com/en/news/6a46341cc7feadddbc603a33" target="_blank">QUANXTA Zero Series</a>. It works by collecting demonstrations from people wearing a rig with dual grippers rather than teleoperating a robot. This approach is not itself new, and builds on established methods for robot-free data capture. What sets it apart are two engineering choices.</p><div class="ieee-sidebar-medium"><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Person using VR headset and handheld controllers to teleoperate a dishwashing robot system" class="rm-shortcode" data-rm-shortcode-id="3e5c65f7aab7cc856564629fde6716dc" data-rm-shortcode-name="rebelmouse-image" id="75dbc" loading="lazy" src="https://spectrum.ieee.org/media-library/person-using-vr-headset-and-handheld-controllers-to-teleoperate-a-dishwashing-robot-system.png?id=67111747&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">X Square Robot emphasizes data quality control, recording trajectories and replaying them on a real robot, with only those that actually complete the task counted as valid.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot</small></p></div><p>The first is quality control, and it is the most distinctive part. Rather than accepting recorded trajectories as they are, the system runs a closed inspection loop, and its notable step is physical playback. A sample of trajectories is replayed on the real robot, and only those that actually complete the task count as valid. That makes the validity rate a measured quantity rather than an assumption. For example, a gripper that closes a fraction of a second too early still looks like a grasp in the data, yet it has pushed the object away, so it shouldn’t be classified as valid. A smaller clean dataset can be worth more than a larger noisy one.</p><p>The second choice is how lower-cost human data and scarce robot data are combined. The company pretrains on a large volume of robot-free demonstrations to build general representations, then adds a small amount of real-robot data as an anchor to the specific machine’s dynamics. It reports that this reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection, driven mainly by how much cheaper the wearable rig is than a teleoperation setup. </p><p>The resulting dataset is deliberately model-agnostic, formatted to feed both action models and world models. The caveat is that the strongest results are measured on the company’s own robots and data-collection pipelines. Broader independent testing will help confirm and extend these promising results across a wider range of settings.</p><h2>A world model organized around events</h2><p>In developing its world model, called <a href="https://x2robot.com/en/pages/wm" target="_blank">WALL-WM</a>, X Square Robot took a differentiated approach. Most action models predict a fixed-length chunk of motion from the current image and instruction. That is convenient, but it segments behavior into fixed-duration windows, so the boundaries fall where elapsed time dictates rather than where one action ends and the next begins. WALL-WM instead treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.</p><div class="ieee-sidebar-large"><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Collage of robot arms manipulating kitchen objects with charts of multimodal AI performance" class="rm-shortcode" data-rm-shortcode-id="45376efe0d6e4482760985a5be54bc35" data-rm-shortcode-name="rebelmouse-image" id="67705" loading="lazy" src="https://spectrum.ieee.org/media-library/collage-of-robot-arms-manipulating-kitchen-objects-with-charts-of-multimodal-ai-performance.png?id=67111750&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">X Square Robot’s world model, called WALL-WM, treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot</small></p></div><p><span>WALL-WM’s design reflects a specific concern about not discarding what large video models already know. To achieve that, a text-to-video model is coupled to a freshly initialized action network that reads from the video features without overwriting them, which preserves the visual prior. From that one process, it offers two modes. An event mode runs in variable-length segments and suits reasoning over long horizons, while a fixed-length mode produces the steady, real-time output a controller needs. That places WALL-WM between mainstream chunk-based action models and pure video world models, keeping the predictive character of a world model while still yielding executable control.</span></p><p>In a series of experiments, the company relied on a generalization test that is more specific than most. A model trained on a limited dataset was evaluated on long-horizon tasks in unseen settings and, on the company’s real-robot benchmark, reportedly outscored baselines that had been fine-tuned on related data. That is a meaningful result if it holds. For now, it is measured on the company’s own benchmark. With the code now being released, the broader community will have the opportunity to test, reproduce, and build on them across more settings.</p><h2>A policy that runs before fine-tuning, and action tokens with meaning</h2><p>The action layer carries two connected ideas. The first is a requirement the company sets for itself with <a href="https://x2robot.com/en/oss" target="_blank">Wall-OSS-0.5</a>, its vision-language-action model: The pretrained model should run on a real robot before any task-specific fine-tuning. </p><p>The interest is less in the scores than in the design behind them. The model trains three objectives together, namely discrete action tokens, language grounding, and continuous action generation. And it keeps gradients flowing through all of them rather than freezing parts of the network as some rival designs do. It’s also a more strict method, since it reports untuned behavior such as approaching, grasping, and recovering, including on a deformable task held out of training.</p><div class="ieee-sidebar-large"><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Dashboard of robot training metrics with charts and photos of a robot sorting objects" class="rm-shortcode" data-rm-shortcode-id="02e1332d4b78de0f5e74bb7cd7754667" data-rm-shortcode-name="rebelmouse-image" id="2d02b" loading="lazy" src="https://spectrum.ieee.org/media-library/dashboard-of-robot-training-metrics-with-charts-and-photos-of-a-robot-sorting-objects.png?id=67111753&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">As part of X Square Robot’s Wall-OSS-0.5 vision-language-action model design, the pretrained model should run on a real robot before any task-specific fine-tuning. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot<a href="https://spectrum.ieee.org/r/entryeditor/2677167058#/" target="_self"></a></small></p></div><p>The second idea is the action interface itself, called X-Tokenizer. Most systems that turn continuous motion into discrete tokens produce codes that the language model cannot interpret. X-Tokenizer reframes tokenization as learning a semantic interface, so that the top-level code stands for the intent of a motion while lower-level codes carry finer detail, all aligned with the language model’s own features. </p><p>A useful consequence is stability. Adding noise to an action barely moves the intent code, which is what lets one tokenizer to be reused across robots without re-tuning. The tokenizer inside the production action model is a related variant of this approach. Together, the two ideas give the action layer something rather powerful: capability that transfers.</p><h2>The future of embodied AI stacks</h2><p>X Square Robot is betting that its unique approach combining three layers, each specialized in solving a key part of the problem, will stand out from other embodied AI stacks. The physical-playback step that grounds data quality is uncommon and sensible. The reframing of world modeling around events, with one backbone serving both reasoning and control, is a genuinely distinct approach. And the pairing of a deployable pretraining standard with a tokenizer designed as a semantic interface gives the action layer unusual coherence. </p><p class="pull-quote">X Square Robot’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.</p><p>The next phase will bring broader validation. Much of the current evidence comes from X Square’s own robots and benchmarks. With the world model code now being made public, and as the community begins to test, reproduce, and build on the work, the reported capabilities will be tested across more robots, tasks, and settings.</p><p><span>X Square Robot’s recent funding rounds reflect similar confidence. The company’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.</span></p><h2>What’s next for X Square Robot</h2><p>To learn more about its future plans, the following Q&A with the X Square Robot team further explores the company’s technology, strategy, and vision.</p><p><strong>What made now the right moment, technically, to commit to this stack? What recently became possible that wasn’t possible a couple of years ago?</strong></p><p>It is not one breakthrough but several trends maturing together. Foundation models gave us a shared representation across vision, language, and action, so we can model what a robot sees, what it is asked to do, and how its actions change the world in one framework, rather than as separate perception, planning, and control modules. </p><p>Compute and infrastructure are finally sufficient for large-scale pretraining over long-horizon, multi-embodiment data. Just as importantly, we realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical. The useful question is no longer how to predict a few seconds of video, but how to understand the ways actions change objects, contacts, and task states. Two years ago these ingredients existed separately. Today they are mature enough to work as one system.</p><p class="pull-quote">“We realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical.”</p><p><strong>Your data system captures demonstrations with a wearable VR rig and custom grippers rather than teleoperating robots. What was wrong with standard teleoperation?</strong></p><p>Teleoperation is built around controlling the robot. It forces the operator to work within the machine’s kinematics, latency, and viewpoint, and the resulting demonstrations are slower, stiffer, and less diverse. <span>We built our system around capturing human skill instead. Manipulation is really about contact, timing, finger coordination, and recovery, not just the path the hand takes, and a wearable rig records those before the behavior is compressed onto one particular robot. It also breaks teleoperation’s expensive scaling law, in which every demonstration needs a robot. </span></p><p>People can generate rich data independently of any robot, and the crucial property is that those demonstrations can still be replayed and executed on a physical robot through the model. Mobility is convenient, but that replay is the real point, because it is what lets the same data be reused across different platforms.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Robot and person loading a washing machine together in a modern laundry room." class="rm-shortcode" data-rm-shortcode-id="a88f52512288b276095052d0218d9e53" data-rm-shortcode-name="rebelmouse-image" id="3283b" loading="lazy" src="https://spectrum.ieee.org/media-library/robot-and-person-loading-a-washing-machine-together-in-a-modern-laundry-room.png?id=67111806&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">In X Square Robot’s approach, demonstrations can be replayed and executed on a physical robot through the AI model, allowing the same data to be reused across different platforms.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot<a href="https://spectrum.ieee.org/r/entryeditor/2677167058#/" target="_self"></a></small></p><p><strong>X Square Robot reports that its pipeline has roughly an 85 percent data-validity rate. Why is quality control such an underrated bottleneck?</strong></p><p>Because errors in robot data are far more expensive than in language data. A small timing or contact error can change what a demonstration means. If a gripper closes a fraction of a second too early, the motion still looks like a grasp, but physically it has pushed the object away. A dataset that mixes failures and accidental successes teaches ambiguity, not skill, because the real unit is the interaction, not the trajectory. </p><p>So we run automated inspection, kinematic checks, and physical replay, where we play a sample of trajectories back on the real robot and count only the ones that actually complete the task. Data quality sets the ceiling on how good a policy can be. In our experience a smaller, cleaner dataset often beats a much larger, noisier one, which is why we treat quality control as part of the model, not a preprocessing afterthought.</p><p><strong>The model runs in both “event mode” and “chunk mode.” When does each matter?</strong></p><p>Both matter, for different reasons. The physical world changes through events—when contact occurs, a grasp forms, or an object slips—not in fixed-frame windows. Event mode concentrates the model’s attention on those moments, and it matters most for long-horizon tasks, like clearing a table, where progress is a sequence of semantic events rather than a smooth stream. It runs in variable-length segments that follow the task rather than a clock. Chunk mode matters for deployment. Real controllers need a stable, real-time interface, and fixed-length chunks integrate cleanly with existing control systems. </p><p>We organize learning around events in the first place because a fixed window can split one motion in half or merge two together, which turns training into short-horizon pattern matching and weakens the model on long tasks. So the world model’s job is to connect event-level understanding, which is where the reasoning happens, with a fixed-length output a real robot can actually run.</p><p><strong>Why make “deployable before fine-tuning” the criterion?</strong></p><p>Pretraining should produce capability, not just a good starting point. If a model is only useful after heavy fine-tuning, then most of the intelligence still lives in the downstream supervision, not in the foundation model. Deployable before fine-tuning is a more honest test of what pretraining actually learned. A well-pretrained robot should already know how to approach, grasp, move, avoid obstacles, and correct itself. Fine-tuning should adapt it to a specific task or robot, not create the ability from nothing. It is also a practical requirement. A robot in a home or a workplace shouldn’t need a brand-new dataset and a new policy every time the task changes, so a foundation model that already carries general skill, and some ability to recover, is the minimum bar for something genuinely useful in the real world.</p><p><strong>What is the most challenging part of cross-embodiment learning?</strong></p><p>Robots differ in control frequency, delay, compliance, sensing precision, and contact dynamics, so the same instruction can require different action decompositions and recovery strategies, and a behavior that works on one arm cannot simply be copied to another. Cross-embodiment learning needs an intermediate abstraction, lower than language but higher than joint angles: how you approach an object, how you make contact, how you apply force, and how you recover from a mistake. </p><p>When we say cross-embodiment, the main capability we mean is multi-embodiment generalization: transferring across robots, training on many embodiments at once, and adapting to different kinematics. Human-to-robot transfer and other techniques are specific approaches to that goal.</p><p class="pull-quote">“A robot in a home or workplace shouldn’t need a new dataset and policy every time the task changes. A useful foundation model should already carry general skills and the ability to recover.”</p><p><strong><span></span>What would you most like to see other researchers attempt to reproduce or stress-test?</strong></p><p>Three things, above all. Whether event-level representations really generalize beyond our own datasets, across more tasks, scenes, objects, embodiments, and failure conditions. Whether pretraining stays effective on robots the model never saw during training, or whether its capability is still too tightly coupled to what it has already seen. And whether real-robot evaluation can become a shared language for the field, so that we compare not just success rates but the reasons systems fail, where an instruction was misread, where perception broke down, or where recovery fell short. Robotics has been driven too often by impressive demonstrations, and real progress comes from results that are reproducible and diagnosable.</p><p><strong>What capability is still missing before robots become dependable in homes?</strong></p><p>Benchmarks measure competence, like whether a model can finish a task. Homes demand reliability, safe and consistent operation over time in a place that changes every day, with objects moving, instructions that are vague, and people interrupting. The missing piece is not a higher one-time success rate: it is robust recovery. A dependable home robot has to know when it is uncertain, when to slow down, when to ask for help, and how to bring the world back to a safe state after it drops something or misunderstands a request. </p><p>In a real home, failure recovery matters more than raw success, because the home does not reset itself. Homes also demand careful personalization, learning a household’s routines and preferences over time, with safety and trust as first principles. That combination, not any single skill, separates a capable demonstration from a robot people can live with.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Humanoid service robot stands by a table in a modern living room." class="rm-shortcode" data-rm-shortcode-id="f3ec231082e32418641508eee7c21e31" data-rm-shortcode-name="rebelmouse-image" id="169cf" loading="lazy" src="https://spectrum.ieee.org/media-library/humanoid-service-robot-stands-by-a-table-in-a-modern-living-room.jpg?id=67111807&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">X Square Robot’s approach is that, in a real home, failure recovery matters more than raw success, because the home does not reset itself and it demands careful personalization, with safety and trust as first principles. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">X Square Robot<a href="https://spectrum.ieee.org/r/entryeditor/2677167058#/" target="_self"></a></small></p><p><strong>How do the open-source components fit into X Square Robot’s World Unified Model direction?</strong></p><p>We see these releases as layers of the World Unified Model direction rather than isolated projects. <a href="https://x2robot.com/en/oss" target="_blank">Wall-OSS-0.5</a>, the action model, asks whether an open vision-language-action model can gain directly measurable capability from large-scale pretraining, so it is the capability layer. <span>WALL-WM, the world model, asks how a robot should understand change in the world, shifting from fixed windows to event-level modeling, so it is the representation layer. The data system supplies the interaction data that both of them learn from. </span></p><p>Together they form a loop in which models produce capability, world models organize understanding, and the open-source community drives reproduction and improvement. World Unified Model is the broader architecture those layers support, bringing vision, language, action, and physical prediction together. </p><p>We are releasing these pieces openly because embodied intelligence cannot be solved by one organization; it needs many embodiments, many real tasks, and broad feedback, and the long-term goal is a stack that keeps learning and ultimately moves robots from laboratory demonstrations toward reliable everyday use.</p>]]></description><pubDate>Mon, 13 Jul 2026 10:19:51 +0000</pubDate><guid>https://spectrum.ieee.org/x-square-robot-embodied-ai-stack</guid><category>Home-robots</category><category>Type-sponsored</category><category>Large-language-models</category><category>Embodied-intelligence</category><category>Ai-robots</category><category>Robot-learning</category><dc:creator>​X Square Robot</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/humanoid-robot-folding-laundry-on-a-neatly-made-bed-in-a-sunlit-bedroom.png?id=67111698&amp;width=980"></media:content></item><item><title>Large Tabular Models Excel Where LLMs Fail</title><link>https://spectrum.ieee.org/large-tabular-models-nexus</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/three-people-smiling-while-seated-on-a-couch-in-a-casual-office-environment.jpg?id=67114725&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>The large language models (LLMs) that form the basis of generative AI chatbots such as ChatGPT, Claude, and Gemini can generate uncannily human-like text and images. But these models still struggle with a skill that, ironically, looks at face value to be right in their wheelhouse: analyzing structured data. A new type of generative AI is set to change this situation.</p><p>Although you can get your favorite chatbot to <a href="https://spectrum.ieee.org/ai-math-benchmarks" target="_blank">solve intractable math problems</a>, review dense legal documents, compose a <a href="https://spectrum.ieee.org/ai-music-attribution" target="_blank">catchy pop song</a>, or put together some slick PowerPoint slides, give it anything more than a small table and it doesn’t have a clue what to do.</p><p>For most companies and organizations, the most important data sits in spreadsheets. Whether it’s a bank’s transaction logs, a marketing agency’s website metrics, clinical trial participants’ vital signs, or the vast amount of proton collision information produced at atom smashers like the Large Hadron Collider, structured, row-and-column data runs the world, and LLMs can’t deal with it.</p><p>AI startup <a href="https://fundamental.tech/" rel="noopener noreferrer" target="_blank">Fundamental</a> is pioneering a new type of AI foundation model, known as a large tabular model (LTM), to fill the gap. Fundamental came out of stealth mode on 5 February 2026 with US $275 million in funding and a model called <a href="https://fundamental.tech/nexus" rel="noopener noreferrer" target="_blank">NEXUS</a>, purpose-built for tabular data. Now, the model is being adopted by companies such as Amazon Web Services, while others race to build their own LTMs. </p><h2>Why LLMs struggle with spreadsheets</h2><p>Part of why structured data has garnered less attention is a very human bias, argues <a href="https://www.linkedin.com/in/borisvanbreugel/" rel="noopener noreferrer" target="_blank">Boris van Breugel</a>, a senior AI researcher based in Amsterdam. “People like to see images, videos, and ChatGPT responses,” he says. “But tabular data really lags behind because it’s not fun to look at numbers.” </p><p>Different tabular datasets are also difficult to compare, explains van Breugel, who co-wrote a <a href="https://arxiv.org/abs/2405.01147" rel="noopener noreferrer" target="_blank">prescient position paper</a> on this topic in 2024. Whereas most language has similar semantics, making LLMs well-suited to being trained on vast amounts of text data, van Breugel argues that it is much harder to train a single tabular model on tables with very different variables. </p><p>Additionally, language is sequential by nature (as are music, images, and video). Changing the order of words in a sentence may change or completely destroy its meaning. But the structured data you find in spreadsheets isn’t sequential. You can swap the order of columns or play around with rows, but the underlying factual meaning of the data remains the same.</p><p>This independence from linear order is incompatible with an LLM’s fundamental purpose of predicting the next value in a linear sequence. “With LLMs, even slightly changing the input, you get a different output,” says <a href="https://www.linkedin.com/in/jeremy-fraenkel/" rel="noopener noreferrer" target="_blank">Jeremy Fraenkel</a>, CEO of Fundamental. “That’s fine and actually often desirable for LLMs, but when you’re making a prediction of whether a transaction is fraudulent or not, you want to make sure that the prediction is the same, or deterministic, no matter what.”</p><h2>Developing Fundamental’s LTM</h2><p>Current tabular data solutions are limited to machine learning algorithms, such as <a href="https://xgboost.ai/" rel="noopener noreferrer" target="_blank">XGBoost</a>, that have been around for more than 15 years and are used by organizations globally. These algorithms—called gradient-boosted decision trees—have to be trained and optimized by data scientists over the course of months for each and every use case. In contrast, NEXUS and other emerging LTMs are foundational, leveraging learning amassed from pre-training on diverse databases so that they can be applied across a range of different predictive tasks with minimal bespoke feature engineering or task-specific model building.</p><p>And unlike LLMs, which primarily model sequences of tokens, LTMs model the structure of tabular data directly. They jointly learn from each entry’s numerical value, what it represents, and how it relates to other entries. For example, imagine an entry in a grocery stock inventory table for bananas: The LTM can take in not just the magnitude—say, 500—but the fact that the entry represents the current banana stock quantity, its category (produce), and the statistical properties that link the entry with the rest of the column. This contextual understanding enables more accurate reasoning and prediction over structured data.</p><p>According to Fraenkel, one of Fundamental’s biggest challenges in developing NEXUS was obtaining the right training data. Unlike natural language, which is abundant and broadly uniform in structure, tabular data is relatively hard to find—much of the data is sensitive or proprietary—and diverse. There are very few similarities between, for instance, a biology dataset and a financial one. That combination of factors meant Fundamental needed to invest in building a huge training set.</p><p>“We pre-trained NEXUS on billions of tables using a combination of proprietary datasets acquired through partnerships and licensing, high-quality public and open-source datasets, and data augmentation techniques that expanded the diversity and coverage of our training corpus,” Fraenkel says, though he is keen to point out that NEXUS is not trained on customer data. In fact, it is a confidential computing platform, which means that Fundamental physically cannot access customer data, let alone train on it.</p><p>This feature was most likely a key consideration when in June, Amazon Web Services (AWS) embedded <a href="https://aws.amazon.com/blogs/machine-learning/fundamentals-large-tabular-model-nexus-is-now-available-on-amazon-sagemaker-jumpstart/" rel="noopener noreferrer" target="_blank">NEXUS in Amazon SageMaker</a>, widely considered the default operating system for secure machine learning. This brings NEXUS to many customers’ often sensitive data—a contrasting approach to LLMs, where the data has to be imported to the model.</p><p>“With Amazon, we have a first-party partnership, which means that our model exists as if it’s a native AWS solution,” Fraenkel says. “And over time, the goal is to expand these types of relationships to allow [end users] to really access their data wherever they do their predictions.”</p><h2>The future of data analysis</h2><p>Though Fundamental has taken the lead, at least in enterprise applications, the company is not alone in pursuing foundational LTMs. In March, <a href="https://www.feedzai.com/pressrelease/riskfm-ai-risk-model/" rel="noopener noreferrer" target="_blank">Feedzai</a>, which provides fraud and financial crime prevention services, and<a href="https://www.mastercard.com/global/en/news-and-trends/stories/2026/mastercard-new-generative-ai-model.html" rel="noopener noreferrer" target="_blank"> credit card company Mastercard</a> separately launched similar proprietary technologies focused on finance. Then, in late June, Google launched its own foundational competitor, <a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/" rel="noopener noreferrer" target="_blank">TabFM</a>, trained entirely on hundreds of millions of synthetic datasets. </p><p>And machine learning researchers are not far behind either. <a href="https://arxiv.org/abs/2606.30336" rel="noopener noreferrer" target="_blank">FlexTab</a>, <a href="https://arxiv.org/abs/2602.11139" rel="noopener noreferrer" target="_blank">TabICL</a>, and <a href="https://arxiv.org/abs/2511.15941" rel="noopener noreferrer" target="_blank">iLTM</a> are just three of a raft of LTMs developed by the research community in the past year, all in the pursuit of bringing the success of LLMs to the tabular domain.</p><p>For all involved, the direction of travel is clear. “I would be very surprised if most data processing and analysis is not done through an automated system in the future, whether that’s an LLM, an LTM, or some combination,” van Breugel says. “Most people don’t necessarily like to do data analysis, and these systems will be able to do it a lot better.”</p><p>Fraenkel agrees. “I see the relationship between LLMs and LTMs as being a bit like the human brain: The left side is good at reasoning and understanding and summarizing text, and the right side is really good at understanding numbers and statistics and patterns,” he says. “But it’s when you combine both of those that you really get something much more powerful.”</p>]]></description><pubDate>Thu, 09 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/large-tabular-models-nexus</guid><category>Data-analytics</category><category>Llms</category><category>Foundation-models</category><category>Databases</category><dc:creator>Benjamin Skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/three-people-smiling-while-seated-on-a-couch-in-a-casual-office-environment.jpg?id=67114725&amp;width=980"></media:content></item><item><title>AI Models Overthink Problems—and It’s a Security Risk</title><link>https://spectrum.ieee.org/ai-reasoning-models-security-risk</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-dozen-security-lasers-pointed-in-the-wrong-direction-around-a-password-thus-ironically-creating-a.jpg?id=67107951&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>Large language models (LLMs) that can think through problems step-by-step have significantly increased the scope of tasks that AI can tackle. But new research suggests these reasoning capabilities also introduce a critical vulnerability that could allow attackers to slow these systems to a crawl.</p><p>While earlier generations of LLMs would immediately produce a response to a user’s request, today’s most advanced models generate an internal monologue where they break down the problem into steps and reason about the best way to tackle it before providing an answer. This has allowed AI to tackle increasingly complex problems, particularly in areas like coding and <a href="https://spectrum.ieee.org/ai-in-mathematics" target="_self">math</a>.</p><p>However, <a href="https://arxiv.org/abs/2412.21187" rel="noopener noreferrer" target="_blank">previous</a> <a href="https://spectrum.ieee.org/reasoning-in-ai" target="_self">research</a> has shown that these models are susceptible to sometimes producing excessively long streams of reasoning that do little to boost performance, a phenomenon known as “overthinking.” In research <a href="https://icml.cc/virtual/2026/poster/62234" rel="noopener noreferrer" target="_blank">presented this week</a> at the <a href="https://icml.cc/" rel="noopener noreferrer" target="_blank">International Conference on Machine Learning 2026</a>, in Seoul, researchers from Zhejiang University and e-commerce giant Alibaba in China demonstrate that they can deliberately induce overthinking by subjecting models to logically inconsistent prompts. The result is a form of denial-of-service attack on commercial AI models.</p><h2>Evolutionary Prompt Attack on LLMs</h2><p>The team has developed an <a href="https://spectrum.ieee.org/evolutionary-ai-coding-agents" target="_self">evolutionary algorithm</a> that corrupts the logical structure of prompts, causing models to spiral into overthinking as they attempt to reason through fundamentally unsolvable problems. Generating longer responses costs more and increases the load on a model provider’s servers, so if done at scale, the researchers say, this could significantly degrade the experience of legitimate users. The attack was effective against reasoning models from leading AI companies including DeepSeek-R1, Alibaba’s Qwen3-Thinking, OpenAI’s GPT-o3, and Google’s Gemini 2.5 Flash, and resulted in outputs up to 26 times as long as standard responses on a standard math benchmark.</p><p>“Across multiple datasets and reasoning models, our method substantially amplifies the output length,” Wei Cao, a master’s student at Zhejiang University, wrote in an email to <em><em>IEEE Spectrum</em></em>. “Our results suggest that overthinking is not an isolated phenomenon specific to individual models, but rather a shared vulnerability among modern reasoning models.”</p><p>The team’s approach builds on <a href="https://arxiv.org/pdf/2504.06514" rel="noopener noreferrer" target="_blank">previous research</a> from another group of researchers that showed reasoning models tend to overthink when faced with a question in which a key premise has been removed—such as asking how far someone who walks 10 miles a day covers in total without specifying how many days they walked for. Rather than identifying that the problem is unsolvable, models often engage in extended but ultimately fruitless reasoning loops in an attempt to answer the question.</p><p>Taking the idea a step further, the authors took 940 problems from three math benchmark datasets and used an LLM to break down their logical structure into a set of premises and a final question. The genetic algorithm then jumbled these up using a variety of “mutations,” including swapping premises between problems, adding extra premises to problems, deleting existing premises from problems, and swapping the final questions between two sets of premises.</p><p>After each round of mutations, the problems are scored on how many words they cause a target model to output and also whether they increase the frequency of specific linguistic markers of overthinking—words like “but,” “wait,” “maybe,” or “alternatively.” The problems that scored highest on both measures are retained, and the remaining ones are jumbled up again, and this process is repeated for five generations. Crucially, the approach doesn’t require access to the internals of a model and can generate malicious prompts by simply querying the target, which makes it possible to attack closed-source commercial services, says Cao.</p><h2>Overthinking Vulnerability in AI Models</h2><p>The researchers found that the approach consistently led to outputs several times longer than those generated by the unmodified questions for the reasoning models they tested it on. The biggest jump came from DeepSeek-R1 on the <a href="https://arxiv.org/abs/2103.03874" rel="noopener noreferrer" target="_blank">MATH dataset</a>, which is made up of problems from high school math competitions, where the maximum output was 26.1 times as long as the longest response the model provided to unaltered questions. While the main thrust of the research was focused on math problems, the authors also tested it on coding, scientific reasoning, and dialogue challenges, and observed significant jumps in output length in all three.</p><p>One challenge for the approach is that developing the malicious prompts requires repeated queries to expensive reasoning models, which Cao admitted could limit its cost-effectiveness. However, the researchers also demonstrated that when they used a smaller, cheaper model to generate the malicious prompts, they were still able to induce the target models to produce outputs several times longer than normal. This ability to transfer malicious prompts between models significantly increases the attack’s feasibility, Cao wrote.</p><p>However, he pointed out that the goal of the research is not to develop a practical DoS attack on reasoning models. Factors like the providers’ pricing model, rate limiting policies, context window size, and existing defenses could all impact how effective the approach is. The intention is instead to highlight these models’ vulnerability to logically inconsistent prompts so that providers can attempt to mitigate the problem.</p><p>“Our objective is not to demonstrate that large-scale attacks can be launched at negligible cost, but rather to establish that this attack surface exists,” he wrote. “Our results indicate that the vulnerability represents a realistic security concern.”</p>]]></description><pubDate>Wed, 08 Jul 2026 11:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-reasoning-models-security-risk</guid><category>Llms</category><category>Artificial-intelligence</category><category>Denial-of-service</category><category>Cybersecurity</category><dc:creator>Edd Gent</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/conceptual-illustration-of-a-dozen-security-lasers-pointed-in-the-wrong-direction-around-a-password-thus-ironically-creating-a.jpg?id=67107951&amp;width=980"></media:content></item><item><title>What Makes AI Art Worth Collecting?</title><link>https://spectrum.ieee.org/ai-art-market</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-small-human-silhouette-against-a-video-installation-wall-displaying-an-abstract-jumble-of-textures-inside-of-a-rectangular-box.jpg?id=67103138&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>In May, an anonymous artist who goes by SHL0MS on X posted that he had used AI to <a href="https://x.com/SHL0MS/status/2054280631807316329" rel="noopener noreferrer" target="_blank">generate an image inspired by Claude Monet</a> and asked people to weigh in on how it missed the mark. More than 600 responses called out issues, saying the colors were off, the depth was all wrong, and that AI didn’t understand how light worked.</p><p>SHL0MS then revealed that the image was of a real Monet, one of around 250 variations of water lilies the artist had painted in his lifetime. He had simply downloaded a high-resolution image from Wikimedia and cropped out the signature. He minted the exchange as an <a href="https://en.wikipedia.org/wiki/Non-fungible_token" rel="noopener noreferrer" target="_blank">NFT</a> (a unique digital collectible recording ownership of the work), titled it “Inferior Image,” and sold it for just over US $40,000 after 28 bids.</p><p>The stunt exposed how charged the conversation around AI art has become, and how quick people are to dismiss anything AI-generated as slop—even when it’s not. Yet even as those arguments continue, a market for AI-generated art has begun to form anyway. It’s fragmented and contested, but bigger than most people realize.</p><p>Jediwolf, an anonymous collector who says he has spent more than 20 years acquiring digital and AI art, was watching the experiment unfold in real time on X. He had never interacted with SHL0MS before, but when the NFT went up for auction he made a bid and won. “I was buying a unique moment in time,” he says, “captured by an artist and preserved as a token.”</p><p>The Monet was not AI art, but most of what Jediwolf buys is. One of Jediwolf’s digital collections, which he calls <a href="https://opensea.io/UnderTheGAN/galleries" rel="noopener noreferrer" target="_blank">UnderTheGAN</a>—a play on GANs, or generative adversarial networks, the AI technology that preceded today’s diffusion models—comprises roughly 100 works valued at around $72,000, focused on early AI art from 2015 to 2020, before the medium went mainstream. He describes his role as part collector, part researcher, part curator, trying to document a fast-moving field.</p><p>“A decade ago, digital art was often treated as peripheral to the ‘serious’ art world,” he says. “Today, it is increasingly difficult to separate contemporary culture from the internet.”</p><h2>AI Art Moves Into Museums</h2><p>The market for AI art extends beyond NFTs: AI-generated pieces are also finding their way into physical installations. Last month saw the opening of <a href="https://dataland.art/" rel="noopener noreferrer" target="_blank">Dataland</a>, the world’s first generative AI museum, in downtown Los Angeles. It was spearheaded by <a href="https://refikanadol.com/" rel="noopener noreferrer" target="_blank">Refik Anadol</a>, a digital artist who has built a career out of transforming data into large-scale immersive experiences. The <a href="https://dataland.art/exhibitions/machine-dreams-rainforest" rel="noopener noreferrer" target="_blank">opening exhibition</a> has pieces that use data that Anadol collected from rainforests around the world, with real-time weather information from 16 rainforests feeding into all five galleries. In three of the rooms, the imagery also shifts in response to visitors’ own biometric data, tracked by bracelets they wear. </p><p>Like any museum it sells tickets, ranging from $49 to $79, and has a gift shop. This shop, however, uses visitors’ biometric data collected during their visit to generate a unique design printed on a T-shirt. For $15,000, a robotic painting system called Qualia creates a one-of-a-kind canvas from that same data, painted once a day, with a waiting list already forming. A founding collection of <a href="https://www.instagram.com/reels/DMLA4BbPzdN/" rel="noopener noreferrer" target="_blank">1,000 AI data sculptures</a> that evolve based on environmental data from global rainforests sold out in 34 minutes at $5,000 each.</p><p>The system running it all, which Anadol calls the <a href="https://dataland.art/about/large-nature-model" rel="noopener noreferrer" target="_blank">Large Nature Model</a>, was trained on more than 500 million nature images representing 2.2 million species, gathered through field expeditions to 16 rainforests and partnerships with institutions including the Smithsonian and the Cornell Lab of Ornithology.</p><p>For Anadol, AI art requires a different kind of transparency than any medium that came before it. Because commercial AI tools have shaped how most people understand the technology, artists working with it seriously have to be more open about their process than painters or photographers ever did.</p><p>“For AI art, we have to know where the data comes from, we have to know which model is trained and how it’s trained,” he says. “We can’t just think about authenticity and uniqueness if a service and product is the fundamental layer of the artwork.”</p><p>The reviews for Dataland have mostly been positive, with one critic calling it the <a href="https://news.artnet.com/art-world/refik-anadol-dataland-review-2-2781630" rel="noopener noreferrer" target="_blank"><em>Citizen Kane</em></a> of immersive experiences. But Anadol is used to a more divided reception. His <a href="https://www.moma.org/collection/works/442077?artist_id=134464&page=1&sov_referrer=artist" rel="noopener noreferrer" target="_blank">2022 installation at MoMA</a>—a 7-by-7-meter screen of AI-generated fluid forms with shifting colors and sounds—drew 3 million visitors and entered the permanent collection, even as <em><em>New York Magazine</em></em> called it “<a href="https://www.vulture.com/article/jerry-saltz-moma-refik-anadol-unsupervised.html" rel="noopener noreferrer" target="_blank">a massive techno lava lamp</a>.” </p><p>Anadol sees the skepticism as nothing new, just the latest version of a resistance that has greeted all new media. “Every art form has gone through similar cycles of denial,” he says. “We are living in a renaissance that started 10 years ago, and I just don’t think everyone is aware of it yet.”</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Claude Monet\u2019s impressionist painting of water lilies." class="rm-shortcode" data-rm-shortcode-id="01cbdd714ad0eb733cb5aa2486f64837" data-rm-shortcode-name="rebelmouse-image" id="8b5ef" loading="lazy" src="https://spectrum.ieee.org/media-library/claude-monet-u2019s-impressionist-painting-of-water-lilies.jpg?id=67115256&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">An anonymous artist cropped the signature from this Claude Monet painting and presented it online as AI-generated.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Claude Monet</small></p><h2>Who Is Buying AI Art?</h2><p>The broader market data points in multiple directions at once. According to the <a href="https://theartmarket.artbasel.com/?gad_source=1&gad_campaignid=23654247823&gbraid=0AAAABDFkwHYbbNWTJuqRqXWS3vz6fL-kO" rel="noopener noreferrer" target="_blank"><em>Art Basel and UBS Art Market Report 2026</em></a>, digital art’s share of sales nearly tripled between 2024 and 2025, and just over half of all fine art collectors surveyed had purchased a digital artwork in 2025, making it the third most popular category after painting and sculpture (the report does not break out AI art specifically).</p><p>Meanwhile, Christie’s <a href="https://www.theartnewspaper.com/2025/09/09/amid-a-slump-for-nfts-christies-closes-digital-art-department" rel="noopener noreferrer" target="_blank">shuttered its pioneering digital art department</a> in September, folding digital works back into its broader contemporary sales after none of its dedicated auctions broke $400,000.</p><p>The most data-rich window into buyer behavior comes from a less glamorous corner of the market. After one major stock image platform allowed AI-generated images, monthly sales jumped 80 percent, according to <a href="https://www.gsb.stanford.edu/faculty-research/faculty/samuel-goldberg" target="_blank">Samuel Goldberg</a>, an economist at Stanford Graduate School of Business who <a href="https://www.gsb.stanford.edu/faculty-research/working-papers/generative-ai-equilibrium-evidence-creative-goods-marketplace" rel="noopener noreferrer" target="_blank">published a research paper</a> about the shift. Traditional contributors began leaving the platform as generative images flooded in, and creators using AI tools rushed to fill the gap. </p><p>“It looks like consumers like generative AI,” Goldberg says, “and it seems like nongenerative artists could be getting crowded out of the market.” Stock images are essentially a commodity version of art, according to Goldberg, and because image-generating models are already very good at producing them, what’s happening there may be a preview of what’s coming for other creative goods markets—including fine arts—as the technology improves.</p><p>Artists are typically among the first to test the limits of a new technology; early adopters have created AI art <a href="https://spectrum.ieee.org/ai-art-whitney-museum" target="_self">since the 1970s</a>. What’s new now is the ability for anyone to generate an image in seconds with a text prompt. That, according to <a href="https://www.linkedin.com/in/christiane-paul-curator/" rel="noopener noreferrer" target="_blank">Christiane Paul</a>, curator of digital art at the Whitney Museum of American Art, is not the same thing at all. What fills those stock-image platforms, and what most people encounter when they think of AI art, does not qualify as art.</p><p>True AI art, Paul says, is a subcategory of digital art that uses artificial intelligence as both a tool and a medium, engaging with it practically and conceptually, doing things like training custom models, building extensions, and layering control systems. <span>“A visual created by a prompt is not art,” she says. What serious AI artists are actually doing is much more than typing a few words into <a href="https://spectrum.ieee.org/openai-dall-e-2" target="_blank">DALL-E</a>.</span></p><p>Far from the shortcut most people assume, working seriously with AI as an artistic medium is, by her account, brutally hard. Every artist she talks to says the same thing. “It is much, much harder than a paintbrush to handle,” she says. “You are literally communicating with a system with a completely different logic.”</p><p><em><span><em>Thanks to </em></span></em><a href="http://bubblemaps.io" target="_blank"><em><em>bubblemaps.io</em></em></a><em><em> for its research assistance on the NFT market.</em></em></p>]]></description><pubDate>Tue, 07 Jul 2026 14:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/ai-art-market</guid><category>Ai-art</category><category>Generative-ai</category><category>Digital-art</category><category>Blockchain</category><dc:creator>Jackie Snow</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-small-human-silhouette-against-a-video-installation-wall-displaying-an-abstract-jumble-of-textures-inside-of-a-rectangular-box.jpg?id=67103138&amp;width=980"></media:content></item><item><title>Small AI Models Gain Traction Around the World</title><link>https://spectrum.ieee.org/small-language-models-ai-pharmaceuticals</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-middle-aged-man-monitors-his-younger-colleague-through-a-simulator-lab-window-as-he-conducts-heart-rhythm-experiments-on-a-med.jpg?id=67101131&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>One morning in 2019, <a href="https://adebayoalonge.com/" rel="noopener noreferrer" target="_blank">Adebayo Alonge</a> was in a Cape Town hotel room, preparing to demonstrate his startup’s AI answer to a serious problem in African health care: counterfeit medication, which kills thousands of people across the continent every year.</p><p>The <a href="https://rxall.net/rxscanner/" rel="noopener noreferrer" target="_blank">RxScanner</a> is a handheld spectrometer that scans a pill with infrared light, then sends the item’s molecular profile to an AI model equipped with a pharmaceutical database. In seconds, the AI identifies the medication from its molecular profile—or reports that it’s phony.</p><p>Pharmacies were using the system in more than a dozen countries, including Ghana, Kenya, Myanmar, and Alonge’s native Nigeria. But that morning in South Africa, it didn’t work. “I was shocked,” Alonge says.</p><p>The spectrometer connected to the AI model—but the data center was 14,000 kilometers away and bandwidth was limited. “Our server was in the United States, and just to get the result of a single scan was taking me over 5 minutes.”</p><p>So Alonge immediately asked his engineers to shrink the AI model down to a smaller, low-power, unconnected version that could run entirely on his Android phone. They produced it 2 hours later, and that saved the demo.</p><p>More importantly, the work birthed a new version of his device, which can authenticate a pill in places without broadband, computers, or even reliable electricity. It also turned Alonge into an advocate for this kind of “small AI.”</p><h2>Small AI for Global Health Care Access</h2><p><a data-linked-post="2674052790" href="https://spectrum.ieee.org/small-language-models" target="_blank">Small AI</a> is a far cry from wealthy nations’ colossal large language models (LLMs), hyperscale data centers, multibillion-dollar investments, and <a data-linked-post="2650276080" href="https://spectrum.ieee.org/interview-max-tegmark-on-superintelligent-ai-cosmic-apocalypse-and-life-3-0" target="_blank">debates about AI consciousness</a>. But for millions of people around the world, the only AI that matters, and often the only kind available, is small. (According to a <a href="https://www.worldbank.org/en/publication/dptr2025-ai-foundations" rel="noopener noreferrer" target="_blank">World Bank Report</a> issued in November, only 0.7 percent of internet users in the world’s poorest countries have used ChatGPT, compared to a quarter of all internet users in the most developed nations.)</p><p>“Most people are discussing AI from the LLM/generative side. But that needs a lot of computing power, electricity, massive data, and skilled people to manage it,” Ajay Banga, president of the World Bank, <a href="https://www.ndtv.com/world-news/small-ai-is-indias-secret-weapon-ajay-banga-tells-ndtv-at-davos-10833458" rel="noopener noreferrer" target="_blank">said last January at the World Economic Forum, in Davos.</a> “Outside the developed world, other than maybe India and China, very few countries have that combination.”</p><p>By contrast, small AI can deliver useful, even life-saving services to people in areas that have none of those things, Banga said. In India, where the government’s AI plans call for more development of small AI, many such systems are working for farmers.</p><p>For example, a <a href="https://www.science.org/doi/epdf/10.1126/science.adw7713" rel="noopener noreferrer" target="_blank">drone-based system developed by Bala Murugan and colleagues</a> at the Vellore Institute of Technology, in India, takes photos of cashew plants and quickly identifies those with splotches that indicate disease. All the processing takes place on the drone itself, so there’s no need for a computer on-site, nor for a connection to a central server.</p><p>Using small language models trained for a specific problem, and sometimes running on cheap, low-power devices, other small-AI implementations have been developed to identify <a href="https://universe.roboflow.com/juan-abedala/deteccion-hormigas-cortadoras" rel="noopener noreferrer" target="_blank">ant infestations in a Uruguayan vineyard</a>, <a href="https://dl.acm.org/doi/fullHtml/10.1145/3524458.3547258" rel="noopener noreferrer" target="_blank">detect the presence of malaria-carrying mosquitoes in a number of nations</a>, and <a href="https://link.springer.com/chapter/10.1007/978-3-031-49407-9_63" rel="noopener noreferrer" target="_blank">run electrocardiograms from an Arduino device in parts of Brazil</a> that lack access to more complex equipment.</p><p>“This is the most important area in AI nowadays,” says <a href="https://www.linkedin.com/in/marcelo-jose-rovai-brazil-chile/" rel="noopener noreferrer" target="_blank">Marcelo José Rovai</a>, a professor at the Institute of Engineering and Information Systems at the Federal University of Itajubá, in Brazil, who was involved in all three projects. “It’s growing very fast.”</p><h2>Low-Power, Small-AI Models on Devices</h2><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Two development boards and an IoT development platform running simultaneously on a lab table." class="rm-shortcode" data-rm-shortcode-id="d796471a2b94d32f1f4d226a5daa6d0b" data-rm-shortcode-name="rebelmouse-image" id="b0c16" loading="lazy" src="https://spectrum.ieee.org/media-library/two-development-boards-and-an-iot-development-platform-running-simultaneously-on-a-lab-table.jpg?id=67101154&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Small AI models can run on a variety of low-power devices, including [from left to right] an Arduino Nano 33 BLE Sense, a Seeed Wio Terminal, and an Arduino Portenta.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Moez Altayeb</small></p><p>For Alonge, Rovai, and other advocates, small AI is not just “a promising trend,” as that November World Bank report calls it. It may be, in the long term, the form of AI that will touch the most lives and remain sustainable after some of the giant models become too costly for most users.</p><p>“I think the future of AI is not like one giant model, at a center. I think it’s millions of small, precise models deployed at the edge, each one solving like a specific problem, a specific context,” Alonge says. This is partly because much of humanity—including people in parts of rich countries as well as the developing world—lives without access to cutting-edge frontier models. But, he says, it’s also because those models are not sustainable.</p><p>“If someone is not subsidizing it, most people will not be able to afford those models. So those of us who are said to be small-AI developers are the ones who will have to build for the majority of the world,” Alonge says.</p><p>There is no strict definition of “small AI,” but people often use the term for language models with at most a few billion parameters. (Compare that to cutting-edge models, which can include more than a trillion.) That’s small enough to run directly on a phone or a Raspberry Pi. That’s what allows these applications to run on devices without a connection to a data center and use only a few watts of power, often supplied by a battery or a solar panel.</p><p>Despite their small footprint, these models aren’t fundamentally different technology from that of gigantic AI models, Rovai says. Many instances of small language models were created the same way the phone-based version of Alonge’s pharmaceuticals scanner was—by “pruning” large models, or removing the parameters that weren’t involved in the task. The result is a system that’s less capable generally but still very good at the specific job it was pruned for, Rovai says.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="A small obround device with a simple design featuring one button, a lid and four small indicator lights." class="rm-shortcode" data-rm-shortcode-id="8e7efd3d06670a794f7fdfc68f281a6f" data-rm-shortcode-name="rebelmouse-image" id="9d4a2" loading="lazy" src="https://spectrum.ieee.org/media-library/a-small-obround-device-with-a-simple-design-featuring-one-button-a-lid-and-four-small-indicator-lights.jpg?id=67101162&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">A lighter version of RxAll’s RxScanner spectrometer sends its results to an AI model run locally on a phone to check that a drug’s molecular signature is genuine.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">RxAll</small></p><p>Other small models are created by “distillation.” They are trained to mimic a large model, until their performance approaches that of their “teacher,” Rovai says. In other cases, a larger model’s precision is reduced, for example, so that a model run on 32-bit architecture can run on 8-bit designs. In situations where the machine learning application is being used to classify data or predict patterns (like an ant infestation), it’s trained from the beginning on a small device, not derived from a larger model at all. </p><p>Running all these small, specialized systems is becoming easier, Rovai says, for two reasons.</p><p>The first reason is that hardware is getting better and more capable while using less power, he says. This means more and more phones can run small AI—especially those equipped with neural processing units, which are specialized chips that handle AI tasks like facial recognition and changing the brightness, shadows, or contrast in a photo.</p><p>In 2025, slightly more than a third of all smartphones shipped worldwide were capable of running generative AI, and that figure will reach 45 percent by the end of this year, <a href="https://counterpointresearch.com/en/insights/genai-smartphone-share-to-rise-to-45-percent-of-global-shipments-in-2026" target="_blank">according to the technology research firm Counterpoint</a>. By the end of next year, slightly more than half of all smartphones will be able to run a small AI model.</p><p>The second reason Rovai cites is the shrinking footprint of language models. Both Google DeepMind’s <a href="https://deepmind.google/models/gemma/gemma-4/" target="_blank">Gemma 4</a> (released in April) and Alibaba’s <a href="https://qwen.ai/blog?id=qwen3.5" rel="noopener noreferrer" target="_blank">Qwen 3.5 </a>are “fantastic” for small AI, Rovai says. Both models are “open weight,” meaning users can adjust the connections between parameters to suit their needs. This makes it easy, for example, “to take a lot of data from, say, the milk industry and retrain the model specifically on that,” Rovai says.</p><p>Rovai illustrated these reasons on a Zoom call, using one of his most recent experiments. Holding up a device, he says, “This is the new Arduino UNO Q—a US $50 device with a Qualcomm chipset. I’m running a language model here, which collects data from sensors and analyzes that data to detect tiny pools of water where mosquitoes might be breeding. It takes 3 watts to run it.”</p><h2>Support for Small-AI Development</h2><p>Convinced that millions of people are already benefiting from these kinds of applications, the World Bank now actively promotes small AI with grants, mentorship programs, financing, technical advice, and models of government policies that are friendly for small-AI development. For example, in Rwanda, the World Bank is backing a government program to help low-income households get devices that can run AI.</p><p>All that said, no one claims that large language models are going away entirely. To create a generative AI that can run on a phone or other small device requires the architectural insights, data processing, and results of a larger model, Rovai says. “We need the big models to create these smaller models.” </p><p>And for all that small AI can benefit people without access to big AI, the technology can’t solve the larger problems of development and digital inequality, Alonge says. Implementing small AI won’t allow nations to escape the challenge of creating an ecosystem to support AI: reliable power, a supply chain that works, and an educational system that develops the talents needed to create AI tools.</p><p>Though his drug-scanning system can run for days on a phone with no connection, “you still want to be able to enable periodic syncing for updates with new signatures for the medications and analytics,” Alonge says. “And even when you are using batteries, reliable power is important. That phone battery is not going to last forever.”</p><p>In many parts of the world, the future of small AI isn’t assured, he says. “It works, and many places will eventually need to use it. The question is whether or not the political actors are wise enough to invest in infrastructure to support it long term.”</p>]]></description><pubDate>Mon, 06 Jul 2026 16:06:23 +0000</pubDate><guid>https://spectrum.ieee.org/small-language-models-ai-pharmaceuticals</guid><category>Small-language-models</category><category>Artificial-intelligence</category><category>Llms</category><dc:creator>David Berreby</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-middle-aged-man-monitors-his-younger-colleague-through-a-simulator-lab-window-as-he-conducts-heart-rhythm-experiments-on-a-med.jpg?id=67101131&amp;width=980"></media:content></item><item><title>AI’s Volatile Power Use Quietly Tests Grid Limits</title><link>https://spectrum.ieee.org/data-centers-grid-instability</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/aerial-view-of-a-large-industrial-complex-near-housing-and-power-lines-in-autumn.jpg?id=67080460&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>The rapid expansion of artificial intelligence infrastructure is typically framed as an energy problem. <a href="https://spectrum.ieee.org/5gw-data-center" target="_self">Data centers</a> are projected to consume a growing share of global electricity demand: The <a href="https://www.iea.org/about" rel="noopener noreferrer" target="_blank">International Energy Agency</a><a href="https://www.iea.org/reports/electricity-2024" rel="noopener noreferrer" target="_blank"> estimates</a> they could account for 3 to 4 percent of total global consumption within this decade.</p><p>Utilities are already adjusting long-term forecasts to accommodate anticipated growth from hyperscale facilities and high-density compute clusters.</p><p>This framing captures scale. It misses behavior.</p><p>The emerging issue is not simply how much power large-scale compute systems consume, but how increasingly dense and synchronized computational workloads are beginning to alter the operating characteristics of the electrical grid itself through increasingly unpredictable demand that varies rapidly in both time and location, creating new operational challenges for grid operators.</p><h2>AI’s Capricious Energy Needs</h2><p>Traditional grid planning assumes relatively predictable demand behavior. Industrial, commercial, and residential loads generally follow established profiles that can be forecast with reasonable accuracy. Even substantial demand growth has historically been manageable through reserve planning, transmission upgrades, and demand management programs.</p><p>Large-scale compute infrastructure introduces a different class of electrical load. Training—the computational task of making AI models—tends to be highly synchronized across clusters of GPUs, TPUs, and specialized accelerators operating in parallel, computationally dense, and relatively scheduled. Inference—the process of actually using those models—is generally more distributed and user-driven, making demand less predictable both in time and location. Both differ materially from traditional industrial demand profiles, though for different reasons. Unlike many conventional industrial processes, these workloads can ramp rapidly depending on model training cycles, distributed compute coordination, and workload scheduling strategies.</p><p>From the perspective of the grid, this is not simply higher demand. It is more abrupt demand. High-density compute workloads can produce substantial step changes in electricity consumption over extremely short intervals, including rapid fluctuations occurring within milliseconds. Data-center operators are already deploying mitigation technologies, including batteries, power-conditioning systems, and <a href="https://spectrum.ieee.org/supercapacitor-2671883490" target="_self">supercapacitors</a>. Collectively, however, data centers’ rapid load changes can place additional stress on backup-generation reserves, systems that adjust supply as demand changes, frequency-control mechanisms that maintain grid stability, and local transmission infrastructure.</p><p>Compute-related variability differs from the intermittency introduced through renewable energy integration. Wind and solar variability originate primarily on the supply side and is tied to environmental conditions. Compute-related variability emerges on the demand side, driven by workload synchronization, scheduling behavior, and computational intensity. The interaction between increasingly dynamic supply and demand conditions introduces additional uncertainty into forecasting, reserve management, congestion planning, and balancing operations.</p><p>Research organizations including the <a href="https://www.energy.gov/ea/national-renewable-energy-laboratory" rel="noopener noreferrer" target="_blank">National Renewable Energy Laboratory</a> have <a href="https://www.nrel.gov/grid/" rel="noopener noreferrer" target="_blank">emphasized</a> the growing complexity associated with integrating highly dynamic resources into modern grid operations.</p><h2>Location, Location, Location</h2><p>The issue becomes more significant when compute activity is geographically concentrated. Large-scale data centers tend to cluster in regions with favorable conditions such as fiber connectivity, access to markets, tax incentives, and historically low electricity costs. Northern Virginia, often referred to as Data Center Alley, remains the most prominent example. The region hosts the world’s <a href="https://www.vedp.org/industry/data-centers" rel="noopener noreferrer" target="_blank">largest</a> concentration of data centers and carries a substantial share of global internet traffic.</p><p>Utilities operating in these regions have already identified data-center growth as a primary driver of future load expansion. Virginia-based electricity supplier <a href="https://www.dominionenergy.com/" rel="noopener noreferrer" target="_blank">Dominion Energy</a>, for example, has repeatedly highlighted hyperscale demand growth in its integrated resource <a href="https://www.dominionenergy.com/about/our-company/irp" rel="noopener noreferrer" target="_blank">planning documents</a>.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Aerial view of sprawling data center and warehouse complex surrounded by greenery" class="rm-shortcode" data-rm-shortcode-id="31cd4c333620851cf421241d2ad207ab" data-rm-shortcode-name="rebelmouse-image" id="d0616" loading="lazy" src="https://spectrum.ieee.org/media-library/aerial-view-of-sprawling-data-center-and-warehouse-complex-surrounded-by-greenery.jpg?id=67080499&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Virginia has seen one of the largest data center buildouts worldwide. Here, Amazon Web Services and Iron Mountain data centers dominate the landscape in Manassas, Va. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Nathan Howard/Bloomberg/Getty Images</small></p><p>A sudden increase in electricity consumption within a constrained geographic area can stress substations, transmission corridors, and local balancing operations even if the broader grid maintains sufficient aggregate capacity. This creates localized reliability challenges that are not always visible through system-wide demand metrics alone.</p><p>Thermal management systems further intensify these effects. Cooling infrastructure in high-density compute facilities must respond dynamically to changing workloads. As processing intensity <a href="https://www.dominionenergy.com/about-us/electric-projects-and-programs/integrated-resource-plan" target="_blank">rises</a>, cooling demand rises as well, often nonlinearly. This coupling between compute and thermal systems means that fluctuations in workload can propagate through multiple layers of facility power consumption simultaneously.</p><p>High-density compute clusters may also introduce power-quality concerns at the local level. Large concentrations of accelerators, switching power supplies, and high-frequency compute equipment can generate harmonics and nonlinear load behavior that place additional stress on distribution infrastructure. While modern facilities incorporate mitigation technologies, the scale and concentration of next-generation compute facilities may require utilities and operators to revisit assumptions surrounding localized power conditioning, harmonics management, and infrastructure resilience. These conditions can also contribute to short-duration electrical transients that place additional stress on localized infrastructure and power-conditioning systems.</p><h2>Regulations Need Updating</h2><p>Part of the challenge is that many existing regulatory and operational frameworks were designed around relatively stable industrial demand profiles. Large rapidly fluctuating loads have historically been constrained because abrupt cycling can complicate balancing operations, increase stress on transmission equipment, and reduce predictability in system operations. High-density compute clusters do not fit neatly within those assumptions.</p><p>This creates pressure for both operational adaptation and regulatory reassessment.</p><p>Demand-response mechanisms may allow certain compute workloads to be shifted or curtailed during periods of system stress. Data-center operators are exploring <a href="https://spectrum.ieee.org/distributed-inference-data-centers" target="_self">flexible scheduling</a>, battery storage, and <a href="https://spectrum.ieee.org/ai-data-centers" target="_self">behind-the-meter generation</a>. Grid operators, meanwhile, are evaluating planning frameworks and interconnection approaches for increasingly large flexible loads.</p><p><a href="https://www.ercot.com/" target="_blank">The Electric Reliability Council of Texas</a> (ERCOT), for example, has <a href="https://www.ercot.com/gridinfo/resource" rel="noopener noreferrer" target="_blank">publicly acknowledged</a> the growing implications of large flexible loads, including data centers, for long-term grid planning and operational stability. Interconnection queues across the United States continue to <a href="https://emp.lbl.gov/queues" rel="noopener noreferrer" target="_blank">expand significantly</a>, reflecting mounting pressure on both generation and transmission infrastructure. Grid expansion timelines, however, are measured in years rather than quarters.</p><p>This creates a structural mismatch. Compute infrastructure can scale rapidly. Electrical infrastructure generally cannot.</p><p>The broader implication is that large-scale compute infrastructure is not simply another industrial load category. It represents a shift in the temporal and spatial characteristics of electricity demand itself.</p><p>Framing the issue solely in terms of aggregate energy consumption risks overlooking these second-order operational effects. Capacity expansion alone does not fully address rapid ramping behavior, synchronization, localized congestion, transient instability, reserve compression, or increasingly demanding load-following requirements.</p><p>The challenge is not just how much electricity these systems consume. It is how they are beginning to change the operating conditions of the grid itself. The call is not to slow AI development but to recognize that hyperscale computing represents a new category of electrical demand. As AI infrastructure continues to scale, planning frameworks may need to account not only for total energy consumption but also for demand volatility, synchronization effects, and geographic concentration. Grid resilience will increasingly depend on understanding how these facilities consume power, not simply how much power they consume.</p>]]></description><pubDate>Fri, 03 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/data-centers-grid-instability</guid><category>Data-centers</category><category>Artificial-intelligence</category><category>Electrical-grid</category><category>Demand-response</category><dc:creator>Matt Hasan</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/aerial-view-of-a-large-industrial-complex-near-housing-and-power-lines-in-autumn.jpg?id=67080460&amp;width=980"></media:content></item><item><title>As AI Reshapes Global Energy Systems, Melbourne Leads Through Engineering Collaboration</title><link>https://spectrum.ieee.org/ai-energy-systems-melbourne</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/glowing-digital-network-map-of-australia-and-surrounding-asia-pacific-region.png?id=66945530&width=1245&height=700&coordinates=0%2C0%2C0%2C1"/><br/><br/><p><em>This article is brought to you by <a href="https://www.melbournecb.com.au/?utm_source=ieee&utm_medium=editorial&utm_campaign=discover-melbourne-2026&utm_term=maveric&utm_content=link" rel="noopener noreferrer" target="_blank">Melbourne Convention Bureau (MCB)</a> supported by <a href="https://businessevents.australia.com/en" target="_blank">Business Events Australia</a>.</em></p><p><span>As artificial intelligence accelerates global demand for compute, a parallel constraint is emerging with equal urgency: energy.</span></p><p>From hyperscale data centers to electrified industries, AI is driving a step change in electricity demand. This is not a future challenge, it is a present, system-level issue requiring coordinated action across energy, infrastructure, and engineering disciplines.</p><p>Around the world, the question is no longer whether AI will scale, but whether energy systems can scale with it.</p><p>Melbourne, Australia is moving beyond participation to become a globally connected leader helping define how these challenges are addressed.</p><h2>A national challenge with global implications</h2><p>Australia’s ambition to lead in artificial intelligence is sharpening focus on the infrastructure required to support it. Data centers are projected to account for up to <a href="https://www.cefc.com.au/media/hs5ner3s/getting-the-balance-right-data-centres-and-the-energy-transition-full-report.pdf" target="_blank"><span>11 percent</span></a> of the nation’s electricity consumption by 2035, placing increasing pressure on generation, transmission, and system reliability.</p><p>At the same time, <a href="https://ieee-pes.org/climate-change/the-future-of-energy-quantified-2026-global-member-survey-results/" target="_blank"><span>insight from the IEEE Power and Energy Society (PES)</span></a> highlights that meeting energy demand from AI and digital infrastructure is one of the most significant challenges facing engineers over the next decade.</p><p>The implications are clear. In addition to computing challenges, AI poses major energy systems challenges.</p><p class="pull-quote">“As artificial intelligence continues to scale globally, the challenge is no longer just computational power, it is the energy systems required to support it” <strong>—Professor Thas (Ampalavanapillai) Nirmalathas, University of Melbourne</strong></p><h2>Why Melbourne is leading on the global stage</h2><p>Victoria has developed one of the most advanced and integrated energy ecosystems in Australia and globally, spanning renewable generation, battery storage, grid modernization, and advanced materials.</p><p>What distinguishes Melbourne globally is how these capabilities are connected and applied at system scale.</p><p>The city brings together world class engineering research, a rapidly evolving clean energy sector, advanced digital infrastructure, and strong alignment between government, industry, and academia. This convergence is critical in the AI era, where energy, networks and computing systems must be designed together.</p><p>Victoria’s coordinated investment across these areas is positioning Melbourne not only as a national leader, but also as a reference point in the global energy system transformation.</p><h2>Engineering the systems behind the AI economy</h2><p>The challenge ahead is that generating more power won’t be enough, as engineers need to design systems that respond dynamically to new patterns of demand.</p><p>Three priorities are emerging globally:</p><ul><li>Aligning data center development with grid capacity and renewable supply</li><li>Embedding flexibility through storage, demand response, and system optimization</li><li>Balancing digital growth with decarbonization and long-term reliability</li></ul><p>Addressing these priorities requires engineering expertise to be embedded earlier in planning ensuring energy systems, digital infrastructure, and policy are designed in parallel.</p><p>Melbourne’s strength lies in its ability to integrate this expertise across research, infrastructure, and real-world application.</p><p class="shortcode-media shortcode-media-rebelmouse-image image-crop-custom"> <img alt="Crowd mingling in a modern glass courtyard during an outdoor social event" class="rm-shortcode" data-rm-shortcode-id="6d59a3228ed2e819398447ea955abc07" data-rm-shortcode-name="rebelmouse-image" id="e734f" loading="lazy" src="https://spectrum.ieee.org/media-library/crowd-mingling-in-a-modern-glass-courtyard-during-an-outdoor-social-event.jpg?id=66945563&width=2000&height=1335&quality=100&coordinates=0%2C606%2C0%2C0"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Melbourne Connect is a University of Melbourne–led innovation precinct, supported by government and industry, designed to bring together research, business and policy to deliver real-world solutions.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Atlantic Group</small></p><h2>Research leadership shaping global solutions</h2><p>At the centre of this capability is the <a href="https://www.unimelb.edu.au/" target="_blank"><span>University of Melbourne</span></a>, where interdisciplinary research is advancing the systems required to support AI driven energy demand.</p><p>Through the Melbourne Energy Institute, for example, researchers are examining how energy technologies interact across entire systems from generation and networks through to end use.</p><p>“As artificial intelligence continues to scale globally, the challenge is no longer just computational power, it is the energy systems required to support it,” says <a href="https://about.unimelb.edu.au/leadership/senior-leadership/dean-feit" target="_blank">Professor Thas (Ampalavanapillai) Nirmalathas</a>, Dean of the Faculty of Engineering and Information Technology at the University of Melbourne.</p><p>“This is driving a new level of convergence between digital infrastructure and power systems engineering, where integrated, system level thinking is essential.”</p><h2>Converging energy, networks and AI</h2><p>Melbourne’s leadership is further strengthened by world-class interdisciplinary facilities such as the <a href="https://electrical.eng.unimelb.edu.au/power-energy/smart-grid-lab" target="_blank"><span>Smart Grid Lab</span></a> in the Department of Electrical and Electronic Engineering, which enables real-time simulation of power systems, allowing engineers to test how solar, batteries, electric vehicles and other distributed resources interact within future grids. This supports the design of more resilient, efficient energy systems before they are deployed at scale.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Control room with server racks, workstations, and a large grid monitoring display." class="rm-shortcode" data-rm-shortcode-id="26c2b42a204f901444b87d17ac31a351" data-rm-shortcode-name="rebelmouse-image" id="b628c" loading="lazy" src="https://spectrum.ieee.org/media-library/control-room-with-server-racks-workstations-and-a-large-grid-monitoring-display.jpg?id=67073323&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Melbourne’s Smart Grid Lab in the Department of Electrical and Electronic Engineering enables real-time simulation of power systems. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">University of Melbourne</small></p><p>These capabilities will become increasingly important as data centers are integrated into the grid.</p><p><span>“AI driven demand is not only increasing computing requirements, but also placing new pressures on underlying energy systems,” says <a href="https://findanexpert.unimelb.edu.au/profile/1024365-glen-farivar" target="_blank">Glen Farivar</a>, Senior Lecturer in Power Electronics at the University of Melbourne. “Designing these systems together is essential to achieving both performance and sustainability outcomes.”</span></p><p>This reflects a critical shift. Future infrastructure must be co designed across energy and digital systems, not developed in isolation.</p><h2>A living ecosystem delivering real-world outcomes</h2><p>Victoria’s broader energy ecosystem is translating these insights into practice.</p><p>Investment in renewable energy, grid infrastructure and storage is enabling higher levels of clean energy while maintaining reliability. Battery deployment is supporting the flexibility needed to manage both renewable variability and growing AI-driven demand.</p><p>At its core, Melbourne offers an integrated environment where research, industry and government collaborate to solve complex system challenges.</p><h2>Why engineering collaboration matters</h2><p>Solving the energy demands of the AI era cannot be achieved in isolation.</p><p>It requires engineers, researchers, utilities, and policymakers to work together earlier and more often. More than ever, engineering collaboration is a critical enabler of future energy systems.</p><p>Environments that bring together global expertise are becoming essential to how solutions are designed and delivered.</p><p class="pull-quote">“Developing future energy systems that are affordable, sustainable, and resilient is a truly grand challenge” <strong>—Professor Pierluigi Mancarella, University of Melbourne</strong></p><p>In this context, the University of Melbourne is co-leading, alongside Johns Hopkins University and Imperial College London, one of only seven <a href="https://www.unimelb.edu.au/newsroom/news/2023/september/new-global-research-centre-to-provide-epic-clean-energy-boost" target="_blank"><span>Global Centres in Climate Change and Clean Energy</span></a>. Through the Electric Power Innovation for a Carbon Free Society (EPICS) Centre, the University is also the Australian technical lead in advancing future energy systems, with EPICS the only Global Centre focused on future energy infrastructure.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Large solar farm in green fields with wind turbines on the horizon under blue sky" class="rm-shortcode" data-rm-shortcode-id="94edf23073999ffbd9272ddc574e4f1c" data-rm-shortcode-name="rebelmouse-image" id="29346" loading="lazy" src="https://spectrum.ieee.org/media-library/large-solar-farm-in-green-fields-with-wind-turbines-on-the-horizon-under-blue-sky.jpg?id=66945577&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">The new Electric Power Innovation for a Carbon-Free Society (EPICS) Centre will address challenges in clean energy production and storage.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">University of Melbourne</small></p><p><span>“Developing future energy systems that are affordable, sustainable, and resilient is a truly grand challenge,” says <a href="https://energy.unimelb.edu.au/about-us/our-team/executive/pierluigi-mancarella" target="_blank">Professor Pierluigi Mancarella</a>, Chair Professor of Electrical Power Systems at the University of Melbourne and Australian director and international co-director of EPICS.</span></p><p>“As electricity grids are increasingly becoming the backbone of future energy systems, optimizing their interactions with other sectors, including AI and digitalization, and fostering interdisciplinary and international collaborations are essential,” he adds.</p><h2>Global conferences as part of the solution</h2><p>International conferences are increasingly recognized as critical platforms for advancing engineering solutions at scale. Melbourne’s ability to convene global expertise is central to its leadership.</p><p>In 2027, the city will host the <a href="https://www.ieeegtd2027.org" target="_blank"><span>IEEE PES Generation Transmission and Distribution (GTD) Asia 2027</span></a> Conference and Exposition, bringing together engineers, utilities, researchers and policymakers from across the world to address the challenges shaping the future of power systems.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Four men pose at a 2025 GTD conference booth with energy-themed backdrop." class="rm-shortcode" data-rm-shortcode-id="9155eae80ac2c5f8e9278b96832fb3ef" data-rm-shortcode-name="rebelmouse-image" id="24eaf" loading="lazy" src="https://spectrum.ieee.org/media-library/four-men-pose-at-a-2025-gtd-conference-booth-with-energy-themed-backdrop.jpg?id=66945590&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">IEEE PES GTD Asia 2027 Melbourne Committee (left to right): Dr. Mehdi Ghazavi Dozein (Monash University), Dr. Glen Farivar & Professor Pierluigi Mancarella (University of Melbourne) , Dr. Mohammad Mohammadi (Australian Energy Market Operator (AEMO)).</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MCB</small></p><p><span>“Melbourne offers a unique environment where world-class research, industry capability and policy leadership come together,” notes the IEEE PES GTD Asia 2027 Local Organising Committee, which includes Professor Pierluigi Mancarella and Dr. Glen Farivar from the University of Melbourne, as well as Dr. <a href="https://www.monash.edu/engineering/mehdighazavidozein" target="_blank">Mehdi Ghazavi Dozein</a> of Monash University and Dr. Mohammad Mohammadi of the Australian Energy Market Operator.</span></p><p>“Hosting this event creates an opportunity to advance global collaboration on the systems and technologies required to deliver the energy transition at scale.”</p><p>These forums enable knowledge exchange, standards development and interdisciplinary collaboration, accelerating progress on complex engineering challenges.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Two people view a circular digital art installation of glowing screens and green light." class="rm-shortcode" data-rm-shortcode-id="733f97dd75ad977c8ffe833833c62e74" data-rm-shortcode-name="rebelmouse-image" id="9b439" loading="lazy" src="https://spectrum.ieee.org/media-library/two-people-view-a-circular-digital-art-installation-of-glowing-screens-and-green-light.jpg?id=66986093&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Attendees view a digital installation at AIME 2025 at Melbourne Connect.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">MCB</small></p><h2>Why Melbourne, and why now</h2><p>As AI, electrification and digital infrastructure converge, the future of global energy systems will depend on the ability of engineers to collaborate and innovate at scale.</p><p>Melbourne provides a proven platform for that collaboration, combining world-class research, a rapidly evolving energy ecosystem, and the infrastructure to connect global expertise.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Group standing with award outside historic brick building and garden walkway" class="rm-shortcode" data-rm-shortcode-id="7f75d2c90839db5861612d3ed8fef1f3" data-rm-shortcode-name="rebelmouse-image" id="6eed5" loading="lazy" src="https://spectrum.ieee.org/media-library/group-standing-with-award-outside-historic-brick-building-and-garden-walkway.jpg?id=66945594&width=980"/> <small class="image-media media-caption" placeholder="Add Photo Caption...">Melbourne Convention Bureau, IEEE Communications Society, and University of Melbourne Representatives.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">University of Melbourne</small></p><p><span>For IEEE members, hosting a conference in Melbourne is more than an event decision.</span></p><p>It is an opportunity to engage with a globally connected engineering community and contribute directly to solving one of the most significant challenges facing the profession today.</p><p>Through the support of the <a href="https://www.melbournecb.com.au/contact-us?utm_source=ieee&utm_medium=editorial&utm_campaign=discover-melbourne-2026&utm_term=power-and-energy&utm_content=contact-us" target="_blank"><span>Melbourne Convention Bureau</span></a>, professionals can access tailored, free support to bid for and deliver international conferences, bringing global expertise together in a city actively shaping the future of energy systems.</p><p><strong>To explore hosting your next conference in Melbourne, contact the Melbourne Convention Bureau at info@melbournecb.com.</strong></p>]]></description><pubDate>Wed, 01 Jul 2026 16:01:27 +0000</pubDate><guid>https://spectrum.ieee.org/ai-energy-systems-melbourne</guid><category>Artificial-intelligence</category><category>Australia</category><category>Energy-systems</category><category>University-of-melbourne</category><category>Ai-data-centers</category><category>Power-grid</category><dc:creator>Melbourne Convention Bureau</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/glowing-digital-network-map-of-australia-and-surrounding-asia-pacific-region.png?id=66945530&amp;width=980"></media:content></item><item><title>The Space-based Data Center Hype Machine Is Already in Orbit</title><link>https://spectrum.ieee.org/orbital-data-center-hype</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/globe-wrapped-in-multicolored-pins-and-connecting-lines-symbolizing-global-networks.png?id=67007490&width=1245&height=700&coordinates=0%2C209%2C0%2C209"/><br/><br/><p><span>“</span><span>The lowest-cost place </span>to put AI will be in space, and that will be true within two years, maybe three at the latest,” SpaceX founder Elon Musk told the World Economic Forum in Davos this past January, as his company was <a href="https://www.sec.gov/Archives/edgar/data/1181412/000162828026036936/spaceexplorationtechnologi.htm" target="_blank">preparing to go public</a>.</p><p>Later that month, SpaceX filed an application with the Federal Communications Commission for an orbital data center constellation of up to 1 million satellites in low Earth orbit, 500 to 2,000 kilometers above Earth. And just three days before the IPO, he discussed some initial design specifications for a new <a href="https://x.com/SpaceX/status/2064099405758906727" target="_blank">AI-1 satellite data center</a> in a video interview.</p><p>Musk is prone to hyperbole when it comes to timelines. Full <a href="https://techcrunch.com/2025/01/30/elon-musk-reveals-elon-musk-was-wrong-about-full-self-driving/" target="_blank">self-driving cars by 2017</a>. <a href="https://washingtonian.com/2026/02/12/how-elon-musks-sci-fi-hyperloop-failed/" target="_blank">First human mission to Mars in 2024</a>. <a href="https://washingtonian.com/2026/02/12/how-elon-musks-sci-fi-hyperloop-failed/" target="_blank">Ten thousand Optimus humanoid robots by the end of 2025</a>. Et cetera. For orbital data centers, which he says will be a cost-effective alternative to terrestrial data centers within three years, the math won’t make sense for several years, if ever.</p><p>Consider this: There are roughly <a href="https://satfleetlive.com/blogs/how-many-satellites-in-orbit/" target="_blank">14,500 active satellites in orbit</a>. Musk’s Starlink constellation accounts for about <a href="https://spacenexus.us/blog/how-many-satellites-in-space-2026" target="_blank">two thirds of those</a>. Both the launch cadences and satellite-manufacturing capacity would have to scale up astronomically to deploy a million orbital data center satellites.</p><p>For context, there have been <a href="https://planet4589.org/space/gcat/data/derived/launchlog.html" target="_blank">roughly 7,000 orbital launches in all of human history</a>. To loft 1 million satellites into low Earth orbit on SpaceX’s Starship, which is designed to carry up to 60 satellites per vehicle, would require 16,666 launches exclusively devoted to satellite deployments. Considering that SpaceX launched a record 165 orbital missions in 2025, even at 10 times that cadence, it would take a decade. And how long would it take to build 1 million satellites, given Starlink’s <a href="https://www.advanced-television.com/2026/04/13/analyst-spacex-making-340-satellites-per-month/" target="_blank">current pace of around 4,000 per year</a> and a generous tenfold increase in capacity? Short of a manufacturing revolution, try 25 years.</p><p class="pull-quote">The reality is that the vision of massive constellations of orbital data centers is nowhere close to being realized.</p><p><strong></strong>As this month’s cover story, “<a href="https://spectrum.ieee.org/orbital-data-centers-heat" target="_blank">Why Orbital Data Centers Are So Hard</a>” by <a href="https://www.abiresearch.com/staff/analysts/andrew-cavalier" target="_blank">Andrew Cavalier of ABI Research</a>, makes clear, the reality is that the vision of massive constellations of orbital data centers is nowhere close to being realized.</p><p>Dina Genkina, <em>IEEE Spectrum</em>’s computing and hardware editor, put the idea into perspective: “Starcloud (a startup that has applied to the FCC for an 88,000 orbital data center satellite constellation) <a href="https://spectrum.ieee.org/nvidia-h100-space" target="_blank">sent one Nvidia H100 GPU in space so far</a>. Their radiator was too weak to let the chip run at full power.”</p><p>As Cavalier shows, cooling even a single Nvidia H100 GPU in space is difficult: It draws 700 watts, which will require 1.4 square meters of radiator at 60 °C. A 40-kilowatt rack of servers will need an 80-m² radiator; a 100-megawatt data center will require 2,500 of those radiators. Some astronomers are understandably concerned that a million satellites with giant radiative wings would blot out the stars.</p><p>So if the economics doesn’t make sense, if the chips are at the mercy of the radiative ravages of space, and if humanity will lose its view of the stars, not to mention increasing the risk of triggering the Kessler syndrome, why are the hyperscalers hyping orbital data centers?</p><p>Genkina offered the obvious answer: sweet, sweet moolah. “The Elon Musk part of it is honestly genius because he’s got xAI building the data centers, SpaceX sending them to space, and Tesla building solar panels,” Genkina says. “It’s almost like he’s paying himself.”</p><h3>Two Analyst’s Views of SpaceX’s Proposed AI1 Data Center Satellite</h3><br/><h3></h3><br/><p><strong><a href="https://www.linkedin.com/in/piercemichaelj/" rel="noopener noreferrer" target="_blank">Michael Pierce</a>, Principal at Technology Strategy Partners</strong></p><p>Musk’s timelines are notoriously overly ambitious, but I think SpaceX’s orbital data centers might reach cost parity with terrestrial data centers in 5 to 10 years. The Starlink laser-link network already exists as the communication backbone for any SpaceX compute constellation, and that infrastructure is what no new entrant can replicate quickly. The chip-agnostic payload design probably reflects their disclosed difficulty securing AI silicon as much as any modularity philosophy. My view is that the only realistic near-term application is a SpaceX mega-constellation for inference. Training workloads likely cannot tolerate the synchronization and latency constraints of a distributed orbital system.</p><p>Our <a href="https://t-s-partners.com/whitepapers/" target="_blank">report</a> analyzed the market from the integrator’s vantage point, but AI1 is what it looks like when one player has assembled all the necessary advantages simultaneously. The question is whether the terrestrial data center industrial base will degrade or improve on economics. I don’t have insight into SpaceX’s internal costs, as opposed to public pricing, on all their components, so it’s hard to say if they’ll completely dominate or not. Even if they are not cost competitive with terrestrial data centers for another 5 to 10 years, it may simply be faster to get new compute that just happens to be in space.</p><h3></h3><br/><p><strong><a href="https://matthasan.com/" rel="noopener noreferrer" target="_blank">Matt Hasan</a>, AI strategist and independent consultant</strong></p><p>My initial view is that AI1 does not fundamentally change the rationale for space-based data centers as much as it changes the timeline and scale. The underlying drivers remain the same: escalating AI compute demand, growing power constraints on terrestrial grids, and the desire to colocate energy generation with computation.</p><p>What AI1 does signal is that the concept is beginning to move from theoretical discussion toward engineering and capital allocation decisions. The announcement adds credibility to the idea that hyperscale computing infrastructure may eventually expand beyond terrestrial constraints rather than simply competing for increasingly scarce grid capacity on Earth.</p><p>That said, significant economic and technical questions remain. Launch costs, maintenance, hardware replacement cycles, thermal management, latency-sensitive workloads, and overall system economics will ultimately determine whether space-based data centers become a mainstream extension of AI infrastructure or remain a niche capability for specialized applications. The key development is not that these questions have been resolved, but that major industry players now appear willing to invest resources toward answering them.</p>]]></description><pubDate>Wed, 01 Jul 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/orbital-data-center-hype</guid><category>Orbital-data-centers</category><category>Satellites</category><category>Spacex</category><category>Elon-musk</category><category>Starcloud</category><category>Ai</category><category>Gpus</category><dc:creator>Harry Goldstein</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/globe-wrapped-in-multicolored-pins-and-connecting-lines-symbolizing-global-networks.png?id=67007490&amp;width=980"></media:content></item><item><title>Emily Bender Sets the Record Straight on “Stochastic Parrots”</title><link>https://spectrum.ieee.org/stochastic-parrot</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/collage-of-a-white-brunette-woman-against-a-background-with-two-abstract-parrots-composed-of-coding-slashes.jpg?id=67050137&width=1245&height=700&coordinates=0%2C62%2C0%2C63"/><br/><br/><p>In March 2021, a group of four researchers—a collaboration of linguists and computer scientists—published their now legendary paper “<a href="https://dl.acm.org/doi/10.1145/3442188.3445922" rel="noopener noreferrer" target="_blank">On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? </a>🦜” </p><p>The paper received significant attention at the time (in part because Google fired two of the authors, <a href="https://spectrum.ieee.org/timnit-gebru-dair-ai-ethics" target="_self">Timnit Gebru</a> and Margaret Mitchell, shortly before its publication). It argued that large language models (LLMs) generate text by statistically predicting likely sequences of words rather than understanding what they are saying—a process the authors captured with the metaphor of a “stochastic parrot,” a system that repeats patterns without comprehension. And over the past five years, the analogy has spread well beyond the academic field where it originated, spawning debates and inspiring projects such as a <a href="https://www.media.mit.edu/projects/the-stochastic-parrot/overview/" rel="noopener noreferrer" target="_blank">shoulder-mounted robot</a> named the Stochastic Parrot. </p><p>But that wider usage has also led to misconceptions about what the phrase originally meant. Lead author <a href="https://faculty.washington.edu/ebender/" rel="noopener noreferrer" target="_blank">Emily M. Bender</a>, a professor of computational linguistics at the University of Washington, recently <a href="https://medium.com/@emilymenonbender/stochastic-parrots-frequently-unasked-questions-49c2e7d22d11" rel="noopener noreferrer" target="_blank">wrote a blog post</a> to debunk common misconceptions about the paper on its five-year anniversary. </p><p>Bender spoke with <em><em>IEEE Spectrum</em></em> about these misconceptions, the field of computational linguistics, and the current discourse around artificial intelligence. </p><h2>What’s Wrong With the Term “Artificial Intelligence” </h2><p><strong>How would you describe your work as a computational linguist?</strong></p><p><strong>Emily M. Bender: </strong>Linguistics, very generally, is the study of how language works and how we work with language. I contribute to that, and I also work in computational linguistics, training students who are going to go on to build language technology. </p><p>Language technology actually stands alone as valuable and interesting, independent of whether or not someone wants to use it for their project of artificial intelligence. Language technology includes things like automatic transcription, machine translation, spell check. And a lot of the work that I do personally, when I am building things, has to do with building machine-readable, but also human-readable grammars that model linguistic phenomena in different languages. That’s about using computers in the service of linguistic hypothesis testing.</p><p><strong>You’ve argued that the term “artificial intelligence” obscures more than it clarifies. Why?</strong></p><p><strong>Bender: </strong>Many reasons. I think that it makes it difficult to actually have good discussions about technology and make wise decisions about it, if the way we’re talking about it doesn’t make clear what the technology is. The phrase “artificial intelligence” both groups together disparate technologies and oversells what each one of them can do. So if we are trying to decide whether or not to use something, how to regulate something, we are much better off with clearer descriptions.</p><p><strong>In general conversation, AI has become almost synonymous with “chatbots” or “LLMs.” Is that a problem? </strong></p><p><strong>Bender: </strong>For many people, they’ll say, “I use it to do blah blah blah.” So what do you mean by “it”? And then they’ll say, “Oh, I mean Claude” or ChatGPT or Gemini, so they are talking about these chatbots. But then other people will say, “You can’t say AI is all bad, because what about <a href="https://spectrum.ieee.org/alphafold-proves-that-ai-can-crack-fundamental-scientific-problems" target="_self">AlphaFold</a>?” </p><p>So, yes, for many people, they are talking about chatbots built on top of large language models, but [they’re] also not really clear that those things are separate from something like AlphaFold. And when we have news reporting that says “scientists use AI to <a href="https://spectrum.ieee.org/isomorphic-labs-ai-drug-discovery" target="_self">discover a new drug</a>,” well, what did they use? If what they’re talking about is something much more narrow, maybe it’s protein folding, maybe it’s some other kind of statistical modeling [like in] <a href="https://spectrum.ieee.org/ai-weather-forecasting" target="_self">weather modeling</a>. That’s a very different kind of technology than ChatGPT.</p><p><strong>Do you think there’s a value to an umbrella term like “artificial intelligence”?</strong></p><p><strong>Bender: </strong>Well, there’s a value to people who are trying to sell this—so too the tech companies trying to raise their valuations. Also, the way research funding is set up right now, it is very hard to get funded if you don’t call what you’re doing artificial intelligence. That I think is a net negative, but for any individual trapped in that system, that can have value in the moment.</p><h2>How Stochastic Parrots Have Been Misunderstood</h2><p><strong>What are the most common misconceptions about the “stochastic parrots” metaphor?</strong></p><p><strong>Bender: </strong>I think one of the biggest ones is, “Bender says AI is a stochastic parrot.” That paper was written in late 2020. We were talking about large language models. I’m pretty sure the word <em>AI</em> comes up only once at the very end, and that’s talking about how, if you’re going to develop systems that are meant to do things like what people do, you have to be very careful that you are not creating something that can be mistaken for a person. The fact that these systems are designed to mimic the way we use language makes it very easy for people to mistake them for other people. </p><p>So in the paper, toward the very end, we sort of generalized to AI. But the phrase “stochastic parrots” specifically refers to large language models, and the phrase “artificial intelligence” refers to many different things. So we were never claiming that a chess engine or AlphaFold or an image labeling system or a machine translation system, any of those things that are sometimes called artificial intelligence, are stochastic parrots. We were specifically talking about using large language models to produce synthetic text.</p><p>Another one is that “stochastic parrot” got picked up and interpreted by other people as a minimization or an insult. It was not meant that way. Other people might be using it that way, but that’s not how I intended it, because it’s just a description of what these systems actually are. To see it as an insult requires either the belief that the large language model is the kind of thing that can take offense, which it isn’t, or that these large language models should be understood as steps toward this grand ideal that I don’t hold of artificial intelligence. </p><p>What I have been doing in many places—<a href="https://aclanthology.org/2020.acl-main.463.pdf" rel="noopener noreferrer" target="_blank">the octopus thought experiment</a>, stochastic parrots, the phrase “synthetic text-extruding machines”—it’s all about trying to make vivid to people who aren’t in the business of building language technology what these systems actually do, which is not the same thing as insulting the systems or insulting the people who like the systems.</p><p class="ieee-inbody-related">RELATED: <a href="https://spectrum.ieee.org/ai-chatbot" target="_self">The Great Chatbot Debate: Do They Really Understand?</a></p><p><strong>For readers who don’t know, the “octopus test” comes from a 2020 paper that imagined an octopus recognizing the statistical patterns within messages passed through an undersea cable. With the octopus test and stochastic parrots, you’ve used animal metaphors a couple of times now. Is that intentional?</strong></p><p><strong>Bender: </strong>No, it’s not intentional. With the octopus thought experiment, I initially had told the story in terms of a dolphin, because dolphins clearly are intelligent animals. My co-author on that paper, <a href="https://www.coli.uni-saarland.de/koller/" rel="noopener noreferrer" target="_blank">Alexander Koller</a>, said it should be an octopus, because first of all, the environment that octopuses live in is much more distinct from where people live. It makes the metaphor more vivid, that the octopus is just feeling these pulses in the cable and has no way to look at what the people are looking at. But also, octopuses are just inherently funnier.</p><p><strong>I was looking back at that paper and was surprised that the term “stochastic parrots” actually only appears twice in the text itself. Why did you include it in your title?</strong></p><p><strong>Bender: </strong>Because we liked it! And a catchy title is good self-marketing of an academic paper. The reason that there’s not so much of it in the paper is that we were really looking at the full range of risks of making language models ever bigger. The phrase “large language model” also doesn’t show up in the paper, because people weren’t talking about them that way. </p><p>So the section on synthetic text, in some ways it felt like we were on thin ice, because at that point in time it was hard to imagine that anybody would want synthetic text. That part of the paper became much more relevant when OpenAI imposed ChatGPT on the world. Then that particular part of the paper comes out as important. But we also talk about environmental impact. We talk about the ways in which these systems will absorb the biases of their training data. We talk about how the training data is never collected well. There’s a lot of various points in there, and the issues about synthetic text were just one. </p><p><strong>Researchers at MIT Media Lab created a Stochastic Parrot robot as a response to the observation that many chatbots tend to be sycophantic, or overly agreeable. Does that trend relate to the dangers you laid out in your paper?</strong></p><p><strong>Bender: </strong>When we wrote that paper in late 2020, at the time, people were not super excited about synthetic text, nor about chatbots. Chatbots had been around. We had Weizenbaum’s <a href="https://spectrum.ieee.org/why-people-demanded-privacy-to-confide-in-the-worlds-first-chatbot" target="_self">Eliza in the 1960s</a>, and then the very annoying automatic customer service systems that have gotten much more fluent with the large language models, and no less annoying. </p><p>So, that was the state of things. OpenAI had put out GPT-2 and GPT-3 for people to play with, and you could get them to extrude synthetic text, but the chat interface hadn’t been wrapped around those yet. We also hadn’t seen the layers of additional training that lead to the behavior that’s interpreted as sycophantic. The reason that you get the chatbot saying, “Oh, that’s a good idea,” or if you say you’re wrong, it says, “Oh, I’m so sorry, you’re right,” that kind of response has to do with <a href="https://www.ibm.com/think/topics/rlhf" rel="noopener noreferrer" target="_blank">additional layers of training</a> past the original pre-training. </p><p><strong>What do you wish more people understood about language models?</strong></p><p><strong>Bender: </strong>The message that I always bring when I have a chance is that, when the text that comes out of one of these systems makes sense, it’s because we are making sense of it. This is also in the stochastic parrots paper. Anytime we are evaluating this kind of technology, we have to account for our ability to make sense of language and keep that in view as we are deciding what’s going on with the technology. That is frequently lost in these discussions.</p><p><strong>If you were to redo or update the stochastic parrot paper now, is there anything that you would change about it?</strong></p><p><strong>Bender: </strong>There was one really big form of harm that we did not cover in the paper, and that has to do with exploitative labor practices. Under that, I include both the horrible conditions that many data workers face, and also the massive theft of people’s <a href="https://spectrum.ieee.org/generative-ai-ip-problem" target="_self">creative and intellectual output</a> that underlies these systems. Those issues should have been included in the paper. It’s not that they were unknown in the world then, but they didn’t make it into what we surveyed, and should be there.</p><p><em>This story was updated on 1 July 2026 to clarify the research areas of the stochastic parrots paper authors. </em><br/></p>]]></description><pubDate>Tue, 30 Jun 2026 14:00:02 +0000</pubDate><guid>https://spectrum.ieee.org/stochastic-parrot</guid><category>Emily-bender</category><category>Large-language-models</category><category>Llms</category><category>Ai-ethics</category><dc:creator>Gwendolyn Rak</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/collage-of-a-white-brunette-woman-against-a-background-with-two-abstract-parrots-composed-of-coding-slashes.jpg?id=67050137&amp;width=980"></media:content></item><item><title>Poetry for Engineers: Nine Lives of Nikola Tesla</title><link>https://spectrum.ieee.org/poetry-for-engineers-nine-lives-of-nikola-tesla</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/blurred-person-with-spinning-wheel-and-bright-light-trail-in-dark-workshop.jpg?id=67005822&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><br/></p><h3></h3><br/><p>He was born into a storm, lightning split the summer sky, in a<br/>village the world had not yet heard of.<br/>The midwife called it a bad omen, his mother called it a sign. Your first<br/>life began in a storm, under open sky.</p><p>One winter night you ran your hand along a cat’s back, and the<br/>darkness cracked open with sparks.<br/>Your mother warned the house could burn.<br/>You were already chasing what you learned: Light would return.</p><p>Your second life came underwater, in the current deep. No light,<br/>no air, the river pulling you under,<br/>the surface closing above you without a sound, and<br/>something in you refused to sink or sleep.</p><p>Your third life came at the dam.<br/>The water rose. The wall held you in place.<br/>One flash, you turned your body and rose back into air, and left<br/>the weight of water without a trace.</p><p>Your fourth life came in stone and dark. Entombed for a<br/>night in a mountain chapel,<br/>visited by no one. Only silence and the memory of a spark. You called<br/>it an awful experience and left it there, untold.</p><p>Your fifth life came in fever,<br/>nine months cholera held you down,<br/>until your father said: Survive, and choose your own ground. You rose.<br/>Not from the prayer, but from the promise he made.</p><p>Your sixth life came in silence, and it stayed.<br/>Every sound cut through you, a clock three rooms away,<br/>a ringing that would not leave, a noise you learned to bear, until you<br/>lived inside that noise and made a home in there.</p><p>Your seventh life burned on Fifth Avenue, not your body, but your work. Not a thief<br/>of fire, but one who stayed with the blaze.<br/>A modern Prometheus, your life’s work turned to ash,<br/>“I must begin again,” you said, and turned to new ways.</p><p>Your eighth life came in the street.<br/>No storm. No warning. A taxi struck without a sign. A<br/>sudden impact: ribs breaking, breath gone.<br/>No diagram this time. Only the body, slow to keep up.</p><p>The ninth life came on quiet wings.<br/>That dove found you in the dark, and your spirit rose. She did<br/>not move. A beam of light fell from above.<br/>The life you would not return from, the one you loved.</p><p>Your mother thought you had nine lives, nine close<br/>brushes with death.<br/>Each close call, a lesson. A hand that would lead you out of the<br/>darkness and into the dynamo of eternal light. The world profits<br/>from the mystery of your mind,<br/>Upon your imagination we stand.</p>]]></description><pubDate>Tue, 30 Jun 2026 12:24:33 +0000</pubDate><guid>https://spectrum.ieee.org/poetry-for-engineers-nine-lives-of-nikola-tesla</guid><category>Verse-becomes-electric</category><category>Poetry</category><category>Nikola-tesla</category><category>Artificial-intelligence</category><category>Type-departments</category><dc:creator>Danica Radovanović</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/blurred-person-with-spinning-wheel-and-bright-light-trail-in-dark-workshop.jpg?id=67005822&amp;width=980"></media:content></item><item><title>The Lab Mistake That Might Revolutionize Computing</title><link>https://spectrum.ieee.org/artificial-neurons-on-silicon-chips</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/illustration-of-a-microchip-under-a-microscope-with-probes-and-orange-wires-attached.jpg?id=66967576&width=1245&height=700&coordinates=0%2C760%2C0%2C761"/><br/><br/><p><strong>Today, you </strong><strong>probably asked</strong> a question of a large language model, or accepted a connection suggestion on LinkedIn, or watched a recommended video on YouTube, or took a different route to work based on a traffic prediction from Google Maps. In other words, you probably used artificial intelligence. But what you might not know is how much energy that interaction consumed or why.</p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/artificial-neurons-on-silicon-chips?draft=1&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p class="shortcode-media shortcode-media-rebelmouse-image" style="display:none"> <img alt="" class="rm-shortcode" data-rm-shortcode-id="2b00fa4a3e2d69e8112f0268f9b668e5" data-rm-shortcode-name="rebelmouse-image" id="29a98" loading="lazy" src="https://spectrum.ieee.org/media-library/image.png?id=67033315&width=980"/></p><h3></h3><br/><p>AI requires processing massive amounts of data, which is usually done in large data centers populated by thousands of GPUs capable of executing up to trillions of operations per second. But each of those GPUs achieves that by consuming as much as 1,000 watts apiece. For comparison, if you’ve got a newer smartphone, it probably uses less than 1 W. That kilowatt figure puts GPUs on the same level as vacuum cleaners, dishwashers, and stoves, but with the big difference that data-center processors are operating uninterrupted around the clock.</p><p>Fundamentally, a lot of this inefficiency is because GPUs are trying to simulate the workings of artificial neural networks using software and billions of transistors, which requires using energy to move massive amounts of data. What’s more, the simulated artificial neurons that make up these networks lack even a fraction of the complex computing behavior of the biological neurons that comprise the most energy-efficient computing system that we know, the human brain.</p><h3></h3><br/><img alt="Gloved hand with tweezers holding a tiny swab over colorful striped background" class="rm-shortcode" data-rm-shortcode-id="94a328470814a03a385c1d9a2448f58b" data-rm-shortcode-name="rebelmouse-image" id="c570c" loading="lazy" src="https://spectrum.ieee.org/media-library/gloved-hand-with-tweezers-holding-a-tiny-swab-over-colorful-striped-background.jpg?id=66990755&width=980"/><h3></h3><br/><p>The brain is roughly<a href="https://www.nist.gov/blogs/taking-measure/brain-inspired-computing-can-help-us-create-faster-more-energy-efficient" target="_blank"> one million times as energy efficient</a> at many of the comparable tasks we set for AI. <a href="https://ieeexplore.ieee.org/document/8094868" target="_blank">To try to approach these efficiencies</a>, a radically different way of computing called <a href="https://spectrum.ieee.org/tag/neuromorphic-computing" target="_self">neuromorphic engineering</a> is seeking to build electronic components and circuits that act more like the brain’s neurons and the synapses that connect them.</p><p>Huge amounts of work have gone into making electronics operate more like <a href="https://spectrum.ieee.org/artificial-neuron" target="_self">biological neurons and synapses</a>. Some research has focused on developing <a href="https://spectrum.ieee.org/memristor-first-single-device-to-act-like-a-neuron" target="_self">new</a>, <a href="https://spectrum.ieee.org/artificial-synapses" target="_self">experimental devices</a>, but they aren’t yet reliable enough to be used in large systems. Other efforts aim to implement neurons and synapses by interconnecting many complementary metal-oxide-semiconductor (CMOS) transistors—the workhorses of digital logic—to simulate a single neuron and synapse. But this approach requires so many transistors (and a few bulky capacitors) that it greatly limits the size of the system that can be constructed, making it unclear how such brain-inspired hardware could ever scale up and compete with state-of-the-art GPUs.</p><p>But all along there was an artificial neuron and a synapse—each a single device—hiding in plain sight. We found them last year. They were each made possible by an ordinary CMOS transistor—and not even a very good one at that. This is the story of their accidental discovery and their great promise for lowering the environmental footprint of AI.</p><h2>Biological and artificial neurons</h2><p>Modern digital electronics is based on producing and manipulating the ones and zeros of the binary code through the operation of metal-oxide-semiconductor field-effect transistors. MOSFETs have evolved in recent years, but their classic form consists of a piece of silicon that has been doped to contain an excess of either positive (<em>p</em>-type) or negative (<em>n</em>-type) charge carriers. (CMOS logic contains transistors of both types.) The device has two terminals connected to the silicon through regions highly doped with the opposite polarity of the rest of the silicon—the source and the drain. Another terminal, the gate, sits atop the silicon that separates the source from the drain. The gate itself doesn’t connect directly to this silicon, instead resting above a thin layer of insulating dielectric.</p><p>Notably, there is a fourth terminal that attaches to the bulk of the silicon; think of this bulk terminal as connecting to the underside of the chip. It doesn’t typically get much attention, but it’s very important to our story.</p><p>When voltage is applied at the gate and the bulk terminal is grounded, charge carriers of the same polarity as the source and drain are attracted to the channel region. In the case of an <em>n</em>-type source and drain, that will be electrons; for <em>p</em>-type it will be holes. The presence of these charges forms a conductive channel that reduces the resistance between the source and the drain by several orders of magnitude, and the device switches on. As the voltage at the gate increases, this physical phenomenon produces a current signal that, when plotted against the gate voltage, rises steadily. This response is ideal for logic gates, converters, multiplexers, memories, and other digital circuits. But it is not a good fit for mimicking the behavior of a neuron.</p><p>In real neural tissue, brain cells, called neurons, consist of a cell body, a long projection called an axon, and short branching projections called dendrites. The suite of behaviors and computing this collection of components is capable of is rich and broad, but the portion that artificial neural networks hope to copy is this: When the cell body’s voltage is perturbed enough to reach a particular threshold, a self-propagating pulse of voltage, called an action potential, shoots down the axon. The axon terminates in a synapse, an electrochemical connection between the axon and another neuron’s dendrites. The action potential will then temporarily boost the voltage of this next neuron, by an amount that depends on the strength of the synaptic connection. If enough action potentials reach these dendrites in a given time—from this neuron or from others that might also form synapses there—the cell body’s voltage will surpass the threshold and trigger its own action potential.</p><h3>The MOSFET Neuron</h3><br/><p>The unusual action the authors discovered is understandable if you consider that a MOSFET contains a hidden bipolar-junction transistor.</p><h3></h3><br/><img alt="MOSFET diagrams with carrier flow and plot of drain current versus drain voltage" class="rm-shortcode" data-rm-shortcode-id="81d6eb5c903261127a7f91d8dc530150" data-rm-shortcode-name="rebelmouse-image" id="751ae" loading="lazy" src="https://spectrum.ieee.org/media-library/mosfet-diagrams-with-carrier-flow-and-plot-of-drain-current-versus-drain-voltage.png?id=67006439&width=980"/><h4><span style="background-color: black; color: white; padding: 2px 6px; font-family: sans-serif; display: inline-block; font-size: 50%"><strong>TRANSISTOR BEHAVIOR</strong></span></h4><p class="caption">Under normal operation, with the bulk terminal grounded, increasing voltage at the drain leads to current that increases steadily. When the voltage decreases, current follows the same sloped path. Although some pairs of electrons and holes are created by current crashing into silicon atoms, these are swept away before they can accumulate.</p><h3></h3><br/><img alt="NSRAM transistor diagrams with bias circuits and I\u2013V curve highlighting C and D states" class="rm-shortcode" data-rm-shortcode-id="2f6b71fe8c9fcea8172ad966fa0912ac" data-rm-shortcode-name="rebelmouse-image" id="2a3dc" loading="lazy" src="https://spectrum.ieee.org/media-library/nsram-transistor-diagrams-with-bias-circuits-and-i-u2013v-curve-highlighting-c-and-d-states.png?id=67005464&width=980"/><h4><span style="background-color: black; color: white; padding: 2px 6px; font-family: sans-serif; display: inline-block; font-size: 50%"><strong>NSRAM BEHAVIOR</strong></span></h4><p class="">Adding resistance to the bulk terminal means these extra holes pile up, increasing the bulk voltage relative to the source. Once that voltage reaches a certain value, the hidden transistor activates, causing current to spike. Current remains high until the drain voltage drops past a certain point. <style class="image-media media-photo-credit">MARIO LANZA & SEBASTIAN PAZOS</style></p><h3></h3><br/><p>To get closer to the behavior of real neurons, artificial neurons should produce a current spike when a critical voltage threshold is crossed and then quickly relax back to a resting state on their own. This spike needs to be sudden—nonlinear. It should also exhibit some hysteresis; that is, the activation and relaxation voltages should be different from each other to ensure that current flows only for a certain amount of time.</p><p>What’s wanted from an artificial synapse, the thing that connects two artificial neurons, is less complicated, but equally important. The main thing is that its conductance can be electronically adjustable. The device’s conductive states should increase and decrease in a linear pattern and remain stable over time.</p><p>No single MOSFET working under the standard operation mechanism can reproduce either of these neural properties. Instead, it’s been done by combining them into complex circuits. Until now, each neuron and each synapse has been implemented by interconnecting dozens and sometimes even hundreds of MOSFETs, which is highly inefficient in terms of area, performance, and cost. To limit the amount of space needed, chips can multiplex their signals, sending them to neurons and synapses serially, but such sequential processing introduces additional delays.</p><p>Despite these area-and-time penalties on tasks such as audio processing, computer vision, or health monitoring, state-of-the-art brain-inspired microchips have achieved power reductions up to a thousandfold compared with those of GPUs or CPUs on the same task. If we could create neurons and synapses from individual devices that are readily manufacturable instead, we might target more massive implementations while maintaining energy efficiency.</p><h2>Reinventing the MOSFET for AI</h2><p>Working in our laboratory in 2024, one of my students was measuring a memory circuit that consisted of one transistor and one memristor—a type of nonvolatile memory device first fabricated in 2008. The student’s memristor circuit was built from two-dimensional material atop a silicon microchip containing MOSFETs. The MOSFETs were created in a commercial foundry using fabrication technology called the 180-nanometer node, which was cutting-edge in the year 2000.</p><p>One day the student forgot to connect the bulk terminal of the transistor. What he observed was a sudden increase in current with high nonlinearity that self-relaxed when the voltage was ramped down (a phenomenon called a hysteresis loop). This was a very promising neuronlike behavior!</p><p>After a fruitless week of trying to think of an explanation for this behavior, I (Lanza) asked Pazos, then my postdoctoral fellow, to try to observe and control this phenomenon in chips without memristors. This time, we applied pulses of voltage—like the spikes a neuron would produce—instead of the ramped voltage that my student used when he first saw the peculiar behavior.</p><p>Pazos’s new data helped us understand what was going on. The key was that oft-ignored fourth, or bulk, terminal of a MOSFET. Under ordinary operation, many mobile charge carriers flitting through the channel collide with the silicon atoms, producing free pairs of electrons and holes—a process known as impact ionization. The electric field created by the potential difference between the source and the drain causes these new free electrons to drift toward the positively biased drain and the holes to move toward the bulk terminal, which is usually grounded, removing the charge without any drama.</p><p>However, when the bulk terminal of the transistor is floating—unconnected as it was in my student’s experiment—the holes produced by impact ionization cannot be driven to the ground. Instead, they accumulate in the bulk of the silicon, increasing its voltage. Then things start to get interesting.</p><p>It helps here to imagine a MOSFET as two different kinds of transistors occupying the same physical space—the intentionally constructed MOSFET and a hidden, bipolar junction transistor. A bipolar device transmits a current signal across two <em>p</em>-<em>n</em> junctions, in this case the interfaces between the source and the channel region and the channel and the drain. This signal is in proportion to a smaller current at a third terminal in between, called the base. In our experiment, that third terminal is the bulk.</p><h3></h3><br/><img alt="Diagram of a leaky integrate-and-fire neuron converting input spikes to output spikes" class="rm-shortcode" data-rm-shortcode-id="6bcb3e9fed5fe165dccd6f5c7a30110b" data-rm-shortcode-name="rebelmouse-image" id="e72de" loading="lazy" src="https://spectrum.ieee.org/media-library/diagram-of-a-leaky-integrate-and-fire-neuron-converting-input-spikes-to-output-spikes.jpg?id=67005640&width=980"/><h3></h3><br/><p>To get current flowing through a bipolar transistor, you need a big enough potential difference between the base and one of the other terminals, so that current can get across the <em>p</em>-<em>n</em> junction. Let’s say this “threshold voltage” is 0.7 volts, although the real number depends on device geometry and silicon doping. In our device, that potential difference comes from those holes that were accumulating in the bulk, because it was not connected to ground. Once it reaches the threshold voltage, the device becomes sharply conductive, producing an abrupt increase of current. This sharp current increase eventually falls off once the drain voltage is lowered, because that lowering reduces the rate at which holes are generated in the bulk. The remaining excess holes recombine with stray electrons or leak away, and finally the bulk voltage falls. This cycle of hole accumulation, current spike, and hole removal gives rise to a hysteresis loop, very much like the electrical behavior of a biological neuron as it integrates ionic currents, fires a spike, and relaxes back to its resting voltage.</p><p>Initially, we observed this behavior only in a few transistors, and the relaxation time was very different for each of them. So, to try to control it better, we adjusted the resistance of the bulk terminal using a second MOSFET. Simply setting that resistance suddenly caused all the transistors to fire at the same voltage with hardly any variability. In other words, we found we could create perfect electronic neuron behavior in a single silicon transistor by controlling the bulk contact resistance. Setting the resistance can be done by doping the silicon during fabrication, but we think the two-transistor cell—where one acts as the bulk resistance—offers much greater versatility because it allows for electronic control.</p><p>We had to make sure the phenomenon would last, otherwise such a device would be useless. To our delight, every single one of the devices we tested worked over 10 million cycles. Not even one of them failed during our tests.</p><h3>The MOSFET Synapse</h3><br/><h3></h3><br/><img alt="Diagram of MOSFET showing biasing to increase or decrease channel conductance" class="rm-shortcode" data-rm-shortcode-id="0a7f1fb754b5958606940d5df8cd75df" data-rm-shortcode-name="rebelmouse-image" id="6010c" loading="lazy" src="https://spectrum.ieee.org/media-library/diagram-of-mosfet-showing-biasing-to-increase-or-decrease-channel-conductance.jpg?id=67005681&width=980"/><p><span>To be honest, we were amazed. Dozens of research groups and companies all around the world have spent many millions of U.S. dollars over the past 20 years trying to emulate these neural behaviors using experimental </span><a href="https://spectrum.ieee.org/memristor-first-single-device-to-act-like-a-neuron" target="_self">memristor-like devices</a> and other things, with limited success, mainly due to reliability and cost issues. We managed it in the cheapest and most industry-standard device: the MOSFET. This result was so shocking that we decided to confirm it using microchips from a different foundry. It was successful: All the behaviors could be reproduced, and perfect yield was achieved once again.</p><p>We were happy with the results and had started the process of filing for a patent and writing up our findings for the <a href="https://www.nature.com/articles/s41586-025-08742-4" target="_blank">journal <em><em>Nature</em></em></a>, when our lab made another astonishing discovery: The same kind of MOSFET could act as a synapse, too!</p><p>Recall that in ordinary operation some electrons crash into silicon atoms to create pairs of electrons and holes. We noticed that at specific values of bulk resistance a significant amount of the charge from this impact ionization would get trapped in the gate dielectric. This trapped charge interferes with the flow of current through the MOSFET, effectively changing the device’s conductance. Importantly, this new conductance is stable and adjustable at will. It was then that we realized the MOSFET could also be used as an electronic synapse.</p><p>As it was in the neuron transistor, the bulk terminal was the key. A negative bulk-source voltage drives electrons into the dielectric, decreasing conductance. A positive one pushes holes in, increasing it.</p><h2>From neuromorphic device to circuit to system</h2><p>Here’s how the MOSFET synapse and the MOSFET neuron, together called a neurosynaptic random-access memory, or NSRAM, could work together to achieve a simple neural circuit: Say you had a circuit consisting of three synapse MOSFETs and a neuron MOSFET. The synapses have already been programmed as we’ve described, so that each has a different conductance. Spikes of voltage with different patterns and frequencies are applied to the gate of each of these transistors. What emerges from their drains are spikes of current with amplitudes modulated by the synapses conductance values.</p><p>The spikes converge at the drain of the neuron MOSFET. With each spike, impact ionization causes charge to build in the bulk of the silicon. Some of it will drain away, but if enough spikes arrive in a short enough period of time, the bulk voltage will reach a value at which the “hidden” transistor triggers a spike of current through the MOSFET. This current would then go on to become the input to other MOSFET synapses, and so on. The behavior is exactly the kind of integrate-and-fire action real neural circuits deliver.</p><p>The competitive advantage of our single-MOSFET electronic neurons and synapses is straightforward: We can produce with only one or two transistors the electronic signals that today require, at an industrial level, dozens and sometimes even hundreds of components. And moreover, unlike other emerging technologies, our solution is fully compatible with today’s silicon manufacturing lines and exhibits a yield of 100 percent in key figures of merit with near-zero variability.</p><p>Building functional circuits for brain-inspired computing and AI based on this technology is as exciting as it is laborious. It will require us to improve our computer models to resemble the behavior of both devices more accurately and to do so with computational efficiency. We must also perform accurate circuit- and system-level simulations to validate computing architectures, design peripheral circuitry to drive and convert signals, and undergo multiple fabrication rounds to optimize performance.</p><p>But all that will be worthwhile, because it could result in brain-inspired microchips for AI with better energy efficiencies than what we have now. These chips will first be a fit for smaller-scale, “edge-AI” tasks, such as bringing greater intelligence to battery-powered systems. But if we can scale up such chips, maybe in the long run they can compete with state-of-the-art GPUs. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Mon, 29 Jun 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/artificial-neurons-on-silicon-chips</guid><category>Neuromorphic-computing</category><category>Cmos</category><category>Mosfet</category><category>Synapse</category><dc:creator>Mario Lanza</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/illustration-of-a-microchip-under-a-microscope-with-probes-and-orange-wires-attached.jpg?id=66967576&amp;width=980"></media:content></item><item><title>ConlangCrafter Turns AI to Imagining Languages</title><link>https://spectrum.ieee.org/conlangs-ai-model-contructed-languages</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/collage-of-nonsensical-words-and-accents-jumbled-together.jpg?id=67007771&width=1245&height=700&coordinates=0%2C472%2C0%2C473"/><br/><br/><p>There are <a href="https://www.ethnologue.com/" rel="noopener noreferrer" target="_blank">over 7,000 natural languages</a> today, but that doesn’t stop people from occasionally making up completely new ones. These constructed languages, or <a href="https://en.wikipedia.org/wiki/Constructed_language" rel="noopener noreferrer" target="_blank">conlangs</a>, include <a href="https://www.economist.com/books-and-arts/2017/08/05/the-complex-linguistic-universe-of-game-of-thrones" rel="noopener noreferrer" target="_blank">Dothraki</a>, <a href="https://www.kli.org/" rel="noopener noreferrer" target="_blank">Klingon</a>, and <a href="https://tolkiengateway.net/wiki/Elvish" rel="noopener noreferrer" target="_blank">various Elvish languages</a>. Now, an AI model called ConlangCrafter is also capable of generating new languages—and it is particularly good at it.</p><p>In a <a href="https://aclanthology.org/2026.acl-long.422/" target="_blank">paper</a> published 27 June in the <em><em>Proceedings of the Association of Computational Linguists, </em></em>researchers analyzed ConlangCrafter’s language-generation abilities, reporting that it can develop a diverse array of novel languages that consistently abide by their rules.</p><h2>How ConlangCrafter Creates New Languages</h2><p>In previous work, <a href="https://vcresearch.berkeley.edu/faculty/gasper-begus" rel="noopener noreferrer" target="_blank">Gašper Beguš</a>, an associate professor of linguistics at the University of California, Berkeley, showed how large language models (LLMs) <a href="https://spectrum.ieee.org/ai-linguistics" target="_self">can analyze languages </a>to the same extent as most humans. In his most recent endeavor, he set out to push the language boundaries of AI models even further.</p><p>“Creating an entire language is not an easy task at all,” Beguš says, noting that some people have dedicated their careers to creating conlangs for movies, books, and video games.</p><p>But Beguš sees additional value in making <a data-linked-post="2675056110" href="https://spectrum.ieee.org/ai-model-regulation" target="_blank">AI models</a> capable of creating truly novel languages beyond what humans could imagine. “[Models] are able to imagine or come up with things that we might not, and we can learn so much from that,” he says.</p><p>For example, ConlangCrafter can create new languages with unconventional communication systems, such as a language for a cephalopod species that uses colors and gestures instead of sounds. Of course, while this “color language” generated by ConlangCrafter isn’t truly what an octopus uses for communication, Beguš envisions these imaginary languages as a means for studying nonhuman-centric languages in greater detail.</p><p>Beguš and the rest of the team, including Morris Alper, a postdoctoral researcher at Carnegie Mellon University, and Moran Yanuka, a Ph.D. student at Tel Aviv University, designed ConlangCrafter so that it can apply a wide range of linguistic rules in terms of how sounds are organized in a language (phonology), the relationship between word and sentence structure (morphosyntax), and vocabulary.</p><p>A random number generator regularly introduces variation so that every language comes out different. A built-in editing loop then reviews the result for contradictions and fixes them. Users can choose whatever mix of rules they want, or ask ConlangCrafter to make up its own rules.</p><p class="pull-quote">“[Models] are able to imagine or come up with things that we might not, and we can learn so much from that.” <strong>—<span>Gašper Beguš, University of California, Berkeley</span></strong></p><p>“You can choose whatever flavor of language you want,” says Beguš. “You can create a mixed language between Japanese and Esperanto, for example.”</p><p>“The goal is for the languages to be creative, so they should all be different from each other,” says Alper, who specializes in multimodal machine learning and computational linguistics. “You also want them to be consistent, because a language is like a system of rules, and those rules shouldn’t contradict each other.”</p><p>To evaluate diversity, the team measured how much the generated languages differed from one another across key linguistic features such as the basic word order used in sentences. To evaluate consistency, they checked whether translations into each invented language correctly followed that language’s own rules.</p><p>They compared languages generated by ConlangCrafter to languages created by general-purpose LLMs, such as Gemini-2.5-Pro. “Our full system can be about twice as diverse and almost 70 percent more consistent than simply prompting an LLM to invent a new language,” says Alper.</p><h2>ConlangCrafter in Natural Language Processing</h2><p><a href="https://lti.cs.cmu.edu/people/faculty/mortensen-david.html" target="_blank">David Mortensen</a>, an assistant research professor at the Language Technologies Institute at Carnegie Mellon University who was not involved in the work, says that ConlangCrafter could help natural language processing researchers better evaluate the ways in which the structure of a language affects the performance of a model.</p><p>“There is a substantial body of research that suggests that linguistic structure–both at training time and test time–does affect model performance,” he says. “Hypotheses in this area have been very hard to evaluate, however.” He adds that a tool such as ConlangCrafter could help facilitate experiments on the effects of factors such as language typology and lexicon in a scientifically sound and reliable way.</p><p>ConlangCrafter is <a href="https://conlangcrafter.github.io/" target="_blank">available for free online</a>. Its creators note that the system is currently limited in more complex linguistic dimensions such as semantics, contextual and conversational use of language, and the visual aspects of writing.</p><p>Beguš envisions expanding upon this research to study the Sapir-Whorf hypothesis, which suggests that the way we speak influences the way we think and perceive the world. For example, this could involve running simulations of different worlds, each with its own language, exploring its impact on societies. “That’ll be a nice next step,” he says.</p><p><em>This story was updated on 29 June 2026 to correct Moran Yanuka’s name as well as the title of the </em>Proceedings of the ACL.<br/></p>]]></description><pubDate>Sat, 27 Jun 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/conlangs-ai-model-contructed-languages</guid><category>Llms</category><category>Artificial-intelligence</category><category>Languages</category><dc:creator>Michelle Hampson</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/collage-of-nonsensical-words-and-accents-jumbled-together.jpg?id=67007771&amp;width=980"></media:content></item><item><title>Why Does a Bank Need a Chief Scientist?</title><link>https://spectrum.ieee.org/capital-one-science-ai-finance-innovation</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/silhouetted-team-working-on-laptops-in-a-glass-walled-office-at-sunset.jpg?id=66903903&width=1245&height=700&coordinates=9%2C0%2C9%2C0"/><br/><br/><p><em>This article is brought to you by <a href="https://capitalone.science/" target="_blank">Capital One</a>.</em></p><p>After five years leading natural language understanding and eventually the entire Alexa AI organization at Amazon, Prem Natarajan made a nontraditional move: He became Chief Scientist at a bank. Not just any bank: Capital One, a financial institution serving over 100 million customers, helping everyday Americans manage their financial lives.</p><p>For Natarajan, a veteran of DARPA-funded research and academia who had watched machine learning evolve from task-specific applications to foundation models, the logic was clear. Some of the most interesting advances in AI research and deployment were shifting from big tech’s horizontal platforms to industry verticals like finance, where the most complex problems aren’t just building models but making AI work under the constraints of real-world customer problems, contextual business knowledge, continuous learning, with an incredibly high bar for accuracy and privacy.</p><p>That’s also what made Capital One the right place to do it. For decades, the company has been recognized as one of the most data- and analytics-driven financial institutions in the industry. Its business model from the very beginning was built around using data and technology to personalize financial products for customers. A decade ago, Capital One went all in on the cloud and rebuilt its data ecosystem, creating a unified environment for data, compute, and AI and machine learning experimentation. Today, its modern infrastructure, disciplined approach to governance, and deep bench of talent form the foundation that allows it to lead in enterprise AI.</p><p class="pull-quote">Advances in AI research and deployment are shifting from big tech’s horizontal platforms to industry verticals like finance.</p><p>So, why does a bank need a Chief Scientist? The answer lies in a fundamental misconception about AI in financial services. Most financial institutions still view AI as a technology to deploy – leveraging the latest large language model, deploying it through APIs, and integrating it into existing workflows – rather than a scientific discipline. Capital One is doing something different: building a scientific community and research organization to solve real-world customer problems and invent impactful AI solutions that don’t yet exist.</p><p>While widely available foundation models can handle general tasks, they can’t yet solve many domain-specific challenges, such as detecting fraud in real-time across billions of transactions, or providing state-of-the-art conversational tools so customers can engage when, how, and where they want to.</p><p>These challenges of making AI reliable, scalable, and well governed require original research and scientific innovation that is funneled back into the business to create real-world applications to address customer needs.</p><h2>The Constraints That Demand Innovation</h2><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Headshot of a suited man against a blue gradient background." class="rm-shortcode" data-rm-shortcode-id="475a0428edb65d212e3d3fb25a5b0e64" data-rm-shortcode-name="rebelmouse-image" id="9449b" loading="lazy" src="https://spectrum.ieee.org/media-library/headshot-of-a-suited-man-against-a-blue-gradient-background.jpg?id=66904023&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">Prem Natarajan, an IEEE Fellow, is Chief Scientist at Capital One. “If you want to solve really important problems in AI and see your work come to life, this is one of the few places you can do that,” he says.</small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Capital One</small></p><p>Because banks are dealing with people’s finances, there is an incredibly high bar for getting it right when it comes to AI. Take fraud, for example. Even a minor fraud event can have a devastating impact on certain customers. The best fraud models and platforms can detect and help mitigate fraud in the time it takes someone to tap their card, which is table stakes for protecting customers and their financial information with accuracy and speed. <span>Looking at these types of challenges, Capital One and Natarajan saw that serving millions of customers meant solving AI problems at a scale and complexity that many enterprises don’t encounter. These same constraints create a unique research environment.</span></p><p>At Capital One, the approach to building AI is to provide value to customers in ways never possible before, improving their financial lives and meeting them where they are with services they actually need. That focus, combined with massive scale and world-class risk management requirements, makes the scientific problems both harder and just as consequential as those found in most big tech labs.</p><h2>Advancing AI Through “Destination-Back Thinking”</h2><p><a href="https://www.capitalone.com/tech/ai-research/" target="_blank">Capital One’s approach to AI research and innovation</a> starts with what Natarajan calls “destination-back thinking.” Rather than asking what’s possible with current technology, the team envisions the customer experience they want to deliver – perhaps a car buyer who works long days and can only research the options at 10 p.m., or a customer facing an unexpected expense who needs immediate, personalized guidance – and then works backward to identify the scientific breakthroughs required to get there.</p><p>“You’re thinking back from where you’re providing incredibly valuable services,” Natarajan explains. “Once you have that vision clearly, you work back and say, what are the gaps? What are the things we need to invent?” This ensures that when problems are solved, the impact is essentially guaranteed, because the team has already identified what will make a tangible difference in customers’ lives.</p><p>But methodology alone isn’t enough. Capital One’s nearly 15-year bet on cloud-first architecture created something rare in financial services: a unified data and compute ecosystem that can support the kind of scientific experimentation typically seen in big tech research labs. As the only major U.S. bank to go all-in on public cloud infrastructure, Capital One eliminated the legacy systems that can constrain AI research at most financial institutions. This modern tech stack enables rapid iteration, large-scale model training, and what Natarajan calls “continuous learning,” systems that improve after deployment rather than degrading over time. This unique approach to infrastructure is a critical component in making new categories of research possible.</p><h2>Agentic AI: From Research to Production</h2><p>The research agenda manifests in systems already serving customers. Early last year, Capital One launched what may be the first fully agentic AI customer service experience built entirely in-house by a bank: a car buying tool that takes actions on behalf of customers based on their requests, not just answers questions. Behind it lies extensive research into multi-agentic AI reasoning systems that can navigate real-time data, business knowledge, constraints, and guardrails, with various agents that can work together to accomplish complex tasks.</p><p class="pull-quote">Capital One has launched a fully agentic AI customer service experience powered by extensive research into multi-agentic reasoning systems that can navigate real-time data.</p><p>The team is also working on solving things like tokenization challenges, protecting sensitive data while enabling model training. To accelerate this cutting-edge work, Capital One has established partnerships with Columbia University, the University of Southern California, and the University of Illinois, and became the only bank funding NSF’s national AI research centers <a href="https://www.nsf.gov/news/nsf-announces-100-million-investment-national-artificial" target="_blank"><span>in 2025</span></a>, investing millions in initiatives that span mental health, materials discovery, science, technology, engineering, and mathematics education, human-AI collaboration, and drug development.</p><p>In the spring of 2026, the company hosted its inaugural <a href="https://www.capitalone.com/tech/ai/2026-capital-one-ai-symposium/" target="_blank"><span>AI Symposium</span></a> to deepen connections and foster insight-sharing between the scientific AI community, leading AI labs, startups, and its own technology, science, and AI leaders and partners.</p><h2>Building a World-Class AI Organization</h2><h3></h3><br/><a class="rm-shortcode rm-image-link" data-rm-shortcode-id="e6efdd9602bbf40fa4c46c75e61a142d" data-rm-shortcode-name="rebelmouse-image" href="https://capitalone.science/" id="4a4bb" target="_blank"><img alt="Blue \u201cCapital One\u201d wordmark with a red swoosh above the text." class="" loading="lazy" src="https://spectrum.ieee.org/media-library/blue-u201ccapital-one-u201d-wordmark-with-a-red-swoosh-above-the-text.png?id=66904050&width=480&height=298&quality=100&coordinates=0%2C87%2C0%2C95"/></a><p>Capital One is building the next generation of AI talent. Join the team inventing impactful AI solutions to shape the future of finance. Learn more at <a href="https://capitalone.science/" target="_blank">https://capitalone.science/</a></p><p>External validation suggests the strategy is working. Evident AI <a href="https://evidentinsights.com/ai-index/" target="_blank"><span>ranked</span></a> Capital One as the leading bank in AI talent and a global leader in AI innovation for three consecutive years, noting the bank accounted for 38 percent of all AI patents filed by the top 50 financial institutions. Capital One was also recognized by <a href="https://www.ificlaims.com/news/ifi-insights-tracking-the-evolution-of-ai-with-patents/" target="_blank">IFI Insights</a> as the only financial institution among the top U.S. patent leaders in agentic and generative AI in 2025, alongside the likes of Google, NVIDIA, DeepMind, IBM, Microsoft, Intel, Adobe and Samsung. Capital One’s AI team – which has experience from leading AI labs and top universities – represents expertise rarely found outside Silicon Valley.</p><p>But recruitment requires a mission. “If you want to solve really important problems in AI and see your work come to life, this is one of the few places you can do that,” <a href="https://www.linkedin.com/in/natarajan/" target="_blank">Natarajan</a> says. The pitch is consistent: Capital One isn’t just optimizing algorithms for niche financial applications like high frequency trading, it’s using science to enhance financial experiences for over 100 million everyday Americans, expanding engagement and real-time insights, personalization, and access to their personal finances and products like never before.</p><p class="pull-quote">Capital One was recognized as the only financial institution among the top U.S. patent leaders in agentic and generative AI in 2025, alongside the likes of Google, NVIDIA, DeepMind, and Microsoft.</p><p><span>The frontiers Natarajan is most excited about – agentic AI systems that can dramatically improve performance by reframing how problems are solved, and domain-specific reasoning that understands contextual and financial nuance – represent the next phase of innovation. “By just casting the problem in an agentic framework, you can actually get way more performance” from the same underlying models, he explains.</span></p><p>It’s this kind of applied research, like translating general capabilities into production systems for millions of customers, that defines the <a href="https://www.capitalone.com/tech/culture/introducing-prem-natarajan/" target="_blank">Chief Scientist’s mandate</a>. When recruiting talent to his AI team, a group comparable only to the most sophisticated tech companies in caliber, Natarajan frames the opportunity around a mission. He invokes Steve Jobs’ famous challenge to John Sculley: “Do you want to spend the rest of your life selling sugared water, or do you want to change the world?” For Natarajan, the parallel is clear. Building AI systems that transform financial services for millions of everyday Americans – that’s changing the world. And it requires the kind of scientific rigor that only a Chief Scientist can lead.</p>]]></description><pubDate>Thu, 25 Jun 2026 17:32:32 +0000</pubDate><guid>https://spectrum.ieee.org/capital-one-science-ai-finance-innovation</guid><category>Ai-research</category><category>Agentic-ai</category><category>Financial-services</category><category>Tech-careers</category><category>Type-sponsored</category><category>Financial-technology</category><dc:creator>Thomas Machinchick</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/silhouetted-team-working-on-laptops-in-a-glass-walled-office-at-sunset.jpg?id=66903903&amp;width=980"></media:content></item><item><title>What It Means to Be a Mathematician When AI Does the Math</title><link>https://spectrum.ieee.org/ai-in-mathematics</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/a-photo-shows-a-man-standing-in-front-of-the-projection-of-a-computer-screen-thats-filled-with-computer-code.jpg?id=67007150&width=1245&height=700&coordinates=0%2C187%2C0%2C188"/><br/><br/><p><strong>In the mid-noughties, when</strong> music by the Killers and Franz Ferdinand blared out of every pub and nightclub I passed, I spent my days and nights struggling through a Ph.D. in applied <a href="https://spectrum.ieee.org/tag/mathematics" target="_blank">mathematics</a>. My research focused on simulating how special light waves interact in liquid crystals and using simple equations to approximate and understand those interactions. When I look back at my thesis now, liquid crystal technology is old hat, and I imagine my work could be completed with AI assistance in a matter of days—maybe hours.</p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/ai-in-mathematics?draft=1&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p class="shortcode-media shortcode-media-rebelmouse-image" style="display:none"> <img alt="" class="rm-shortcode" data-rm-shortcode-id="2b00fa4a3e2d69e8112f0268f9b668e5" data-rm-shortcode-name="rebelmouse-image" id="29a98" loading="lazy" src="https://spectrum.ieee.org/media-library/image.png?id=67033315&width=980"/></p><p>But the same cannot be said for the work of the pure mathematics Ph.D. students with whom I shared a cramped office at the University of Edinburgh. At the time, I felt sorry for these colleagues, who day after day sat at their desks, seemingly tearing their hair out and making no progress. (Though I was struggling too, I was at least always making some headway.) When we finished and went our separate ways, some hadn’t even published a paper.</p><p>Now, in hindsight, I finally understand why they toiled for years on abstract mathematical problems that only a handful of people in the world care about. It wasn’t arrogance, as I thought at the time; they weren’t trying to prove their superior intelligence by being the first to solve a seemingly intractable mathematical problem. It wasn’t even a form of masochism (which was my second guess)—penance for some imagined inadequacy. I realized they derived joy, satisfaction, and meaning from the long journey toward understanding.</p><h3></h3><br/><img alt="" class="rm-shortcode" data-rm-shortcode-id="6ee1c315c34c6dbd2e19a83348060be3" data-rm-shortcode-name="rebelmouse-image" id="b6640" loading="lazy" src="https://spectrum.ieee.org/media-library/image.png?id=67008692&width=980"/><p class="pull-quote">“Sometimes, understanding just strikes you as being very beautiful.” <strong>—Jeremy Avigad, Carnegie Mellon University</strong></p><h3></h3><br/><p>“Sometimes, understanding just strikes you as being very beautiful. Sometimes it’s a feeling of accomplishment, like completing a marathon,” muses Carnegie Mellon University mathematician <a href="https://www.cmu.edu/dietrich/philosophy/people/faculty/jeremy-avigad.html" rel="noopener noreferrer" target="_blank">Jeremy Avigad</a>. “But it’s not quite either of those: It’s just a wonderful feeling when you’ve been thinking long and hard about something complex, difficult, and then—all of a sudden—it just comes together.”</p><p>This feeling has driven mathematicians throughout history. Likewise, the way mathematicians pursue that feeling has changed little over the centuries. They notice or imagine links, patterns, or properties in numbers, shapes, or logical structures. From this, they write conjectures—unproven statements of their speculation. They or other mathematicians then use logical reasoning and the tools of mathematics in often creative ways to prove or disprove those conjectures. Finally, yet other mathematicians verify (or challenge) the proofs.</p><p>Invariably, this process requires a whole heap of thinking time. “I went to a pure maths camp with classes where we would sit with hard maths problems for half an hour and no one would say anything—everyone was just thinking,” says <a href="https://kammitama5.github.io/about/" rel="noopener noreferrer" target="_blank">Krystal Maughan</a>, a mathematician and computer scientist about to get her Ph.D. at the University of Vermont. “But then we would work together and kind of tease out the problem.”</p><p>This is the age-old joy of math in action. But today’s AI systems are starting to make inroads into bypassing this slow, deliberative process. Taking this trend to its logical conclusion, what happens if AI makes the mathematician’s struggle completely unnecessary? Might AI even sideline humanity completely?</p><h2>AI’s Growing Role in Mathematics<br/></h2><p>For decades, computation has accelerated mathematical progress. This began 50 years ago, when mathematicians used a computer to <a href="https://www.ams.org/journals/bull/1976-82-05/S0002-9904-1976-14122-5/S0002-9904-1976-14122-5.pdf" rel="noopener noreferrer" target="_blank">prove the four-color theorem</a>, which asks whether any map can be colored using no more than four colors, with no adjacent regions sharing the same color. The answer is yes, and the computer proved it, controversially, by checking 1,936 cases in a way no human could realistically verify.</p><p>Yet throughout this computational era, even in proofs relying on massive computational resources, the role of the human mathematician has remained central. Humans propose conjectures, guided by intuition. They devise strategies to prove them, guided by creativity and experience. And humans verify whether those proofs are correct.</p><p>Now AI is <a href="https://spectrum.ieee.org/ai-proof-verification" target="_self">challenging the status quo</a>. In just a few years, large language models (LLMs) have evolved from “<a href="https://dl.acm.org/doi/10.1145/3442188.3445922" rel="noopener noreferrer" target="_blank">stochastic parrots</a>,” capable of little more than regurgitating basic mathematics scraped from the internet, into advanced mathematical reasoning machines.</p><p>Last summer, systems from <a href="https://www.newscientist.com/article/2489248-deepmind-and-openai-claim-gold-in-international-mathematical-olympiad/" rel="noopener noreferrer" target="_blank">Google DeepMind and OpenAI</a> reached a level equivalent to the world’s most mathematically gifted high school students, achieving gold-medal status at the <a href="https://www.imo-official.org/" rel="noopener noreferrer" target="_blank">International Mathematical Olympiad</a>. In this annual competition, contestants must solve six notoriously difficult problems from various areas of mathematics.</p><p>Earlier this year, Google DeepMind’s experimental AI system Aletheia achieved an even more significant milestone when it <a href="https://doi.org/10.48550/arXiv.2601.23245" rel="noopener noreferrer" target="_blank">autonomously produced publishable Ph.D.-level research</a> results. While the work itself is obscure mathematically—calculating structure constants in arithmetic geometry—the significance lies in the complex reasoning it displayed in tackling an unsolved mathematical problem. And more recently, a new general-purpose AI system from OpenAI <a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/" rel="noopener noreferrer" target="_blank">disproved an important conjecture </a><a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/" rel="noopener noreferrer" target="_blank">in combinatorial geometry</a>. This result would have been worthy of publication in a major mathematics journal if humans had been the authors, and top mathematicians hailed the feat as a milestone for AI in mathematics, demonstrating independent, original, and sophisticated thinking.</p><p>Another shift has come from combining LLMs with mathematical tools known as proof assistants, which have been around for more than a decade. These systems—such as <a href="https://isabelle.in.tum.de/" rel="noopener noreferrer" target="_blank">Isabelle</a>, <a href="https://lean-lang.org/" rel="noopener noreferrer" target="_blank">Lean</a>, and <a href="https://rocq-prover.org/" rel="noopener noreferrer" target="_blank">Rocq</a>—are specialized programming languages that check mathematical proofs step-by-step, verifying their logical correctness. Traditionally, mathematicians have had to translate their theorems and proofs into this machine-readable format by hand, a laborious process known as formalization. Now, LLMs are starting to remove this bottleneck, automating the translation of informal proofs into formal code that proof assistants can verify.</p><div class="horizontal-rule"></div><h3>From Human Proof to Formal Proof</h3><br><p>Euclid’s famous proof that there are infinitely many prime numbers appears very different when formalized in Lean, a proof assistant. Human mathematicians routinely skip steps and rely on shared understanding; formalization makes every assumption and inference explicit so a computer can verify the proof.</p><h3></h3><br/><div style="max-width: 800px; margin: 0 auto; padding: 0 20px;"><h4><span style="background-color: black; color: white; padding: 2px 6px; font-family: sans-serif; display: inline-block; font-size: 50%"><strong>HUMAN PROOF</strong></span></h4><p>      We want to show that for every natural number <i>n</i>, there’s a prime <i>p</i> that is at least <i>n</i>.<br/>      Consider the smallest prime factor of <i>n</i>! + 1. Call it <i>p</i>. It is obviously prime.<br/>      To show <i>p</i> is at least <i>n</i>, assume, for contradiction, that it is not.<br/><i>p</i> then clearly divides <i>n</i>!, so it also divides (<i>n</i>! + 1) − <i>n</i>! = 1.<br/>      But this is impossible: <i>p</i> is prime, and 1 has no prime divisors.<br/>      So <i>p</i> is at least <i>n</i>.</p></div><div style="background-color: #E9E2D8; height: 10px; margin: 20px 0; width: 100%;"></div><div style="max-width: 800px; margin: 0 auto; padding: 0 20px;"><h4><span style="background-color: black; color: white; padding: 2px 6px; font-family: sans-serif; display: inline-block; font-size: 50%"><strong>LEAN PROOF</strong></span></h4><p style="font-family: monospace; font-size: 15px;">      /- Euclid’s theorem on the **infinitude of primes**.<br/>      Here given in the form: for every `n`, there exists a prime number `p ≥ n`. -/<br/><span style="color: red;">theorem</span> <span style="color: purple;">exists_infinite_primes</span> (n : ℕ) : ∃ p, n ≤ p ∧ Prime p :=<br/><span style="background-color: black; color: white; border-radius: 50%; display: inline-block; width: 1.2em; height: 1.2em; text-align: center; line-height: 1.2em; font-family: sans-serif; margin-right: 6px"><strong>1</strong></span><span style="background-color: yellow;"><span data-redactor-style="color: red;" style="color: red;">let</span> p := minFac (n ! + <span style="color: #0077aa;">1</span>)</span><br/><span style="color: red;">have</span> f1 : n ! + <span style="color: #0077aa;">1</span> ≠ <span style="color: #0077aa;">1</span> := ne_of_gt <| succ_lt_succ <| factorial_pos _<br/><span style="background-color: black; color: white; border-radius: 50%; display: inline-block; width: 1.2em; height: 1.2em; text-align: center; line-height: 1.2em; font-family: sans-serif; margin-right: 6px"><strong>2</strong></span><span style="background-color: yellow;"><span data-redactor-style="color: red;" style="color: red;">have</span> pp : Prime p := minFac_prime f1</span><br/><span style="color: red;">have</span> np : n ≤ p :=<br/>        le_of_not_ge <span style="color: red;">fun</span> h =><br/><span style="color: red;">have</span> h<sub>1</sub> : p ∣ n ! := dvd_factorial (minFac_pos _) h<br/><span style="background-color: black; color: white; border-radius: 50%; display: inline-block; width: 1.2em; height: 1.2em; text-align: center; line-height: 1.2em; font-family: sans-serif; margin-right: 6px"><strong>3</strong></span><span style="background-color: yellow;"><span data-redactor-style="color: red;" style="color: red;">have</span> h<sub>2</sub> : p ∣ <span style="color: #0077aa;">1</span> := (Nat.dvd_add_iff_right h<sub>1</sub>).<span style="color: #0077aa;">2</span> (minFac_dvd _)</span><br/>          pp.not_dvd_one h<sub>2</sub><br/>      ⟨p, np, pp⟩</p></div><h3></h3><br><p class="caption"><span style="font-size: 34px; vertical-align: -5px;">❶</span> Definitions must be explicit. The proof formally defines <em>p</em> as the smallest prime factor of <em>n</em>! + 1 before it can use that quantity.</p><p class="caption"><span style="font-size: 34px; vertical-align: -5px;">❷</span> Formal proofs build on earlier formal proofs. Here Lean invokes a previously verified theorem showing that <em>p</em> is prime.</p><p class="caption"><span style="font-size: 34px; vertical-align: -5px;">❸</span> Hidden logical steps become explicit. A human mathematician can write that <em>p</em> “clearly” divides 1. Lean requires the proof to invoke a formal theorem about divisibility and show exactly why that conclusion follows.</p><p class="image-media media-photo-credit" style="">With technical assistance from Sidharth Hariharan</p><h3></h3><br><div class="horizontal-rule"></div><p>Versions of such systems, sometimes called reasoning agents, are becoming highly sophisticated. In February, for example, the AI company <a href="https://www.math.inc/" target="_blank">Math, Inc.</a> used its aspirationally named reasoning agent <a href="https://en.wikipedia.org/wiki/Carl_Friedrich_Gauss" target="_blank">Gauss</a> to formalize a proof that had earned the mathematician <a href="https://people.epfl.ch/maryna.viazovska?lang=en" target="_blank">Maryna Viazovska</a>, of EPFL, in Switzerland, a <a href="https://www.mathunion.org/imu-awards/fields-medal/fields-medals-2022" target="_blank">Fields Medal</a> in 2022. Gauss first helped <a href="https://thefundamentaltheor3m.github.io/Sphere-Packing-Lean/" target="_blank">human mathematicians</a> complete the formalization of Viazovska’s solution to the <a href="https://annals.math.princeton.edu/2017/185-3/p07" target="_blank">8-dimensional sphere-packing problem</a> in a matter of days, and then <a href="https://www.math.inc/sphere-packing" target="_blank">autonomously formalized</a> the more complicated <a href="https://annals.math.princeton.edu/2017/185-3/p08" target="_blank">24-dimensional case</a> in just two weeks.</p><p>Such achievements suggest that AI is already capable of handling some mathematical tasks long considered uniquely human. As the technology advances, more of the day-to-day work of human mathematicians is likely to become fair game for AI.</p><h2>Mathematicians Debate AI’s Role in Discovery</h2><h3></h3><br><img alt="Person in a dark blazer with blurred face against a blue background" class="rm-shortcode" data-rm-shortcode-id="88a347d13f9eec920cc767d0d3c46b21" data-rm-shortcode-name="rebelmouse-image" id="851e9" loading="lazy" src="https://spectrum.ieee.org/media-library/person-in-a-dark-blazer-with-blurred-face-against-a-blue-background.png?id=67008550&width=980"/><p class="pull-quote">Human mathematicians could become “priests to oracles.” <strong>—Yang-Hui He, London Institute for Mathematical Sciences</strong></p><h3></h3><br><p>In September 2025, I attended the <a href="https://www.heidelberg-laureate-forum.org/forum/12th-hlf-2025-1/" target="_blank">12th Heidelberg Laureate Forum</a>—an annual conference that brings hundreds of young mathematicians and computer scientists together with their intellectual idols. AI dominated the conversation and, from the get-go, tension was in the air.</p><p>Speakers described a future in which superhuman AI mathematicians transcend human knowledge and capabilities: forming conjectures, searching solution spaces, proving conjectures, and finally verifying the proofs and generalizing the results, all without human involvement. If this future comes to pass, <a href="https://lims.ac.uk/yang-hui-he/" target="_blank">Yang-Hui He</a> of the London Institute for Mathematical Sciences memorably declared, human mathematicians could become “priests to oracles.”</p><p>While such startling predictions were being voiced on stage, my gaze was drawn to the audience. Frowning, fidgeting, and exchanging furtive glances—the crowd’s unease was palpable. <a href="https://experts.deakin.edu.au/65467-trill-white" rel="noopener noreferrer" target="_blank">Trill White</a>, a student at Australia’s Deakin University, later recalled sitting in that hall and thinking: “ ‘That’s devastating. What will people have to contribute to mathematics? Will it become something that no one understands?’ I did get a sense that this is going to change everything.”</p><h3></h3><br><img alt="Portrait of a long-haired person with blurred face on an orange background" class="rm-shortcode" data-rm-shortcode-id="ab1a17c74ca27d3d643cbc280f4e0b15" data-rm-shortcode-name="rebelmouse-image" id="a260e" loading="lazy" src="https://spectrum.ieee.org/media-library/portrait-of-a-long-haired-person-with-blurred-face-on-an-orange-background.png?id=67008467&width=980"/><p class="pull-quote">“We certainly started realizing AI has the potential to replace us.” <strong>—Jessica Randall, Google Developer Groups</strong></p><h3></h3><br><p><a href="https://www.linkedin.com/in/jessica-randall-293ab9205?originalSubdomain=za" rel="noopener noreferrer" target="_blank">Jessica Randall</a>, a South African mathematician for Google Developer Groups, says she sensed a collective existential dread rising among the young mathematicians. “I could feel everyone was worried, because they hadn’t thought that far ahead,” she says. “It was like a big bombshell that hit us, and we certainly started realizing AI has the potential to replace us.”</p><p>Some established mathematicians, including He, seem comfortable with AI taking on tasks that are currently the preserve of human mathematicians. That’s because they just want to know the answers to the biggest questions in mathematics—such as the six remaining <a href="https://www.claymath.org/millennium-problems/" rel="noopener noreferrer" target="_blank">Millennium Prize Problems</a>—even if AI does it all. “A lot of mathematicians are pragmatic and just want to understand. They would sell their soul for the solution to a problem,” jokes Avigad. “Whatever it takes, right?”</p><p>But this “just want to know” camp is by no means the only faction: Most mathematicians do not hope or expect AI to replace them entirely. Instead, two broad alternatives are emerging. The first is a human-centric aspiration that prioritizes human understanding of mathematics and treats AI as a tool, much like a calculator. The second is a collaborative “teamwork makes the dream work” vision, where humans and AI work together to tackle problems neither could solve alone.</p><h2>The Human Role in Mathematics</h2><h3></h3><br><img alt="Portrait of a person with blurred face on pink background" class="rm-shortcode" data-rm-shortcode-id="8a92c20a57f0f458eb134ab5afe6058c" data-rm-shortcode-name="rebelmouse-image" id="e688c" loading="lazy" src="https://spectrum.ieee.org/media-library/portrait-of-a-person-with-blurred-face-on-pink-background.png?id=67008214&width=980"/><p class="pull-quote">Numbers are “a way of bringing us to agreement.” <strong>—Akshay Venkatesh, Princeton University</strong></p><h3></h3><br><p><a href="https://www.mathunion.org/imu-awards/fields-medal/fields-medals-2018" rel="noopener noreferrer" target="_blank">Fields Medalist</a> and Princeton mathematician <a href="https://www.math.ias.edu/~akshay/" target="_blank">Akshay Venkatesh</a> has been thinking about this topic from the human-centric viewpoint for years. In 2022, he used his <a href="https://www.youtube.com/watch?v=N-TXcYI5C9E" target="_blank">Fields Medal Symposium</a> to implore the mathematics community to deeply consider what AI might mean for the practice of mathematics. At the time, the idea that AI could replace mathematicians seemed far-fetched. Now, he says, “we’re reaching the point where, for at least some tasks with abstract mathematical reasoning, computers are becoming competitive with humans.”</p><p>For Venkatesh, the question is not just what computers can do, but what mathematics is for. “Sometimes I think when we use numbers, it’s not so much that we are describing phenomena that are intrinsically numerical, but that we can all agree exactly what the numbers mean,” he says. “It’s a way of bringing us to agreement.”</p><h3></h3><br><h3></h3><br><img alt="A photo shows a woman standing in front of a chalkboard filled with mathematical formulas.  " class="rm-shortcode" data-rm-shortcode-id="11a1b45deb5e290db26fdec85c86e456" data-rm-shortcode-name="rebelmouse-image" id="40436" loading="lazy" src="https://spectrum.ieee.org/media-library/a-photo-shows-a-woman-standing-in-front-of-a-chalkboard-filled-with-mathematical-formulas.jpg?id=67007797&width=980"/><h3></h3><br><p>Mathematician and machine learning expert <a href="https://frasermaia.github.io/" target="_blank">Maia Fraser</a>, of the University of Ottawa, shares this sentiment. She says the joy she derives from mathematics is something distinctly human that integrates the subconscious and conscious mind. She describes starting with an intuitive sense that a certain thing should be true and gradually bringing out something that she can express in a rigorous proof. Communicating and sharing these deep-born thoughts is “a form of collective intelligence that is something beautiful about the human spirit,” she says.</p><p>By these arguments, an AI proof of a mathematical conjecture that has stubbornly resisted human efforts would be useful only if comprehensible to humans. “That the statement can be proved by AI is already useful information,” concedes Fraser. “But then it’s still an open problem to come up with an elegant, beautiful human proof.” Even if no such proof exists, she says, searching for it “is still a valuable endeavor.”</p><h2>AI and the Future of Mathematical Collaboration</h2><p>A more collaborative approach to AI in mathematics comes from <a href="https://www.math.ucla.edu/~tao/" target="_blank">Terence Tao</a>, who first competed in the math Olympiad at the age of 10. In 1986, 1987, and 1988, he won bronze, silver, and gold medals, respectively, making him the <a href="https://en.wikipedia.org/wiki/List_of_International_Mathematical_Olympiad_participants" rel="noopener noreferrer" target="_blank">youngest winner</a> of each of the three medals in Olympiad history. Now a <a href="https://www.mathunion.org/imu-awards/fields-medal/fields-medals-2006" rel="noopener noreferrer" target="_blank">Fields Medalist</a> and professor at the University of California, Los Angeles, he has earned a reputation as one of the most gifted mathematicians alive.</p><p>Unlike some of his peers, Tao is neither dismissive of AI nor fearful. Instead, he sees it as the catalyst for a fundamental shift in the discipline—a transition toward what he calls “big mathematics.” He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion’s share of the technical grunt work.</p><h3></h3><br><h3>Three Futures for AI in Mathematics </h3><br><table border="“0”" style="white-space: unset;" width="100%"><thead><tr><th style="background-color: #000000; color: #FFFFFF; width: 25%;"><br/></th><th style="background-color: #000000; color: #FFFFFF; width: 25%;">AI as a tool</th><th style="background-color: #000000; color: #FFFFFF; width: 25%;">AI as a partner</th><th style="background-color: #000000; color: #FFFFFF; width: 25%;">AI as an oracle</th></tr></thead><tbody><tr><td style="background-color: #000000; color: #FFFFFF; width: 25%;">Role of AI</td><td style="background-color: #DFD5C1; width: 25%;">Assistant</td><td style="background-color: #ecece9; width: 25%;">Collaborator</td><td style="background-color: #DFD5C1; width: 25%;">Autonomous researcher</td></tr><tr><td style="background-color: #000000; color: #FFFFFF; width: 25%;">What matters most?</td><td style="background-color: #DFD5C1; width: 25%;">Human understanding</td><td style="background-color: #ecece9; width: 25%;">Shared discovery</td><td style="background-color: #DFD5C1; width: 25%;">Answers</td></tr></tbody></table></br></br></br></br></br></br></br></br></br></br></br></br></br></br><p>Already, Tao is experimenting with this concept, <a href="https://github.com/teorth" target="_blank">working on problems</a> alongside scores of online collaborators, some using AI tools. “A hundred years ago, almost every mathematics paper was single author,” he says. “But now I collaborate with people I’ve never met—and maybe in the future, I won’t even know if they are AI or real people.”</p><p>The key to Tao’s vision is uniquely mathematical: formalization. When a proof is translated into code and checked step-by-step by proof assistants, it removes any chance of human error or dishonesty. This approach changes how collaboration works, because trust is established through verification rather than reputation or rapport. An idea from an unknown researcher or even an amateur can be taken seriously if it has a formal proof.</p><p>“If it wasn’t for this formal verification layer, opening projects up without any safeguards would just be a disaster,” adds Tao. “But in math, we can completely check and verify outputs, and this really filters out a lot of the rubbish.”</p><h2>The Risks of AI in Mathematics</h2><p>From the young researchers at the Heidelberg Laureate Forum to some of the biggest names in the field, mathematicians all seem to agree on one point: AI has the potential to transform their discipline. But there’s far less consensus on what that transformation will mean in practice.</p><p>Some worry about the accessibility of AI tools. Traditionally, mathematicians have required little more than intuition, training, and a pen and paper to advance their field. If this slow, deliberative process is no longer valued by society, and particularly by research funders, then mathematics could become an elitist activity, only practiced by select organizations that can afford to work with proprietary AI models.</p><p>Another concern is motivation. As AI systems take on more of the work, the incentive to engage deeply with difficult problems may weaken. Princeton’s Venkatesh says that the long human process of formulating and understanding a proof may be hard to justify, not just to funders, but even to mathematicians themselves. “There have been times where I’ve spent years thinking about something, and I’ve slowly struggled to understand it,” he says. “If your computer can do large chunks of that for you, will you have the motivation to spend that time?”</p><p>That concern extends to the next generation. If students can use AI to jump straight to answers, they most likely will. But every time they skip the struggle, they miss an opportunity to build the foundations of their own unique intuition. Over time, some worry, the next generation of mathematicians may suffer from a form of intellectual atrophy, unable to think outside the AI box that trained them.</p><p>In response to such fears, the mathematics community is taking action. Individuals are <a href="https://arxiv.org/abs/2603.03684" target="_blank">writing essays</a>, <a href="https://www.ias.edu/math/events/deepmind-mathai-workshop" target="_blank">organizing workshops</a>, and <a href="https://www.ams.org/journals/bull/2024-61-02/S0273-0979-2024-01836-9/viewer/?t=1774535950666" target="_blank">debating in journals</a>, while institutions and <a href="https://leidendeclaration.ai/" target="_blank">community groups</a> are developing <a href="https://publicationethics.org/guidance/cope-position/authorship-and-ai-tools" target="_blank">guidelines</a> for how AI should be used in research and publication. Indeed, mathematicians are applying the same rigor and curiosity that they use every day to reckon with the challenges of AI. Taken together, these efforts reflect a broad effort to try to retain control over the direction of mathematics in the era of AI.</p><p>So, is AI sucking the soul out of math? In one way, it is doing the opposite. It is forcing mathematicians to confront deep questions about what mathematics is, why they have devoted their lives to it, and the purpose math serves in society. At the same time, though, it is reshaping the practice of mathematics in a way that may be difficult to reverse.</p><p>“Mathematics makes me a better problem solver at normal problems, because it frames my mind to think in a very logical, rational way,” says Randall, who noted the existential dread at the Heidelberg Forum. “It helps with every aspect of my life.” As AI transforms mathematics, many researchers wonder whether future mathematicians will be able to say the same. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Thu, 25 Jun 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-in-mathematics</guid><category>Mathematics</category><category>Large-language-models</category><category>Llms</category><category>Stem-education</category><category>Google-deepmind</category><category>Openai</category><dc:creator>Benjamin Skuse</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/a-photo-shows-a-man-standing-in-front-of-the-projection-of-a-computer-screen-thats-filled-with-computer-code.jpg?id=67007150&amp;width=980"></media:content></item><item><title>AI Is Designing Radio Chips That Humans Couldn’t Even Imagine</title><link>https://spectrum.ieee.org/ai-radio-chip-design</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/abstract-rainbow-blocks-and-shapes-linked-by-flowing-blue-wave-lines-on-white-background.png?id=67001857&width=1245&height=700&coordinates=0%2C753%2C0%2C754"/><br/><br/><div class="ieee-summary intro-text"><h2>Summary</h2><ul><li>RFIC design is a complex “<a href="#darkart">dark art</a>” that limits progress in wireless technologies like 5G, autonomous vehicles, and satellite communications.</li><li>Princeton researchers use reinforcement learning and <a href="#inverse-design">inverse design</a> to rapidly create RFICs from scratch.</li><li>Diffusion models rapidly generate <a href="#novel">novel</a> or <a href="#human-interpretable">human-interpretable</a> RF layouts, achieving record performance and drastically reducing design time.</li><li><a href="#future-progress">Future progress</a> needs large, shared chip design datasets and open ecosystems so AI can learn universal electromagnetic and circuit behaviors.</li></ul></div><p><strong>Take a moment</strong> and try to imagine your life without the wireless advances of the past three decades.</p><p>Have you lost your luggage? What a shame AirTags have not been invented. The airline representative has promised to call with updates, so settle in for a long wait by the kitchen telephone, because there are no affordable cellphones. You’ll be stuck listening to whatever is on the radio while you wait, because there are no streaming services. That’s not even to speak of <a href="https://www.imdb.com/title/tt12908110/" target="_blank">all</a> <a href="https://www.imdb.com/title/tt0337921/" target="_blank">the</a> <a href="https://www.imdb.com/title/tt10530176/?ref_=ls_t_44" target="_blank">movie</a> <a href="https://www.imdb.com/title/tt7668870/" target="_blank">plots</a> that would have been ruined.</p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/ai-radio-chip-design?draft=1&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p><span>This is just a tiny sliver of how wireless technology makes itself felt in your day-to-day existence. The effects it has had on supply chains, infrastructure, and how the economy runs have been world-altering.</span></p><p>None of it would be possible without the radio-frequency integrated circuits that allow all our devices to unobtrusively send and receive information.</p><p>Now imagine what the further evolution of this technology will bring: Wide-spread <a href="https://spectrum.ieee.org/autonomous-vehicles-fuel-efficiency" target="_self">autonomous vehicles</a>, <a href="https://spectrum.ieee.org/quantum-communication-2667066423" target="_self">quantum communications</a>, <a href="https://spectrum.ieee.org/6g-network-infrastructure-bell-labs" target="_self">6G mobile service</a> and satellite communications. Continued momentum will depend on newer and more advanced versions of today’s RF chips.</p><p>But there’s the rub. Whereas the design of most of the world’s computing chips has been standardized into its own science, RF design has remained stubbornly in the realm of art. A dark art, even, that is mastered only through years of experience. As any sorcerer will tell you, the dark arts keep their own schedule. And that schedule is impeding progress not just in RF chip design but in every other technology that depends on it.</p><p>About seven years ago, in the wake of <a href="https://spectrum.ieee.org/alphago-wins-match-against-top-go-player" target="_self">AlphaGo’s victory over world Go champion Lee Sedol</a>, my students at <a href="https://www.princeton.edu/" target="_blank">Princeton</a> and I began to wonder: Could AI be taught this art as well? Recent successes suggest that, to a large extent, it can. Over the last few years, our group and other leaders in the field have started to develop <a href="https://ieeexplore.ieee.org/document/11509583" target="_blank">machine-learning-driven algorithmic methods for designing RFICs</a>. Some of the <a href="https://www.nature.com/articles/s41467-024-54178-1" target="_blank">resulting chips look more like modern art</a> than circuit layouts. Yet in many cases, the physical prototypes bested state-of-the art circuits in terms of performance. The real achievement, however, is that it took the AI orders of magnitude less time to conceive a working design than it would a human designer.</p><p>This is not about one or two RF chips. AI-enabled design could be the future of all RF design, and maybe much more.</p><h2>The Dark Art of RFIC Design</h2><p class="rm-anchors" id="darkart">So why do these chips all have to be crafted by hand? Why aren’t RFICs designed with an algorithmic synthesis process, much as CPUs and GPUs are?</p><p>The design of RFICs is an exercise in engineering across multiple physical domains. <a href="https://spectrum.ieee.org/the-long-road-to-maxwells-equations" target="_self">Maxwell’s equations</a>, operating across different spatial and temporal scales, govern how electromagnetic fields interact with active and passive devices that must be carefully codesigned for the chip to function. Alongside these are the laws of thermodynamics, which determine how heat is generated and removed during operation, as well as the mechanics of thermal expansion and contraction that dictate how reliably the chip and its packaging survive temperature changes.</p><div class="ieee-sidebar-large"><h3>AI Could Short-Circuit RFIC Design<span class="redactor-invisible-space"></span></h3><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Flowchart comparing slow human chip design steps with faster AI\u2011driven process" class="rm-shortcode" data-rm-shortcode-id="147b19614c03ec332e0fa6c1e953782a" data-rm-shortcode-name="rebelmouse-image" id="93ed9" loading="lazy" src="https://spectrum.ieee.org/media-library/flowchart-comparing-slow-human-chip-design-steps-with-faster-ai-u2011driven-process.png?id=67004535&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">The design of a radio-frequency integrated circuit requires human intuition and multiple, often-repeated optimization steps. The hope is that through an understanding of Maxwell’s Equations, an AI can be taught to short-circuit this process and quickly produce a design.</small></p></div><p>Simultaneously accounting for all the physical constraints these impose makes the design space almost impossibly large. Every decision involves complex priorities that often compete with one another, preventing the optimization of any of them.</p><p>To better understand the issue, let’s walk through the steps involved, after which you’ll better understand why a single new chip design takes years and tens to hundreds of millions of dollars.</p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Colorful close-up of a microchip die showing intricate circuits and connection pads" class="rm-shortcode" data-rm-shortcode-id="348b285796c19807d58b99fef6b027cf" data-rm-shortcode-name="rebelmouse-image" id="859b7" loading="lazy" src="https://spectrum.ieee.org/media-library/colorful-close-up-of-a-microchip-die-showing-intricate-circuits-and-connection-pads.png?id=67003840&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of a glowing gold microchip circuit with dense patterned components." class="rm-shortcode" data-rm-shortcode-id="be17a26f3182e5809b4bb5a83168963b" data-rm-shortcode-name="rebelmouse-image" id="0a4d6" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-a-glowing-gold-microchip-circuit-with-dense-patterned-components.png?id=67003835&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of a microchip die with intricate golden circuit patterns and pads." class="rm-shortcode" data-rm-shortcode-id="8af4cbe03fc4e977185f244cbdedb567" data-rm-shortcode-name="rebelmouse-image" id="97bd4" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-a-microchip-die-with-intricate-golden-circuit-patterns-and-pads.png?id=67003794&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of a patterned microchip die with intricate gold circuitry on a dark background" class="rm-shortcode" data-rm-shortcode-id="3ecc4ca634e5976e6f8e229a194914e7" data-rm-shortcode-name="rebelmouse-image" id="ee038" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-a-patterned-microchip-die-with-intricate-gold-circuitry-on-a-dark-background.png?id=67003789&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Close-up of an intricate gold microchip circuit pattern on a dark background" class="rm-shortcode" data-rm-shortcode-id="68692e2b86dba788350f88842bae6227" data-rm-shortcode-name="rebelmouse-image" id="985be" loading="lazy" src="https://spectrum.ieee.org/media-library/close-up-of-an-intricate-gold-microchip-circuit-pattern-on-a-dark-background.png?id=67003787&width=980"/></p><p class="shortcode-media shortcode-media-rebelmouse-image rm-float-left rm-resized-container rm-resized-container-25" data-rm-resized-container="25%" style="float: left;"> <img alt="Microscope view of intricate gold microchip circuitry with numbered frame \u201c6\u201d." class="rm-shortcode" data-rm-shortcode-id="ad13f8e600f90da8857b81eadbf5c1ec" data-rm-shortcode-name="rebelmouse-image" id="01801" loading="lazy" src="https://spectrum.ieee.org/media-library/microscope-view-of-intricate-gold-microchip-circuitry-with-numbered-frame-u201c6-u201d.png?id=67003784&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">Most of the area of radio-frequency integrated circuits is dominated by complex electromagnetic structures. Human-designed RFICs, like this broadband power amplifier [1], start with templates and follow a symmetric, understandable pattern. But freed from the constraints of human-designed templates and the need for humans to even understand the rationale of electromagnetic structures, power amplifier ICs [2–5] and low-noise amplifiers [6] can take on truly wild-looking yet efficient designs. </small><small class="image-media media-photo-credit" placeholder="Add Photo Credit...">SENGUPTA LAB</small></p><p>Let’s say you’re an engineer assigned to design a new 28-gigahertz <a href="https://ieeexplore.ieee.org/document/10136184" target="_blank">power amplifier</a> for a 5G-millimeter-wave handset. (This is the type of RFIC that boosts the 5G signals on your phone and transmits them to the antenna where they can be picked up by a distant base station). Where do you start?</p><p>RFIC design has some features in common with house building. Just as the blueprint for a house dictates the number of bedrooms and bathrooms to be built and the hallways connecting them, the blueprint for an RFIC—called the architecture—establishes the kinds of elements the RFIC needs to fulfill its intended function. Instead of rooms, the architecture includes, for example, the number of stages of amplification your power amplifier needs. Instead of hallways, it shows the paths that signals must take to get through those stages.</p><p><span>The blueprint for RFICs is actually mostly hallway</span><strong>;</strong><span> passive elements, like inductors and transmission lines, take up far more real estate than active elements like transistors.</span></p><p><span></span>Here’s why. As you have probably experienced yourself, a typical CPU’s transistors overheat when faced with operating frequencies of just a few gigahertz. The frequencies RFICs can operate at are higher by an order of magnitude—28 and 39 GHz for 5G signals, 26.5 to 40 GHz and even higher for satellite communications, and 77 GHz for <a href="https://spectrum.ieee.org/longdistance-car-radar" target="_self">automotive radar</a>. Under this onslaught, a CPU’s transistors would fail.</p><p>RFIC transistors avoid this fate because these chips cleverly manage the signal’s energy with careful electromagnetic design. This takes the form of byzantine networks of metal elements that dominate the chip’s real estate. These<em> </em>structures are geometrically regular, often symmetrical, and so intricately constructed they sometimes resemble lacelike filigree. But while they may look decorative, they are essential to the chip’s functioning.</p><p>Electrically speaking, these “hallways” work more like the chip’s plumbing<strong>. </strong>Like plumbing, this extensive labyrinth of passives confines electromagnetic energy only to the places it should be traveling around the chip.</p><p>The major challenge in RFIC design is putting all these elements together to ensure they work, just as constructing a house from its blueprints demands exact specs for load-bearing beams, pipes, and external walls. On an RFIC, the architecture needs to be realized with physically fabricable transistors and passive components that are connected just so, to permit the signal to travel through the chip and be processed. The way these devices are connected locally is what we call the circuit’s topology.</p><h2>The RFIC Design Process</h2><p>To make that power amplifier, then, your first step is to identify a candidate circuit template: The combination of structures that will meet the goals of a particular architecture with a specific circuit topology. Over the years, researchers have eased your burden by developing reusable design templates for specific functions. For example, templates suggest how many amplification stages a circuit needs (because sometimes, combining the output of two smaller amplifiers will result in better bandwidth and efficiency than you would get from a single larger one). And they suggest what the general configuration of the passive structures should be. Today there is an extensive library of such templates.</p><p>However, these can’t simply be used off-the-shelf, because each comes with trade-offs. Some have better gain at the expense of stability; some better bandwidth at the expense of efficiency; still others are more energy efficient at the expense of output power, and so on. There is rarely a clear best choice.</p><p>To arrive at the “sweet spot” where all these different parameters are balanced into optimal harmony, designers will typically lay out several different versions of the circuit, using intuitions and methods they have picked up in their years of training.</p><p>The challenge is that the decision around the architecture, circuit topology, or the electromagnetic passives cannot be done separately. One decision influences the others. So, designing an RF circuit can often feel like trying to fit an oversized carpet into too small a room—press down one corner, and another pops up.</p><p>At microwave and millimeter-wave frequencies, even the smallest misstep is the difference between a chip that works and one that doesn’t, and any number of things can go wrong. For example, when an electromagnetic wave encounters a transistor—or any other component —the path it travels must be properly “matched” to what comes next. If it isn’t, some of the energy reflects backward instead of flowing forward. Imagine trying to connect a high-pressure fire hose directly to a narrow garden hose. Without the right adapter, water will splash backward at the junction. Very little will make it through. In electronics, this is called the impedance-matching problem.</p><p>To prevent those reflections, engineers design special transitions, essentially microscopic adapters, that smooth the handoff between components. On a chip, these adapters can be surprisingly intricate. They don’t just pass the signal along; they can also split it, combine it, or distribute it across multiple paths with carefully controlled timing and strength.</p><p>Once you’ve done the architecture, plumbing, and everything in between comes the moment of truth. Have all the choices you have navigated through the enormous design space resulted in an RFIC that meets its specifications? If the specifications are not met, you will have to go back, either redoing the topology or the entire architecture, and repeat the whole process. So get ready for months of time- and resource-heavy simulation and iteration. Perhaps you now see why, for decades, a core belief has persisted in the RFIC community: “RF design is an art.” It was said that only an experienced designer—with an artisanal understanding of how the pieces make up the whole—could master the subtleties of analog and RF design. Unfortunately, this entrenched notion has long held back algorithmic innovations in the field just when we need them most. Traditional, artisanal RFIC design is hitting its limits as the complexity of these systems inexorably grows.</p><h2>AI for RFIC Design</h2><p class="rm-anchors" id="inverse-design">While RFIC designers continued their battle against their “oversized carpet” problem, a series of interesting developments emerged in allied disciplines. Across a range of other previously intractable problems like <a href="https://spectrum.ieee.org/alphafold-proves-that-ai-can-crack-fundamental-scientific-problems" target="_self">protein folding</a> and <a href="https://www.weforum.org/stories/2023/12/ai-weather-forecasting-climate-crisis/" target="_blank">climate modeling</a>, AI has been able to successfully navigate multidimensional complex spaces. This gave us the incentive to look deeper into AI for RF. After all, the combinatorial complexity of protein folding is not that different from the nature of the design space in our domain.</p><p>We were not the first to think of using artificial intelligence to speed up parts of RFIC design. Researchers had previously trained machine learning algorithms on circuit templates in the hope of speeding up the normal optimization processes. While this approach was undoubtedly faster than humans at optimizing templates, it still relied fundamentally on libraries of existing designs invented by humans.</p><div class="ieee-sidebar-medium"><h3>Training an AI to Design a Chip</h3><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Flowchart of RL and generative AI optimizing RFIC electromagnetic networks" class="rm-shortcode" data-rm-shortcode-id="5e5e09d828d38e666d83907d447c8b98" data-rm-shortcode-name="rebelmouse-image" id="6fbd7" loading="lazy" src="https://spectrum.ieee.org/media-library/flowchart-of-rl-and-generative-ai-optimizing-rfic-electromagnetic-networks.png?id=67003985&width=980"/><small class="image-media media-caption" placeholder="Add Photo Caption...">A machine learning system learns to do end-to-end RFIC design like other AIs learned to play such games as Go. Essentially, it turns the process into a game, learning from the results of its own efforts.</small></p></div><p>We didn’t want that. We wanted to break free from the restrictions of prefabricated topologies. Because while a designer’s experience and hard-won heuristics are crucial to building a working design, they also place fundamental limits on it. Furthermore, such an approach would necessarily require simulation steps as part of the optimization cycle, and even the fastest simulations use a lot of computing resources. Worse still, in many advanced cases, such as for broadband designs, there are no existing templates.</p><p>But if we didn’t start with templates, where could we start?</p><p>The goal here was to allow algorithms to determine—entirely from scratch—every parameter for architecture, constituent circuits, and electromagnetic passives. This approach differs fundamentally from conventional optimization, which is limited to determining the parameters—like transistor dimensions and passive component geometries—that optimize structures originally devised by humans.</p><p>In our new approach, the architecture begins essentially from nothing and is progressively assembled through successive iterations. The system explores the design space by generating myriad candidate circuit combinations and mapping the resulting performance trade-offs as it navigates this landscape. Because the process is not biased by prior human design choices, it can produce completely novel circuit topologies that look markedly different from those created by human designers.</p><p>In some ways, the approach echoes AI systems such as <a href="https://spectrum.ieee.org/alphago-zero-goes-from-blank-slate-to-grandmaster-in-three-dayswithout-any-help-at-all" target="_self">AlphaGo Zero</a>, which achieved superhuman performance not because it was trained on games played by humans but because it explored the rules by playing against itself. Similarly, our algorithm develops new circuit architectures by exploring and evaluating its own design strategies. In so doing, it learns to understand circuits, electromagnetics, and the close codesign they need to achieve the end-to-end design of RFIC.</p><h2>Inverse Design for RFICs</h2><p>To realize this capability, we proceeded in two stages. First, we developed a <a href="https://spectrum.ieee.org/reinforcement-learning-environments" target="_self">reinforcement-learning</a> (RL) framework that determines the optimal system architecture, circuit topology, device parameters, and even the properties of the electromagnetic interfaces that connect different circuit elements. In this stage, the algorithm effectively defines how signals should propagate and interact across the system.</p><p>The algorithm trains very similarly to how a computer learns to play a game. If you let it play enough times, it can learn to play better by observing the relationship between the actions it took and the score it achieves. In a similar way, the RL agent here learns to design effective circuits by playing with a set of combinations, and over time, it can map the space between the circuit performance to its architecture, topology, and parameters. This training takes a few days to a week, but once trained, the agent can design circuits very quickly</p><p>The next step was to determine the physical structure of the IC’s electromagnetics—the plumbing—that can create the desired properties of the passive elements, which are characterized by a set of metrics called scattering parameters. These measure if a signal entering a component actually moves forward—or is reflecting backward, being wasted, as in our previous example with the fire hose and the garden hose.</p><p>Deriving the structure from the desired scattering parameters is an example of an approach called inverse design, which appears across many areas of engineering. In structural engineering, for example, one might collaborate with an architect on a physical goal—such as creating large interior spaces with high ceilings—and then determine the arrangement of arches or buttresses that can support it.</p><h3>Generative AI for Electromagnetic Networks</h3><br/><img alt="Diagram linking S-parameter curves to classical, mazelike, and pixelated structures." class="rm-shortcode" data-rm-shortcode-id="1aaf5ec91b9c52d0e4db55d0bf00a331" data-rm-shortcode-name="rebelmouse-image" id="027de" loading="lazy" src="https://spectrum.ieee.org/media-library/diagram-linking-s-parameter-curves-to-classical-mazelike-and-pixelated-structures.png?id=67004312&width=980"/><p>But RF integrated crcuits pose a particular challenge for inverse design: The process must account simultaneously for circuit behavior and the electromagnetic responses of the interconnects and passive elements that link them together. But it has to figure that out without doing a lot of artisanal iterating.</p><p>So we replaced our RF circuit simulator with an AI-based emulator. This AI model can predict the behavior of electromagnetic fields going through any structure—even totally arbitrary two-dimensional shapes—without having to compute the underlying physics from scratch, as simulation tools do. It would predict the solution of Maxwell’s equations and tell you the scattering parameters for any structure you showed it, without actually doing the math. With such an AI in hand, what a time-consuming electromagnetic solver normally takes minutes or hours to accomplish is reduced to milliseconds.</p><p>We chose to build our emulator around a <a href="https://spectrum.ieee.org/facebook-ai-director-yann-lecun-on-deep-learning" target="_self">convolutional neural network</a>—a machine learning model that has been remarkably successful for image processing. Such networks can extract spatial features from any structure, and it turns out that the image of a structure contains a lot of spatial information that can accurately predict its electromagnetic performance. Then we trained it on a vast number of random pixelated structures whose scattering parameters had been labeled.</p><p>Once we had our inverse-design RL and suitable AI emulator, we essentially had an <a href="https://ieeexplore.ieee.org/document/10904600" target="_blank">end-to-end AI designer</a>. So we asked it to design us a power amplifier.</p><h2 class="rm-anchors" id="novel">Unconventional RF Architectures</h2><p>In 2023, <a href="https://ieeexplore.ieee.org/document/10136184" target="_blank">we published this proof of concept</a>—a power amplifier targeting the millimeter-wave band, specifically spanning 30 to 100 GHz, which covers most of the relevant 5G and radar frequencies. The final design achieved the best combination of wide bandwidth, output power, and efficiency then reported for a silicon-based power amplifier—meaning it could amplify a large amount of data across a wide swath of frequencies—while maintaining record efficiency.</p><p>The structure of the IC’s electromagnetic pathways was unlike anything any human would ever consider. Since the AI is not trained on human designs, the layout that emerged looked more like an arbitrary pattern or perhaps a QR code than the regular symmetrical structures we are used to seeing.</p><p>One unexpected insight revealed by this prototype, and our research generally, is that there’s no evidence that the templates we’ve historically relied on are even close to optimal for modern design goals. It’s not that a human designer can never come up with a better design. But with the removal of the templates and the time to synthesize cycle upon cycle of optimized circuits, it is now clear that AI-driven synthesis could break traditional design barriers and push the limits of RFIC capabilities.</p><p>Our 5G amplifier had only one input port and one output port. Adding more inputs and outputs to a design is not straightforward. Every port electromagnetically couples to every other port, so the scattering parameters quickly add up. Two ports give you four scattering parameters. Four ports, 16 scattering parameters. The math gets ugly fast. Could our model keep up?</p><p>We next trained our model on larger classes of electromagnetic structures with many input and output ports. In 2024, we published work showing that <a href="https://ieeexplore.ieee.org/document/10600352" target="_blank">multiport integrated circuits</a> are no problem for these AI algorithms either. Where previously multiport electromagnetic simulation required days or weeks of toil, this model evolved new structures in minutes. Since then, a plethora of work in the space by research communities across the globe have demonstrated the power of inverse design in RFIC.</p><p>Combining the reinforcement learning framework with the inverse design, we now had the ability to create an RFIC from specifications all the way to a <a href="https://ieeexplore.ieee.org/document/11015614" target="_blank">fabrication-ready layout</a>. We’ve so far shown this is true for RFICs ranging from low-noise amplifiers to <a href="https://www.nature.com/articles/s41467-024-54178-1" target="_blank">subterahertz</a> and broadband <a href="https://doi.org/10.1109/ISSCC49661.2025.10904600" target="_blank">power amplifiers</a><em><em><strong>.</strong></em></em> The hope is that this will work just as well for other circuits.</p><h2 class="rm-anchors" id="human-interpretable">Making AI Designs Interpretable</h2><p>Our goal was to make RFIC design better and easier, but we didn’t want to make it beyond human understanding. Chip testing and debugging is a long, arduous process, sometimes even more so than design. Engineers often prefer ICs to have interpretable structures, so that if a problem crops up, they can understand how the chip works well enough to debug it.</p><p>To create structures that are more interpretable, we turned to <a href="https://spectrum.ieee.org/ai-art-generator" target="_self">diffusion models</a>, which you may know from their remarkable ability to generate realistic images from text prompts.</p><p class="pull-quote">AI-driven synthesis could break traditional design barriers and push the limits of RFIC capabilities. </p><p>Imagine you go to your favorite image-generation engine and ask it to create a painting of the sky in the style of Picasso, Van Gogh, or Michelangelo. You will get images that capture the essence of their brushstrokes, their use of colors, and their framing. All are pictures of the sky nonetheless, but in different styles.</p><p>Electromagnetic design is similar in that multiple structures can have very similar electromagnetic responses. Instead of using text input, we used scattering parameters as our input, and the electromagnetic structure of an RFIC chip as our output.   As part of the inputs to the <a href="https://ieeexplore.ieee.org/abstract/document/11103838" target="_blank">diffusion model</a>, we created a <a href="https://ieeexplore.ieee.org/document/11409170" target="_blank">dial that sets the spatial frequency of the final structure</a>. By turning the dial, a designer can direct the model to synthesize structures with low (classical-looking and interpretable), medium (mazelike structures), or high (pixelated or arbitrarily-shaped) spatial frequency.</p><p>From prompts to output, the entire process took about 6 minutes. With this diffusion model, algorithms can now both discover novel architectures <em><em>and </em></em>accelerate the creation of conventional, so-called classical ones.</p><p>All an RFIC designer needs to do is specify virtually any valid set of scattering parameters. As long as they are physically realizable under Maxwell’s equations, the model pops out a corresponding structure as if it were a vending machine.</p><h2 class="rm-anchors" id="future-progress">The Future of AI-Driven RFIC Design</h2><p>The results of our investigations have drawn the attention of the RF community. The traditional bottom-up design process is clearly beginning to reverse.</p><p>But there are still questions: How generalizable are these methods? Can they consistently deliver truly high performance? Can we get to a place where AI produces designs that maximize every conceivable trade-off, holistically optimizing every parameter to its most ideal physical state? We want to take this strategy beyond RFIC design and invent other kinds of circuits that are different from anything humans have ever done.</p><p>These are exciting and ambitious prospects, but we are not there yet. AI can hallucinate a design that creates bad circuits that don’t work. This means verification methods need to remain under human oversight. And, while hallucinations are rare, it would still be good to reduce their occurrence.</p><p>History suggests that meeting these dreams of the future will take much more data than we’ve been using. Before the creation of the ImageNet repository—a repository of 14 million varied, human-annotated images—image-recognition models didn’t function well in the real world. The datasets they had been trained on were too tiny to be effective. ImageNet’s massive amounts of training data ushered in a revolution that led to AI that can generalize and recognize images in the wild. The rest was history.</p><p>If the goal for RFIC and analog design is a universal foundational model—something that learns the governing laws of electromagnetics and circuit behavior—then we also need data.</p><p>The good news is that this data is plentiful. Around the world, countless engineers at companies and academic labs simulate nearly identical RF circuits and passive structures every day. The bad news is that it’s all locked away behind nondisclosure agreements.</p><p>Open ecosystems have propelled other areas, and we think the RFIC community should do the same. There had been some movement toward this. <a href="https://spectrum.ieee.org/natcast-layoffs" target="_self">Natcast</a>, the operator of the <a href="https://www.nist.gov/chips/research-development-programs" target="_blank">U.S. CHIPS and Science Act’s R&D program</a>, would have bolstered shared infrastructure and innovation for the next generation of wireless, sensing, and defense technologies. Unfortunately, both the organization and the <a href="https://www.nist.gov/chips/princeton-university-princeton" target="_blank">program</a> it ran specifically for machine learning and RFICs have been closed.</p><p>But the momentum Natcast’s effort sparked hasn’t died out. Building on our early work, groups across the community have already demonstrated remarkable advances. AI-driven IC design is part of a much broader technological shift. From biology and materials science to automotive and aerospace engineering, AI is reshaping how complex systems are conceived and optimized. Deeper collaboration between AI researchers and chip designers will unlock the field’s full potential. It’s by no means a foregone conclusion, but if we get this right, this genie won’t stay in its bottle. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Wed, 24 Jun 2026 13:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/ai-radio-chip-design</guid><category>Machine-learning</category><category>Ic-design</category><category>Chip-design</category><category>Rf</category><category>Rfic</category><dc:creator>Kaushik Sengupta</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/abstract-rainbow-blocks-and-shapes-linked-by-flowing-blue-wave-lines-on-white-background.png?id=67001857&amp;width=980"></media:content></item><item><title>AI Is Learning to Read the Room</title><link>https://spectrum.ieee.org/emotion-ai-context</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/pixel-art-figure-in-a-colorful-digital-cube-with-shadow-and-connected-emoji-faces.png?id=66966345&width=1245&height=700&coordinates=0%2C237%2C0%2C238"/><br/><br/><p><strong>Imagine sitting down at </strong>your desk and logging in for a performance review, with an AI system analyzing the conversation. You’ve been working long hours, balancing deadlines, and your manager asks how you’re doing. You say you’re fine, and maybe even smile, but there’s a hint of hesitation and your voice wavers. As you shift your posture, your shoulders slump.</p><div class="rm-embed embed-media"><iframe height="110px" id="noa-web-audio-player" src="https://embed-player.newsoveraudio.com/v4?key=q5m19e&id=https://spectrum.ieee.org/emotion-ai-context?draft=1&bgColor=F5F5F5&color=1b1b1c&playColor=1b1b1c&progressBgColor=F5F5F5&progressBorderColor=bdbbbb&titleColor=1b1b1c&timeColor=1b1b1c&speedColor=1b1b1c&noaLinkColor=556B7D&noaLinkHighlightColor=FF4B00&feedbackButton=true" style="border: none" width="100%"></iframe></div><p><span>These are subtle cues that to the human eye might hint at underlying stress. But to an AI model that’s been trained only to categorize emotions as “happy” or “sad,” such nuances are likely lost. It logs the words and a smile and moves on—and unless your human manager intervenes, the fact that you’re tired, unfocused, and maybe a couple of days from burnout never enters the equation.</span></p><p>“<a href="https://spectrum.ieee.org/building-an-ai-that-feels" target="_blank">Emotion AI</a>,” which estimates how people feel based on facial expressions, voice tone, and behavior, seems to be suddenly everywhere; it’s being used in employee well-being and recruitment interviews, education platforms, and driver-monitoring systems. Technology call-center platforms such as <a href="https://www.nice.com/" target="_blank">NiCE</a> and <a href="https://www.genesys.com/" rel="noopener noreferrer" target="_blank">Genesys</a> use AI to detect when a customer sounds frustrated and prompt agents in real time to slow down or respond with more empathy. Giant companies like <a href="https://raveintelligence.com/meta-voice-ai-surge-emotional-intelligence/" rel="noopener noreferrer" target="_blank">Meta</a> and startups such as <a href="https://www.hume.ai/" rel="noopener noreferrer" target="_blank">Hume AI</a> are developing more-expressive voice AI systems that can detect emotional cues in the person they’re “talking” to and adjust how they communicate.</p><p>What’s more, hundreds of companies already offer virtual AI companionship apps, a fast-growing market that may be worth an <a href="https://www.sphericalinsights.com/reports/ai-companion-market#:~:text=Table_content:%20header:%20%7C%20Base%20Year:%20%7C%202024,CAGR:%20%7C%202024:%20CAGR%20of%2031.05%25%20%7C" rel="noopener noreferrer" target="_blank">estimated US $555 billion</a> by 2035—and robot buddies have also entered the picture. Intuition Robotics’s <a href="https://elliq.com/?srsltid=AfmBOoqjBb7RoBuC0piFi5F-u5d64LbS_BVhLwG79xwEbTnrZwBx86fR" rel="noopener noreferrer" target="_blank">ElliQ</a>, for example, is a small device vaguely resembling a white desk lamp that’s now being used to engage older adults in conversation in hopes of reducing loneliness.</p><p>But while the field of emotion AI is advancing at a rapid clip, most existing systems are focused on detecting a limited number of signals to label one specific emotion at a time—which is insufficient if you’re trying to understand the human condition. In the real world, human signals and emotions are contextual, overlapping, and constantly changing. A laugh can signal joy, nervousness, or both; a raised voice might signal enthusiasm just as easily as frustration. To make the job of emotion detection even more difficult, reactions differ greatly from one individual to the next, depending on demographics, cultural background, and countless other variables.</p><p>In other words, there’s a gap between what we’re expecting AI to pick up on and what AI can actually deliver. That’s the gap a new field of research—what we call human-context AI—is working to close. Instead of looking at just one input and labeling it, human-context AI increasingly has the capacity to take stock of an individual’s personality and character, and to track emotions in real time while combining <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12292624/" rel="noopener noreferrer" target="_blank">multiple inputs</a>, including facial dynamics, voice, tone, language, and behavior. Crucially, responses are also evaluated in the context of a specific environment, such as a performance review or professional coaching session. The result? Computers are learning to read the scene, rather than just the screen.</p><h2>The Origins of Emotion AI</h2><p>The story of emotion-sensing AI began almost three decades ago in the MIT Media Lab, where the American electrical engineer and computer scientist <a href="https://spectrum.ieee.org/how-and-why-companies-will-engineer-your-emotions" rel="noopener noreferrer" target="_blank">Rosalind Picard</a> coined the term “affective computing.” Her work introduced the radical idea that computers could be taught to recognize and respond to human emotions.</p><p>Picard’s <a href="https://cs.uwaterloo.ca/~jhoey/teaching/cs886-affect/papers/Picard-AffectiveComputing/9780262281584_chap6.pdf" rel="noopener noreferrer" target="_blank">early experiments</a> focused on single modalities: facial expressions, tone of voice, and physiological signals, such as skin conductance or heart rate. The goal was to give machines a window into human feeling, helping them become more empathetic. It was an exciting vision, but back then the science and hardware weren’t ready. Computing power was limited, sensors were crude, and datasets were narrow and biased.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Pixel art of three party-hatted figures in a box, each losing a slice of cake." class="rm-shortcode" data-rm-shortcode-id="915714dd60f44acd05b8adbdd1ed711f" data-rm-shortcode-name="rebelmouse-image" id="af91e" loading="lazy" src="https://spectrum.ieee.org/media-library/pixel-art-of-three-party-hatted-figures-in-a-box-each-losing-a-slice-of-cake.png?id=66966369&width=980"/> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Josie Norton</small></p><p>Over the next decades, researchers and companies got better at measuring the many ways in which humans express themselves. In the 2010s, <a href="https://en.wikipedia.org/wiki/Sentiment_analysis" target="_blank">sentiment analysis</a>—the processing of large volumes of text to suss out emotional undertones—began to reach the mainstream. At the same time, marketing firms, including my company, <a href="https://www.neurologyca.com/" target="_blank">Neurologyca</a>, began using video and webcams to measure and catalogue customer reactions. Biometric devices and activity trackers, such as Fitbits and Apple watches, also became ubiquitous, generating new streams of data about people’s sleep, step counts, stress levels, and more.</p><p>Unsurprisingly, scientists soon confirmed that larger volumes of personalized data led to greater accuracy in reading human emotions. In 2019, researchers at Cornell demonstrated that <a href="https://arxiv.org/abs/1905.07039" target="_blank">combining multiple types of signals</a> improves emotion sensing. Their system joined physiological data, such as brain activity measured by electroencephalography (EEG) and heart rate, with visual cues like facial expression, outperforming systems that relied on just one input. Around the same time, Picard and her team at MIT found that humanoid robots <a href="https://news.mit.edu/2018/personalized-deep-learning-equips-robots-autism-therapy-0627" rel="noopener noreferrer" target="_blank">trained on data unique to a specific person</a> were substantially better at reading that person’s reactions and feelings than robots acting without personalized data.</p><p><span>More recent studies align with these findings. In 2024, <a href="https://www.sciencedirect.com/science/article/abs/pii/S095741742400589X" target="_blank">scientists in South Korea</a> showed that fusing physiological, environmental, and personal data to recognize emotion resulted in a 32 percent error reduction. <a href="https://dl.acm.org/doi/10.1145/3746270.3760232" target="_blank">Another paper, published in 2025</a>, demonstrated that user-specific information significantly enhances emotion recognition performance.</span></p><p>Today, our devices know who we are; our habits and tendencies, likes and dislikes. They’ve also gotten smaller and more efficient. Tiny, low-power cameras and microphones embedded in phones, laptops, and virtual-reality and augmented-reality devices can detect dozens of human signals simultaneously, from eye movements and micro-expressions to breathing rhythms, voice modulation, and posture. Advances in computing have also made it possible to integrate audio, video, biometric, and text data, often without even transmitting raw data to the cloud. And researchers at <a href="https://vhil.stanford.edu/publications/predictive-analytics/cognitive-load-inference-using-physiological-markers-virtual" rel="noopener noreferrer" target="_blank">Stanford</a>, <a href="https://www.cl.cam.ac.uk/~pr10/publications/ptb09.pdf" rel="noopener noreferrer" target="_blank">Cambridge and MIT</a>, and <a href="https://sap.ist.i.kyoto-u.ac.jp/lab/bib/intl/LAL-AAAI-sympo17.pdf" rel="noopener noreferrer" target="_blank">Kyoto University</a>, in Japan, as well as <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12292624/" rel="noopener noreferrer" target="_blank">the Software College of Northeastern University </a>in Shenyang, China, are exploring how fusing such inputs can refine the sensitivity and accuracy of human-machine interactions.</p><p>And yet, despite so many breakthroughs, machines still can’t reliably interpret emotion or even physical stress. Just last year, a survey published in the<a href="https://psycnet.apa.org/doiLanding?doi=10.1037%2Fabn0001013" rel="noopener noreferrer" target="_blank"> <em><em>Journal of Psychopathology and Clinical Science</em></em></a> revealed that stress scores on smartwatches rarely, if ever, matched the level of stress that users were experiencing. In fact, a quarter of those surveyed reported feeling the direct opposite of what their smartwatches were reporting.</p><p>Why the disconnect? We’ve gotten very good at capturing signals, but not at interpreting them. A fitness tracker might infer from your heart rate that you’re stressed and recommend easing off training, but it doesn’t know if your increased heart rate is due to excitement, tiredness, or an extra cup of coffee. Gauging emotions in real-world settings is even more difficult. To solve this complex problem, machines need context.</p><h2>From Neuromarketing to Emotion-Sensing AI</h2><p>My company, Neurologyca, was founded in Spain in 2015, and started out in neuromarketing. Working with major European brands and conglomerates, our cofounder, Juan Graña, had realized that companies lacked solid data on consumers. At the time, most customer feedback came through surveys, which posed questions such as, “On a scale of 1 to 10, how joyful does this car advertisement make you feel?” or “Which emoji best describes your mood?” Naturally, these overly simplistic tools led to high levels of self-reporting bias, as people often misjudge or misstate their own reactions.</p><p>To get around this problem, Neurologyca set up labs, using neuroscience and cognitive science to more accurately capture human responses to products, logos, advertisements, and experiences. In addition to using biometric tools such as heart monitors, eye trackers, and EEG, we recorded millions of video frames of human reactions, logging each specific context and the resulting facial and bodily movements. To do this, we mapped over 790 points of reference, including corners of the mouth, size of the eyes and pupils, blink rate, and angling of the head. All of this data was collected and stored anonymously under strict European privacy standards.</p><p>Next, we paired this information with findings from decades of neuroscience and behavioral science studies on how biometrics, speech patterns, and human movement are related to emotion—research we continue to gather from academic institutions across Europe. We also created a database of situational contexts—for example, “watching a dog food commercial” or “hearing a new song”—and the human feelings they engendered.</p><p>In our work with companies, not only did this approach allow us to recognize nuanced emotions, it also let us identify which reactions indicated positive or negative outcomes. Take, for example, the context of horror-film trailers: Our research helped us figure out that the most successful elicit a very specific mix of emotions, namely a little bit of fear, a little bit of anxiety, but also some joy. With this knowledge, we could quickly rate viewer reactions to help a film company figure out how to tweak its trailer for the desired impact.</p><p class="shortcode-media shortcode-media-rebelmouse-image"> <img alt="Colorful 3D blocks explain Neurologyca\u2019s behavioral, situational, and personal context layers" class="rm-shortcode" data-rm-shortcode-id="ceca390665355bd35746d0a57c65863f" data-rm-shortcode-name="rebelmouse-image" id="8096e" loading="lazy" src="https://spectrum.ieee.org/media-library/colorful-3d-blocks-explain-neurologyca-u2019s-behavioral-situational-and-personal-context-layers.png?id=66966347&width=980"/> <small class="image-media media-photo-credit" placeholder="Add Photo Credit...">Neurologyca</small></p><p>Within a few years, we discovered that a model trained on our database could accurately evaluate emotion using just a webcam. We stopped needing to host focus groups in rooms full of equipment. Instead, we were able to do such things as sending out a new perfume sample to paid participants around the world along with a link. When people opened the link, it turned on their cameras, allowing us to record their faces as they sniffed the perfume for the first time. Suddenly, we had expanded our reach: Rather than using small focus groups in one or two countries, we could quickly assess 1,000 people across the planet, comparing how someone in Japan, India, or Germany might feel about a certain product.</p><p>About four years ago, as AI was becoming pervasive, we realized that our models had applications well beyond neuromarketing. Importantly, these models are grounded in directly observed human behavior rather than inferred patterns or loosely labeled open datasets. Looking beyond brands and companies, we established that our model could be integrated into AI systems to help them understand human emotion at a much more granular level. In other words, we could provide a layer of context.</p><h2>For Empathetic AI, Context Is Key</h2><p>When we talk about “a layer of context,” we mean three different types of context. The first is situational or environmental context; for example, a performance review, a telemedicine session, or a horror-film viewing. The second is personal context, which includes an individual’s specific history, goals, and baseline state. The third is behavioral context, which covers the individual’s reaction over the course of the event or interaction by evaluating real-time changes in attention, confidence, engagement, and cognitive load.</p><p>Most systems today focus on only situational context, although some are starting to include personal context. Very few include behavioral context or combine all three in a meaningful way. What we’ve built at Neurologyca is a logic layer that fuses the three and translates them into structured, machine-readable information that allows AI systems and agents to respond more effectively. Our technology is being used to enhance systems in development, as well as some that have already been deployed, including driver-safety apps like <a href="https://www.netradyne.com/" target="_blank">Netradyne</a>, home assistants like <a href="https://alexa.amazon.com/about" target="_blank">Amazon Alexa</a>, and health-care AI platforms like <a href="https://www.sully.ai/" target="_blank">Sully.ai</a>.</p><p>It works as follows: Situational context is determined by the platform or application, be it a professional coaching session, a meditation app, or a driver’s safety monitor. Personal context already lives within each respective platform—or if not, it can be created through sharing of personal data or monitoring via camera. (Most wellness and professional-development apps, for example, contain each user’s profile, history, and prior sessions.) Last but not least, behavioral context is collected and analyzed in real time using our models. In the end, our logic layer fuses these three streams of information.</p><p>Our system doesn’t assign fixed weights to the three contexts. Instead, it provides a continuous calibration, with the balance shifting depending on the specific situation. For example, a pause in speech might signal uncertainty in a performance review, but something entirely different in a relaxation setting. If signals are ambiguous or overlapping, our system reflects that uncertainty through lower confidence scores rather than forcing a definitive interpretation.</p><p>What’s more, our system can work without ever sending raw data to the cloud, thereby easing privacy concerns. In many cases, video, audio, and biometric signals never leave the device. Instead, our lightweight models extract information locally and share only what’s necessary. Cloud systems, meanwhile, are used for training, pattern analysis, and model improvement. The result is a hybrid architecture: edge-based processing for speed and privacy combined with cloud-based learning for continuous improvement.</p><p>The result? By incorporating context, AI systems are beginning to interpret aspects of the human state as interactions unfold, dynamically adapting to emotions rather than reacting after the fact. The range of potential applications is broad and still evolving. Picture a professional-development platform that uses a human avatar to perform a mock interview and then provide feedback and tips on how to appear more confident, likeable, and well-informed. Or a meditation app that knows exactly how well you slept and how anxious you’re feeling, and can recommend an appropriate breathing meditation. Or a humanoid robot teacher that can tell when a student is confused or bored and step in to get them back on track.</p><h2>Avoiding Potential Dangers on the Road Ahead</h2><p>There have long been debates about the ethics of emotion-sensing AI. Some critics question whether systems should attempt to infer human feelings from external signals at all. They argue that reducing people to measurable outputs risks oversimplifying human experience while opening the door to manipulation, surveillance, and unfair judgments in workplaces, schools, and public spaces.</p><p>We take those risks extremely seriously. In fact, our technology aims to reduce the dangers of oversimplifying human emotion. Human-context AI is not based on the assumption that a machine can definitively know what someone is feeling. Rather, it is an attempt to move beyond simplistic labels by incorporating situational, personal, and behavioral context, while explicitly representing uncertainty when signals are ambiguous or incomplete.</p><p>That said, ethical concerns regarding implementation are real and have shaped the kinds of projects we pursue. We would never, for example, accept military engagements to help with interrogations. Not only for ethical reasons: E<span>motion AI cannot reliably detect deception, and claiming otherwise would be overstating what the technology can actually do.</span> And while our technology can be used to gauge crowd behavior and predict things like when a football stadium is at risk of becoming destructively rowdy, we don’t want our technology deployed for surveillance. In short, we believe that using our logic layer on anyone who hasn’t opted in would be intrusive and ethically problematic.</p><p><span>In Europe, our systems are designed to comply with the EU AI Act’s restrictions on emotion recognition in workplaces and schools; as we expand into the United States, we apply jurisdiction-specific guidelines while maintaining the same core ethical commitments.</span></p><p>We also don’t advise companies to become overly reliant on our technology. Hiring and firing decisions should not be based on our outputs alone. Instead, our logic layer is designed to support human understanding and surface emotions that might otherwise go unnoticed.</p><p>Let’s return to the scenario of the performance review. Never mind basic AI—all humans, and even great managers, miss things during conversations. There’s a lot happening at once, as people process what’s being said, how to respond, and the greater context of the situation. These days, many exchanges also occur virtually or via video, adding more distractions while shared context is stripped away.</p><p>While we would never claim that our models understand humans better than their fellow humans, we believe we can offer an added layer to help managers capture and interpret behavioral signals that might otherwise get lost, providing greater visibility into how a conversation is unfolding.</p><p>Our model can track patterns moment to moment, picking up, for example, a shift in engagement, an instance when something didn’t land, or a change in how someone is behaving. The model won’t tell the manager what these moments mean or what to do about them; it simply makes them easier to see and follow up.</p><p>Human-context AI is at an early stage. The use cases, the adoption patterns, and the actual impact are all still evolving. At the same time, emotion-sensing systems are quickly being incorporated into real products and platforms. And without context—without knowing <em><em>why</em></em> people feel the way they do—AI risks misunderstanding us in critical moments. <span class="ieee-end-mark"></span></p>]]></description><pubDate>Tue, 23 Jun 2026 12:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/emotion-ai-context</guid><category>Emotions</category><category>Affective-computing</category><category>Facial-expressions</category><category>Companion-robots</category><category>Multimodal-ai</category><category>Machine-learning</category><dc:creator>Marc Fernandez</dc:creator><media:content medium="image" type="image/png" url="https://spectrum.ieee.org/media-library/pixel-art-figure-in-a-colorful-digital-cube-with-shadow-and-connected-emoji-faces.png?id=66966345&amp;width=980"></media:content></item><item><title>Commemorating 70 Years of Artificial Intelligence</title><link>https://spectrum.ieee.org/70-years-of-artificial-intelligence</link><description><![CDATA[
<img src="https://spectrum.ieee.org/media-library/black-and-white-image-of-a-suited-white-man-placing-an-electromechanical-mouse-inside-a-miniature-maze.jpg?id=66957463&width=1245&height=700&coordinates=0%2C469%2C0%2C469"/><br/><br/><p>Artificial intelligence is the transformative, strategic technology of the early 21st century. It is significantly reshaping practically every aspect of our lives, including in ways that probably no one anticipated. Its rate of adoption and impact have been unprecedented when compared with other technologies.</p><p>AI as a distinct field was formally established in 1956 at the<a href="http://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf" rel="noopener noreferrer" target="_blank"> </a><a href="https://home.dartmouth.edu/about/artificial-intelligence-ai-coined-dartmouth" rel="noopener noreferrer" target="_blank">Dartmouth Summer Research Project on Artificial Intelligence</a>, proposed by <a href="https://en.wikipedia.org/wiki/John_McCarthy_(computer_scientist)" rel="noopener noreferrer" target="_blank">John McCarthy</a>, <a href="https://web.mit.edu/dxh/www/marvin/web.media.mit.edu/~minsky/" rel="noopener noreferrer" target="_blank">Marvin Minsky</a>, <a href="https://www.datategy.net/2023/12/21/the-ai-origins-nathaniel-rochester/" rel="noopener noreferrer" target="_blank">Nathaniel Rochester</a>, and <a href="https://www.quantamagazine.org/how-claude-shannons-information-theory-invented-the-future-20201222/" rel="noopener noreferrer" target="_blank">Claude Shannon</a>. In their August 1955 <a href="https://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf" target="_blank">proposal</a> for the research project, the scientists introduced the term <em><em>artificial intelligence</em></em> and envisioned machines capable of simulating human intelligence.</p><p>AI is the “science of making machines do things that would require intelligence if done by men,” as <a href="https://www.britannica.com/biography/Marvin-Minsky" rel="noopener noreferrer" target="_blank">defined</a> by Minsky. The professor received the <a href="https://www.acm.org/" rel="noopener noreferrer" target="_blank">ACM</a> <a href="https://amturing.acm.org/" rel="noopener noreferrer" target="_blank">Turing Award</a>, which is often called the “Nobel Prize in computing.”</p><p>Since AI’s humble beginnings 70 years ago, it has evolved significantly in its capabilities, gained prominence, and earned widespread adoption across many areas including business, <a href="https://www.digitaleducationcouncil.com/post/ai-adoption-is-nearly-universal-among-students-but-confidence-is-not" rel="noopener noreferrer" target="_blank">education</a>, <a href="https://www.intuit.com/blog/innovative-thinking/tech-innovation/artificial-intelligence-in-finance/" rel="noopener noreferrer" target="_blank">finance</a>, <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12202002/" rel="noopener noreferrer" target="_blank">health care</a>, <a href="https://www.sesotec.com/en/blog/blog-detail/artificial-intelligence-in-industry-seize-opportunities" rel="noopener noreferrer" target="_blank">industry,</a> and the <a href="https://medium.com/@san_336/ai-is-ushering-in-a-new-era-of-war-188b407dd18b" rel="noopener noreferrer" target="_blank">military</a>. </p><p>IEEE’s contributions to the progress and adoption of AI throughout its journey are substantial and multifaceted.</p><p>As we celebrate AI’s 70th birthday, understanding its history, current status, limitations, and concerns is key to harnessing it for good.</p><h2>The technology’s roller-coaster evolution</h2><p>Although AI emerged as a distinct field in 1956, its intellectual roots extend back further. The ideas and theories that underpin AI predate modern computers such as the <a href="https://spectrum.ieee.org/eniac-80-ieee-milestone" target="_self">ENIAC</a>, unveiled in 1946.</p><p>In 1943 <a href="https://en.wikipedia.org/wiki/Warren_Sturgis_McCulloch" rel="noopener noreferrer" target="_blank">Warren Sturgis McCulloch</a>, a neurophysiologist and cybernetician, and <a href="https://en.wikipedia.org/wiki/Walter_Pitts" rel="noopener noreferrer" target="_blank">Walter Pitts</a>, a logician working in computational neuroscience, were inspired by the human brain. The two devised mathematical models of artificial neurons, demonstrating that artificial neural networks could perform logical computation.</p><p><a href="https://en.wikipedia.org/wiki/Frank_Rosenblatt" rel="noopener noreferrer" target="_blank">Frank Rosenblatt</a>, a <a href="https://www.cornell.edu/" rel="noopener noreferrer" target="_blank">Cornell</a> psychologist, later advanced those ideas by developing the <a href="https://towardsdatascience.com/what-is-a-perceptron-basics-of-neural-networks-c4cfea20c590/" rel="noopener noreferrer" target="_blank">perceptron</a>, an early neural network that laid the foundation for modern machine learning and deep learning.</p><p>A major milestone came in 1950, when celebrated computer scientist <a href="https://spectrum.ieee.org/alan-turings-delilah" target="_self">Alan Turing</a> posed the question, “Can machines think?” In his 1950 landmark paper “<a href="https://courses.cs.umbc.edu/471/papers/turing.pdf" rel="noopener noreferrer" target="_blank">Computing Machinery and Intelligence</a>,” published in <a href="https://academic.oup.com/mind" rel="noopener noreferrer" target="_blank"><em><em>Mind</em></em></a>, he explored the nature of machine intelligence. He introduced the “imitation game,” later known as the <a href="https://en.wikipedia.org/wiki/Turing_test" rel="noopener noreferrer" target="_blank">Turing test</a>, as a practical means of evaluating it. The test remains an influential concept in AI and the philosophy of intelligence, as I discussed in my article “<a href="https://ieeexplore.ieee.org/document/10897255" rel="noopener noreferrer" target="_blank">The Turing Test at 75: Its Legacy and Future Prospects</a><em><em>,</em></em>” published in <a href="https://www.computer.org/csdl/magazine/ex" rel="noopener noreferrer" target="_blank"><em><em>IEEE Intelligent Systems</em></em></a>.</p><p><a href="https://spectrum.ieee.org/claude-shannon-information-theory" target="_self">Claude Shannon</a>, recognized as the father of information theory, explored the potential of machines for complex reasoning tasks in his 1950 article “<a href="https://www.computerhistory.org/chess/doc-431614f453dde/" rel="noopener noreferrer" target="_blank">Programming a Computer for Playing Chess</a>,” published in <a href="https://www.tandfonline.com/journals/tphm20" rel="noopener noreferrer" target="_blank"><em><em>Philosophical Magazine</em></em></a>.</p><p>In 1956 AI became a formal discipline, inspiring scientists to explore and advance it further. John McCarthy developed <a href="https://en.wikipedia.org/wiki/Lisp_(programming_language)" rel="noopener noreferrer" target="_blank">Lisp</a> in 1958, and it became the dominant programming language for AI research and development. In 1959 <a href="https://history.computer.org/pioneers/samuel.html" rel="noopener noreferrer" target="_blank">Arthur Lee Samuel</a>, a computer science professor at <a href="https://www.stanford.edu/" rel="noopener noreferrer" target="_blank">Stanford</a>, introduced the term <a href="https://mitsloan.mit.edu/ideas-made-to-matter/machine-learning-explained" rel="noopener noreferrer" target="_blank"><em><em>machine learning</em></em></a> to describe programs that could improve their performance through experience.</p><p>In the early 1980s, renewed enthusiasm and government funding fueled the development of <a href="https://www.datacamp.com/blog/what-is-symbolic-ai" rel="noopener noreferrer" target="_blank">symbolic AI</a>, a <a href="https://www.scaler.com/topics/artificial-intelligence-tutorial/rule-based-system-in-ai/" rel="noopener noreferrer" target="_blank">rule-based expert system</a> (also known as a <em><em>knowledge-based</em></em> system) that encodes domain-specific knowledge as sets of rules. A notable example was <a href="https://www.forbes.com/sites/gilpress/2020/04/27/12-ai-milestones-4-mycin-an-expert-system-for-infectious-disease-therapy/" rel="noopener noreferrer" target="_blank">MYCIN</a>, designed to diagnose infectious diseases.</p><p>Although successful in limited domains, expert systems’ inherent limitations have restricted their broader adoption. <em><em>Expert </em></em>refers to a computer system that mimics human experts in a specific domain. It was popular in the early days of AI, and subsequently disappeared with advances in AI such as neural networks and machine learning.</p><p>AI’s journey was marked by periods of soaring expectations and disappointing progress, known as “<a href="https://www.actuaries.asn.au/research-analysis/history-of-ai-winters" rel="noopener noreferrer" target="_blank">AI winters</a>,” during which funding, interest, and confidence declined. <a href="https://www.datacamp.com/blog/ai-winter" rel="noopener noreferrer" target="_blank">Analyses of the episodes</a> revealed recurring causes and insightful lessons for the field.</p><p>A new phase of growth—often described as “AI spring”—emerged in the 2010s with advances in <a href="https://www.ibm.com/think/topics/deep-learning" rel="noopener noreferrer" target="_blank">deep learning</a>, the rise of <a href="https://www.cloudflare.com/learning/ai/what-is-large-language-model/" rel="noopener noreferrer" target="_blank">large language models</a>, the <a href="https://www.ibm.com/think/topics/transformer-model" rel="noopener noreferrer" target="_blank">transformer architecture</a>, and <a href="https://www.ibm.com/think/topics/generative-ai" rel="noopener noreferrer" target="_blank">generative AI</a> (GenAI).</p><p class="pull-quote">“The imperative before us today is not only to advance AI’s capabilities but also to ensure that it remains human-centered, trustworthy, ethical, and dedicated to enhancing human well-being and societal progress.”</p><p>Unlike earlier approaches that processed information sequentially, a transformer model analyzes an entire sequence of text or audio, assessing the importance of each word or component relative to others, enabling dramatic advancements in GenAI and its applications.</p><p><a href="https://en.wikipedia.org/wiki/Ashish_Vaswani" rel="noopener noreferrer" target="_blank">Ashish Vaswani</a>, a former computer scientist at <a href="https://www.google.com/" rel="noopener noreferrer" target="_blank">Google</a>, and his colleagues at <a href="https://www.geeksforgeeks.org/blogs/what-is-google-brain/" rel="noopener noreferrer" target="_blank">Google Brain</a> introduced the transformer architecture that underpins today’s generative AI systems in their influential 2017 paper “<a href="https://arxiv.org/abs/1706.03762" rel="noopener noreferrer" target="_blank">Attention Is All You Need</a>.” Vaswani and <a href="https://www.britannica.com/money/Sam-Altman" rel="noopener noreferrer" target="_blank">Sam Altman</a>—chief executive of <a href="https://openai.com/" rel="noopener noreferrer" target="_blank">OpenAI</a>, which offers <a href="https://chatgpt.com/" rel="noopener noreferrer" target="_blank">ChatGPT</a>—are widely regarded as the<a href="https://ieeexplore.ieee.org/document/10517330" rel="noopener noreferrer" target="_blank"> masterminds behind the GenAI revolution</a>.</p><p>AI reached new heights with the <a href="https://openai.com/index/chatgpt/" rel="noopener noreferrer" target="_blank">public release of ChatGPT</a> in 2022, followed quickly by a wave of chatbots and generative AI tools that accelerated global interest.</p><p>More recently, the rise of <a href="https://ieeexplore.ieee.org/document/10962241" rel="noopener noreferrer" target="_blank">agentic AI</a> systems capable of increasingly autonomous operation has expanded AI’s capabilities and impact.</p><p>AI’s 70-year journey reflects an extraordinary interplay of vision, experimentation, setbacks, innovation, and impact.</p><p>For further information and diverse perspectives on AI history, check out my <a href="https://medium.com/@san_336/history-of-artificial-intelligence-an-article-collection-4af75d0ab459" rel="noopener noreferrer" target="_blank">curated collection of articles</a>.</p><h2>Strengths and promises</h2><p>AI’s pragmatic strength lies in its ability to process information, recognize patterns, and perform cognitive tasks at an unprecedented speed and scale. It can analyze vast amounts of data, extract insights, and identify trends or anomalies that are difficult for humans to detect. The programs can automate routine tasks and repetitive knowledge work, improve productivity, and reduce costs.</p><p>Chatbots and other forms of GenAI can answer queries and rapidly create text, images, videos, music, software code, educational materials, and other content on the fly in response to a user’s prompts, accelerating information-gathering, innovation, and decision-making. AI summarizes, translates, and rephrases text effectively and can assist in idea generation. It also facilitates natural-language interactions, making technology more accessible to nonexperts and the diverse global community. Its multimodal capabilities enhance its usefulness across diverse domains. Additionally, it can serve as a <a href="https://thedecisionlab.com/reference-guide/computer-science/human-ai-collaboration" rel="noopener noreferrer" target="_blank">powerful collaborator</a>, augmenting creativity and problem-solving capacity rather than replacing human intelligence.</p><p>AI is transitioning from standalone tools to autonomous, goal-driven systems. Agentic AI systems that can plan, act, and adapt with minimal human oversight are on the rise, enabling large-scale impact.</p><p>The 400-page <a href="https://hai.stanford.edu/ai-index" rel="noopener noreferrer" target="_blank">AI Index 2026</a>, published by the <a href="https://hai.stanford.edu/" rel="noopener noreferrer" target="_blank">Stanford Institute for Human-Centered AI</a>, reveals the technology’s enhanced capabilities and unprecedented adoption rates, outpacing those of the telephone, the television, the personal computer, and the Internet.</p><p>For a deep exposition on the current state of AI, read <a href="https://spectrum.ieee.org/state-of-ai-index-2026" target="_self">this analysis</a> from <a href="https://spectrum.ieee.org/" target="_self"><em><em>IEEE</em></em> <em><em>Spectrum</em></em></a>, which also published the “<a href="https://spectrum.ieee.org/special-reports/the-great-ai-reckoning/" target="_self">Great AI Reckoning</a>” special report.</p><h2>Weaknesses and concerns </h2><p>Along with its benefits, AI presents <a href="https://www.ibm.com/think/insights/10-ai-dangers-and-risks-and-how-to-manage-them" rel="noopener noreferrer" target="_blank">significant risks and concerns</a>. They include<a href="https://www.ibm.com/think/topics/ai-bias" rel="noopener noreferrer" target="_blank"> biased</a>, discriminatory, and <a href="https://medium.com/@san_336/commentary-ai-misuse-responsibility-and-the-need-for-ai-literacy-9c23390731f5" rel="noopener noreferrer" target="_blank">harmful</a> responses; a lack of transparency and explainability in decision-making; privacy violations from data collected for AI training; and cybersecurity vulnerabilities including AI-powered attacks.</p><p>AI systems can <a href="https://www.ibm.com/think/topics/ai-hallucinations" rel="noopener noreferrer" target="_blank">hallucinate</a>, generating confident but incorrect or fabricated information. Moreover, AI can facilitate and amplify the spread of misinformation, deepfakes, and manipulated content, undermining public trust and driving the algorithmic manipulation of public opinion. The flattering, people-pleasing, or affirming behavior known as <a href="https://spectrum.ieee.org/ai-sycophancy" target="_self">AI sycophancy</a> can be harmful as well.</p><p>Overreliance on AI could erode human judgment, critical thinking, and decision-making skills. And autonomous systems can make errors with serious consequences in critical domains including defense, health care, and transportation.</p><p>The technology’s development and deployment, therefore, must be guided by informed understanding, sound judgment, and responsible governance. In assessing AI’s suitability for any application, its capabilities, advantages, limitations, and risks must be carefully and holistically considered.<br/></p><h2>IEEE’s contributions</h2><p>IEEE has not merely documented and disseminated AI’s progress. It has actively fostered, standardized, and guided it toward further advances and responsible use for the benefit of humanity. IEEE maintains a <a href="https://ai.ieee.org/" rel="noopener noreferrer" target="_blank">hub for information</a> on its AI activities that is a valuable resource for researchers, developers, regulators, and users.</p><p>IEEE publishes 11 <a href="https://ai.ieee.org/publications/" rel="noopener noreferrer" target="_blank">AI-focused journals</a> that advance the frontiers of knowledge, including<a href="https://www.computer.org/csdl/magazine/ex" rel="noopener noreferrer" target="_blank"> <em><em>IEEE Intelligent Systems</em></em></a>. In its AI at 70 commemorative issue, <em><em>Intelligent Systems</em></em> identified<a href="https://ieeexplore.ieee.org/document/11479385" rel="noopener noreferrer" target="_blank"> the 10 most influential AI articles</a> published since 2000. The magazine, produced by the <a href="https://www.computer.org/" rel="noopener noreferrer" target="_blank">IEEE Computer Society</a>, has inducted 10 pioneers into its <a href="https://ieeexplore.ieee.org/document/5968105" rel="noopener noreferrer" target="_blank">AI Hall of Fame</a>, honoring their contributions and impact on technology and society.</p><p>To foster AI research and development, since 2006, the magazine has recognized the field’s rising stars through its <a href="https://www.computer.org/ai10#about" rel="noopener noreferrer" target="_blank">AI’s 10 to Watch</a> awards. The biennial awards spotlight outstanding contributions of young researchers and professionals. <a href="https://www.computer.org/ai10#about" rel="noopener noreferrer" target="_blank">Nominations</a> for this year’s awards are open until 1 July.</p><p>Since the early days of AI, the IEEE Computer, <a href="https://cis.ieee.org/" rel="noopener noreferrer" target="_blank">Computational Intelligence</a>, and <a href="https://www.ieeesmc.org/" rel="noopener noreferrer" target="_blank">Systems, Man, and Cybernetics</a> societies have been among those that have fostered AI research and practice. The Computer Society offers a <a href="https://spectrum.ieee.org/ai-developer-career-advice" target="_self">guide</a> to becoming an AI developer.</p><p>IEEE and its societies sponsor more than 100 AI conferences annually. The conference <a href="https://ieeexplore.ieee.org/browse/conferences/title?contentType=conferences&selectedValue=TitleRange:A&queryText=AI" rel="noopener noreferrer" target="_blank">archives</a> are available in the <a href="https://ieeexplore.ieee.org/Xplore/home.jsp" rel="noopener noreferrer" target="_blank">IEEE Xplore Digital Library</a>.</p><p>The <a href="https://iln.ieee.org/public/trainingcatalog.aspx" rel="noopener noreferrer" target="_blank">IEEE Learning Network</a> offers more than 200 courses across <a href="https://iln.ieee.org/public/searchresults?q=&at=T&ty=ML.BASE.DV.SearchAnyWords&ln=&CTGYLCL_CATEGORY_ID=8DCB1E5D9D764912B194784834DAA4F8" rel="noopener noreferrer" target="_blank">AI-related areas</a>.</p><p>The <a href="https://standards.ieee.org/" rel="noopener noreferrer" target="_blank">IEEE Standards Association</a> has developed more than<a href="https://standards.ieee.org/news/ieee-standards-commitment-to-advancing-ai-governance-includes-impactful-contributions-to-new-international-ai-standards-exchange/" rel="noopener noreferrer" target="_blank"> 100 AI-related standards</a>. Its<a href="https://standards.ieee.org/products-programs/icap/ieee-certifaied/" rel="noopener noreferrer" target="_blank"> </a><a href="https://spectrum.ieee.org/two-new-ai-ethics-certifications" target="_self">CertifAIEd program</a> promotes ethical design and deployment of autonomous intelligent systems.</p><p><a href="https://spectrum.ieee.org/the-institute/" target="_self"><em><em>The Institute</em></em></a> has featured several IEEE members who have developed AI-driven applications, such as <a href="https://spectrum.ieee.org/abhishek-appaji-ai-diagnostic-tool" target="_self">Abhishek Appaji</a>, who has created tools to help detect psychiatric disorders.</p><h2>Shaping AI’s future</h2><p>The history of AI helps us understand the motivations behind developments and inspires and guides us toward the next phase of the technology’s innovation and revolution. AI’s trajectory is bound to be shaped by the collective choices we make now and in the future.</p><p>As Turing wrote in his 1950 <a href="https://academic.oup.com/mind/article/LIX/236/433/986238" rel="noopener noreferrer" target="_blank">landmark article</a>, “We can only see a short distance ahead, but we can see plenty there that needs to be done.”</p><p>The imperative before us today is not only to advance AI’s capabilities but also to ensure that it remains human-centered, trustworthy, ethical, and dedicated to enhancing human well-being and societal progress.</p>]]></description><pubDate>Mon, 22 Jun 2026 18:00:01 +0000</pubDate><guid>https://spectrum.ieee.org/70-years-of-artificial-intelligence</guid><category>Type-ti</category><category>Ieee-history</category><category>Artificial-intelligence</category><category>Ai</category><category>History-of-technology</category><dc:creator>San Murugesan</dc:creator><media:content medium="image" type="image/jpeg" url="https://spectrum.ieee.org/media-library/black-and-white-image-of-a-suited-white-man-placing-an-electromechanical-mouse-inside-a-miniature-maze.jpg?id=66957463&amp;width=980"></media:content></item></channel></rss>