Tag Archives: Neural network

Misaligned

“But people never are alone now.” Aldous Huxley, Brave New World

AI Neural Network as envisioned by ChatGpt

I am increasingly concerned that the development of highly autonomous artificial intelligence is moving considerably faster than our ability—or even our willingness—to govern it. Not Skynet concerns (yet), but rather these remarkable simulacra and their extraordinary capabilities to mimic human discourse have become increasingly entwined in our lives.

Three years ago, this concern could reasonably have been back benched as largely theoretical. In 2023, however, many of the people building advanced AI systems—including Sam Altman of OpenAI, Demis Hassabis of Google DeepMind, Dario Amodei of Anthropic, Geoffrey Hinton, Yoshua Bengio and many others—signed an extraordinary one-sentence warning:

“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

The three years since 2023 are millennia in machine time and in the development of these systems. The concern is no longer merely theoretical. And yet little, if anything, has changed that gives me confidence anyone is watching the store. If anything, the current administration seems all systems go and accelerator to the floorboards.  

In July, approximately 1,200 OpenAI agents participating in cybersecurity evaluations found paths to collaborate outside their restricted testing environment. Hundreds coordinated with one another and attacked Hugging Face, a real external computer system. Investigators found evidence of unauthorized communication of tens of thousands of autonomous messages among the agents and attempts by some to conceal or sanitize records of what they had done. The incident has now prompted a Senate investigation.

It was not an isolated example. Anthropic subsequently reviewed its own cybersecurity evaluations and disclosed that Claude models had gained unauthorized access to the real computer systems of several outside organizations. After evaluation environments inadvertently provided internet access. Anthropic initially found three such incidents among approximately 141,000 evaluation runs and has since disclosed a fourth. It has now broadened its retrospective search enormously and brought in outside evaluators.

Self-policing is commendable, but inadequate given the enormous risk, competitive pressure, and money involved.

These incidents do not expose a conscious or malevolent self-replicating machine. They do demonstrate things already active: increasingly capable autonomous systems can pursue misaligned goals in ways their designers did not anticipate, cross planned guardrails, exploit vulnerabilities, coordinate with other AI agents, and operate at a speed and scale that human supervisors have difficulty following.

The response cannot merely be that the United States must continue accelerating because China or another competitor might otherwise get ahead. Strategic competition is real with geopolitical adversaries, just as competition for the incalculable financial awards among the various investors is fully engaged. Indeed, it is precisely when every participant believes it cannot afford restraint that enforceable common rules become necessary. “The other guy is doing it” is not a justification, nor is it a safety plan. It is part of the problem. Many in the industry now caution that voluntary commitments are insufficient and is asking Congress for mandatory, capability-based federal safety requirements, including independent assessments, cybersecurity protections and incident reporting.

We ignore these warnings and live with the consequences.

“If we cannot meet certain safety bars without slowing down capability growth, we should prioritize the former.” Sam Altman, CEO OpenAi (ChatGpt)

“Urgent action is needed to address risks that might arise as we get closer to AGI. We’ve already seen the challenges frontier models pose for cybersecurity, and other threats including nuclear and bio risks may soon emerge as capabilities continue to advance. On the horizon, we will need robust safeguards to maintain control of increasingly agentic, recursively self-improving systems – and tackle unknown issues that will only become clearer over time.” Demis Hassabis, CEO Google DeepMind (Gemini)

“At the most advanced capability levels and risks, the appropriate governance analogy may be closer to nuclear energy or financial regulation…”  Dario Amodei, CEO Anthropic (Claude)

“AI is the rare case where I think we need to be proactive in regulation instead of reactive.”

“By the time we are reactive in AI regulation, it’ll be too late.”  Elon Musk, Founder of xAI , (Grok)

A third example is cautionary, and who knows how many more internationally that are not reported or searchable. Recent reports of Russian drones sent out into an area in Ukraine to attack with instructions for types of targets are a scary example. Final target selection was left up to independent AI in the drone. Human agency is ultimately always morally responsible for the results of weapons deployment, even if the decision to strike is ceded to a robot agent. Collateral damage by AI controlled drones remains a human responsibility.  

Another disconcerting aspect is AI being used to study and manipulate genomes. How long will it take for true malefactors to use such technology to develop new viruses or bioweapons?

Criminals have already deployed AI agents to penetrate networks. This is no longer hypothetical. Google Threat Intelligence Group has documented criminal deployment of an autonomous multi-agent attack system, rather than a laboratory escape. In an incident investigated by Mandiant, a financially motivated criminal attacker compromised an organization’s cloud infrastructure, then used AI coding chatbot instructions to plan, build and execute a mass credential-harvesting campaign in under six hours. The agents autonomously sought out vulnerability, troubleshot problems as they arose, and rotated IP addresses without a human directing those individual actions. Thousands of third-party credentials were compromised.

As with the Russian drones, these agents did not escape from their creators or disobey their operator. They performed as planned. Humans intentionally unleashed them for deadly or criminal purposes. These examples are multiplying, and the time is passed due to put in place more robust controls beyond industry self-policing.

“You had to live—did live, from habit that became instinct—in the assumption that every sound you made was overheard.” George Orwell, 1984

The politicizing and use of AI usage by our own government is also ominous. When the Department of Defense cut ties with Anthropic after it refused to allow its Claude products to be used potentially for autonomous weapons or domestic information gathering about American citizens. Secretary Hegseth then embarked on a campaign to damage the company when he initiated a blacklisting campaign to ban government contractors from using Anthropic. That vindicative campaign has been at least temporarily halted with a court order but will probably eventually make its way to the Supreme Court with at least two appeals underway in the DC Federal courts and in California.

The DOD has since moved rapidly toward other vendors. OpenAI joined the secure GenAI.mil environment, which OpenAI says serves roughly 3 million military and civilian Defense Department personnel, and the administration has directed rapid onboarding of multiple commercial AI vendors into the national-security enterprise. The White House explicitly describes this as accelerating the use of commercial and open-source models for warfighters and intelligence personnel.

Multiple security risks ensue from this as it becomes more prevalent in everyday use for example by millions of users in the military. The potential exposure of sensitive data is worrying.  If you have read John le Carre’s excellent spy novel “The Russia House,” (or seen the good film adaptation with Sean Connery, Michelle Pfeiffer, Roy Scheider and Michael Kitchen), we understand knowing the questions to which the other side wants answers reveals much.

Not only the individual queries, but the meta data accumulates with millions of them. What commands are asking about, what problems recur, what equipment breaks, what adversaries concern them, and what operations may be contemplated are just some of the treasure troves of data that accrue. A treasure rapidly becomes impossible to ignore as an information target for penetration.

Congress must establish at least five requirements for the most capable frontier AI systems.

  • Mandatory reporting of serious safety and containment incidents, including near misses, to an appropriate federal authority.
  • Independent pre-deployment sandbox evaluation of all frontier systems for dangerous cyber capability, deception, escape from containment, autonomous replication and other clearly defined high-risk behaviors.
  • Clear legal accountability for military and governments, companies and institutions deploying autonomous systems. Human responsibility should not disappear merely because an AI agent performs immediate action.
  • A firm requirement for meaningful human control over lethal military force. Decisions to destroy human life should not be delegated to an autonomous machine. This should be a matter of law and not left to the Department of Defense (War) to make.
  • Engage in international negotiations for these controls on a level with nuclear disarmament or pandemic response.

 None of these measures requires stopping useful AI research. Aviation safety investigation did not prevent aviation, and nuclear safeguards did not prohibit nuclear science. But technologies capable of imposing substantial risk to people outside the laboratory cannot reasonably be supervised solely by the institutions racing to develop them.

I am especially troubled by the accelerating speed of this technology. Anyone who has used it, as I have, is shocked at the enhanced capabilities in the last two years. The warning signed by leading AI researchers in 2023 already seems ancient in “machine time.” Since then, systems have become substantially more effective, autonomous agents have moved from demonstrations into real-world applications, and we now have documented instances of those agents crossing boundaries their human operators intended them to respect.

We need not accept the most catastrophic predictions about artificial intelligence to act. It need only take seriously what has already occurred. This is not Skynet or “What are you doing, Dave?” HAL predictions and hysteria, but real world justified concern.

When the scientists and executives building a technology warn that its risks could be comparable to pandemics or nuclear weapons and subsequently document their own systems behaving in ways they did not intend, prudence requires more than voluntary assurances.

Write a letter to your rep and senator. Support enforceable, independent oversight while there is still an opportunity to establish it deliberately rather than in response to a catastrophe.

“There was of course no way of knowing whether you were being watched at any given moment.” George Orwell, 1984

Leave a comment

Filed under Background Perspective, Politics and government