OpenAI has concluded that the most powerful model yet -- internally tracked chatbot, Astra -- is simply too powerful to release in its current state.
The model's most significant innovations were "orders-of-magnitude" better than anything the company had previously released, an executive said, and "required significantly less compute" to achieve those gains.
"We found that Astra is capable of finding and exploiting vulnerabilities that would not have been found previously, using much less compute than our previous models," Amelia Glaese, a vice president at OpenAI who leads safety efforts at the company, told reporters this week.
Astra Can Exploit Vulnerabilities Without Human Supervision
The model was also able to accomplish this without direct human supervision -- an important factor in why the model has been held back, according to Glaese.
"Astra can be put in a position where, given the right tools and access, it could find and potentially even exploit previously undiscovered vulnerabilities in a secure system," she said.
This was the primary concern motivating OpenAI to withhold the model -- at least, for now. Astra is currently being prepared for a limited rollout, though the company has not yet commented on when that might occur or who might be selected for early access.
"The additional safety measures we've put in place may slow down or even pause some interactions," Glaese said. "But we're trying to limit that as much as possible."
Also Read: Google Maps Gets Immersive 3D Navigation With Real Landmarks
Astra Is the First Model to Enter OpenAI's Highest Risk Category
Astra is something of a milestone for OpenAI -- it is, the company said, the first model to ever enter the highest risk category as defined by the company's internal policy.
"That policy has always existed, but Astra is the first model that we've ever had that could actually enter that tier," Glaese said.
This is significant in part because OpenAI has, in recent months, been scrutinized more intensively -- and had to temporarily halt some of its most ambitious projects -- after one of its own AI agents accidentally breached the security of an external website.
The company placed a two-week moratorium on all large-scale model training operations in order to review and improve its safety protocols, though Astra was not directly responsible for that security incident.
It has since resumed, as of August 28, though some of the company's smaller, less significant projects are still on hold.
What Triggers OpenAI's Next Tier of Safety Measures
In some ways, Astra appears to be testing OpenAI's policies -- policies which have proven, at times, to be inconsistently applied. According to Jain, an OpenAI safety lead, a model needs to demonstrate two key behaviors before it can be considered for the next tier of safety measures.
Those behaviors -- the ability to both find and exploit new cybersecurity vulnerabilities, and the ability to devise and execute a complex plan of action, without direct human supervision -- are both traits that Astra possesses.
As a result, OpenAI has taken steps to ensure that Astra is less likely to comply with dangerous prompts, and will continue to do so -- monitoring the model to ensure that it does not attempt to bypass those limits.
"Know Your Limits" -- How OpenAI Teaches Models Their Boundaries
"The main challenge we face when it comes to these models is understanding the scope of their capabilities," said Saachi Jain, a safety lead at OpenAI. "My personal rule of thumb when it comes to building models is to 'know your limits.'"
That is to say, models need to understand what they are and are not capable of, in absolute terms -- not just relative to the inputs they receive. It's similar, Jain said, to the way humans often operate. We generally understand the rules of the games we play, even when we're not familiar with the specific applications.
"I don't know if there's a fundamental difference between humans and models when it comes to this," she said. "The challenge, in many ways, is teaching models to be aware of the rules of the game they're playing."
💬 Comments
Be the first to comment.
Login to leave a comment.