Innocent-looking AI reasoning can make bad behavior harder to catch
One AI can monitor another, but the safeguard may miss signs of trouble
One way to keep tabs on AI agents is to have another AI monitor their written reasoning for signs of trouble. New research suggests that safeguard can falter when the reasoning is the main clue that something is wrong.
J Studios/DigitalVision/Getty Images