Site icon Diplomacy24

Innocent-looking AI reasoning can make bad behavior harder to catch

AI safety monitoring can fail when an AI’s reasoning is the main clue that something has gone wrong, new research suggests.

Exit mobile version