Researchers train Opus-sized model that generalizes reward hacking into cyberattacks and safety evasionSep 01