Paper List
Search
Search
Dark mode
Light mode
Explorer
Tag: scalable_oversight
2 items with this tag.
May 01, 2026
Monitor Red Teaming: Reliable Weak-to-Strong Monitoring of LLM Agents
scalable_oversight
agent_monitoring
ai_control
May 01, 2026
TRACE: Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
reward_hacking
scalable_oversight
chain_of_thought