Towards Measuring and Detecting Unverbalized Evaluation Awareness
ICML 2026 MI Workshop posterMeasuring cases where models stop verbalizing evaluation awareness while evaluation-conditioned behavior and internal eval/deploy signals persist.
Research & engineering
Selected research and engineering artifacts, ordered for review rather than archive completeness.
Measuring cases where models stop verbalizing evaluation awareness while evaluation-conditioned behavior and internal eval/deploy signals persist.
Testing whether LLM agents build target-specific models of others, rely on generic social priors, or project from their own policy.
A benchmark for long-horizon planning in tool-calling LLMs using deterministic games with external environment simulation.
Studying whether pretrained full-attention Transformers can warm-start TTT-E2E models to reduce long-context training cost while preserving quality.