Evaluate AI agents in realistic environments
Benchmarking AI agents in simulated operational environments.