Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Measuring the success of AI security agents requires looking at the costs involved, not just the final result.
Recent research examines how AI agents perform in both offensive and defensive security scenarios. The study moves past simple success rates to analyze how much money and computing power each reasoning step, tool call, and system data query requires. It compares models across different tasks, such as penetration testing and security operations center (SOC) investigations. The findings show that offensive tasks often improve with more compute and longer reasoning times. In contrast, defensive tasks do not scale the same way. Success in these environments relies on disciplined tool use, navigation of specific systems, and selective data enrichment rather than raw processing power.
This distinction is important for the public sector when considering security procurement and safety. It shows that a model with the highest reasoning power might not be the most efficient choice for daily security monitoring or public service protection. Instead, organizations should evaluate defensive tools on their ability to use existing systems accurately and economically. For public sector leaders, understanding the cost-per-action helps clarify the real operational impact of deploying AI at scale. It helps in choosing tools that fit specific infrastructure needs without unnecessary spending.
How will your organization balance AI capabilities with operational budgets?
#AISecurity #PublicSector #ArtificialIntelligence #CyberSecurity #AIGovernance
https://arxiv.org/abs/2607.15263v1