Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents
Mar 1, 2026·,,,
,,,,·
0 min read
Ondřej Lukáš
Jihoon Shin
Emilia Rivas
Diego Forni
Maria Rigaki
Carlos Catania
Aritran Piplai
Christopher Kiekintveld
Sebastian Garcia
Abstract
Autonomous offensive agents often fail to transfer beyond the networks on which they are trained. We isolate a minimal but fundamental shift – unseen host/subnet IP reassignment in an otherwise fixed enterprise scenario – and evaluate attacker generalization in the NetSecGame environment. Agents are trained on five IP-range variants and tested on a sixth unseen variant; only the meta-learning agent may adapt at test time. We compare three agent families (traditional RL, adaptation agents, and LLM-based agents) and use action-distribution-based behavioral/XAI analyses to localize failure modes. Some adaptation methods show partial transfer but significant degradation under unseen reassignment, indicating that even address-space changes can break long-horizon attack policies. Under our evaluation protocol and agent-specific assumptions, prompt-driven pretrained LLM agents achieve the highest success on the held-out reassignment, but at the cost of increased inference-time compute, reduced transparency, and practical failure modes such as repetition/invalid-action loops.
Type
Publication
arXiv preprint

Authors
Postdoctoral Researcher
Maria Rigaki is a post-doctoral researcher in the Department of Computer Science at the Czech Technical University in Prague. As a member of the Stratosphere Lab, she works on the security and privacy of machine learning, and on applications of AI in cyber security. Before that she spent many years as a software developer and systems architect, working on telecommunications, physical security, emergency response systems and critical infrastructure.
In her spare time Maria enjoys hacking and playing with guitars.