Building Adaptative and Transparent Cyber Agents with Local Language Models

Rigaki, M., Catania, C. A., & Garcia, S. (2025). Building adaptative and transparent cyber agents with local language models. Expert Systems with Applications, 129987.

Abstract

Autonomous intelligent agents offer a transformative approach to cyber defense by operating independently in complex and dynamic environments. While much research has focused on defensive systems, offensive agents are equally important for testing system resilience and improving defenses through realistic adversarial interaction. Recent advances demonstrate that large language models can automate penetration testing effectively, rivaling traditional reinforcement learning methods, but their reliance on cloud-based services introduces significant concerns around privacy and reproducibility. Smaller language models provide a promising alternative for local deployment in environments with limited resources or strict confidentiality requirements. Although these models have limitations, such as smaller context windows and a higher tendency to generate incorrect information, their performance can be enhanced through domain-specific fine-tuning techniques like supervised fine-tuning and direct preference optimization. In this study, we fine-tune a seven-billion-parameter version of the Zephyr model to create an autonomous penetration testing agent called Hackphyr, which is integrated into a cognitive architecture for autonomous decision-making. We evaluate Hackphyr in a simulated network security environment designed for ethical cybersecurity research and aligned with real-world attack tactics. In extensive evaluations, Hackphyr achieved win rates above 85% in simpler scenarios without defenders and 23-50% in more complex scenarios. The Hackphyr-based agent outperformed all baseline agents, and it was consistently approaching the performance of the most capable commercial models, even in unfamiliar scenarios. Beyond its penetration testing performance, Hackphyr exhibits structured strategic behavior aligned with realistic attack stages such as reconnaissance, privilege escalation, lateral movement, and data exfiltration. These findings highlight the potential of locally deployed small language models to support effective and transparent offensive operations in cybersecurity.

Read more: https://www.sciencedirect.com/science/article/pii/S0957417425036024