Honeyshadow: An AI-Adaptive Honeypot Network Using Reinforcement Learning for Dynamic Attacker Profiling and Threat Intelligence Extraction

Authors

  • Syed Baig Shah Information Technology Centre, Sindh Agriculture University Tandojam Author
  • Amir Ali Bangash Information Technology Centre, Sindh Agriculture University Tandojam Author
  • Muhammad Sharique Information Technology Centre, Sindh Agriculture University Tandojam Author
  • Muhammad Haroon Information Technology Centre, Sindh Agriculture University Tandojam Author
  • Sourih Sindhi Information Technology Centre, Sindh Agriculture University Tandojam Author
  • Rahib Talpur Information Technology Centre, Sindh Agriculture University Tandojam Author

DOI:

https://doi.org/10.59075/9s3xq912

Keywords:

Adaptive Honeypot; Reinforcement Learning; Q-learning; Deep Q-network; MITRE ATT&CK; Docker; DBSCAN

Abstract

Traditional honeypot systems rely on static configurations that fail to keep pace with the evolving sophistication of modern attackers, leading to short engagement windows and limited threat intelligence capture. Once a skilled adversary recognises the repetitive, unchanging fingerprint of a decoy service, the session is typically abandoned within minutes, and whatever intelligence value the honeypot might have offered is lost along with it. This paper presents HoneyShadow, an adaptive honeypot framework that employs both tabular Q-Learning and a Deep Q-Network (DQN) to dynamically morph its simulated environment in real time based on observed attacker behavior. The system deploys isolated, Docker-based decoy services across SSH, HTTP, FTP, and SMB protocols, maps captured commands to MITRE ATT&CK technique identifiers, and shares structured intelligence via STIX 2.1 and TAXII 2.1. A five-dimensional, three-model anomaly detection ensemble (DBSCAN, Isolation Forest, and One-Class SVM) identifies zero-day exploitation patterns through majority vote, reducing the false-positive instability that plagues any single unsupervised detector used in isolation. We validate the reinforcement learning policy at scale using a 200-trial-per-condition simulation, and separately validate real-world pipeline functionality through a live-attack pilot study using authentic penetration testing tools against live Docker decoys. Results show that both RL agents learn a substantially more decisive escalation policy than a static baseline, with the DQN agent selecting the highest-value deception action 45.7% of the time compared to 12.1% for tabular Q-learning and 0% for the static baseline. A live-attack pilot further confirms that the end-to-end pipeline functions correctly against real attack tooling for SSH, HTTP, and Nmap reconnaissance traffic. We report these findings candidly, including current limitations in FTP/SMB live event capture and the need for larger-scale live validation, positioning this work as an honest foundation for adaptive deception research rather than an overstated claim of efficiency.

Downloads

Published

2026-03-30

How to Cite

Honeyshadow: An AI-Adaptive Honeypot Network Using Reinforcement Learning for Dynamic Attacker Profiling and Threat Intelligence Extraction. (2026). The Critical Review of Social Sciences Studies, 4(1), 8065-8079. https://doi.org/10.59075/9s3xq912