PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement

Through the multi-agent task splitting and RAG-enhanced PLC high-interaction honeypot system, the traditional protection solution has solved the shortcomings in simulation accuracy, economic cost and intelligence level, and achieved efficient virtual and real isolation and active trapping defense, which has improved the security defense capabilities of the industrial control system.

CN120433955APending Publication Date: 2025-08-05GUANGZHOU UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510400471.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing industrial control system protection solutions are difficult to effectively deal with APT attacks and zero-day vulnerability threats. Traditional honeypot technology has shortcomings in simulation accuracy, economic cost and intelligence levels, and it is impossible to build an efficient virtual and real isolation and active trapping defense paradigm.

Method used

The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement is adopted to simulate the real industrial control environment through the high simulation equipment layer, and combine multi-agent collaborative analysis and RAG knowledge base to achieve high interactive and high concealment defense.

Benefits of technology

Build a highly realistic industrial honeypot environment that can accurately simulate protocol behavior and logic control functions, counter advanced threats in real time, and improve the security defense level of industrial control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433955A_ABST
    Figure CN120433955A_ABST
Patent Text Reader

Abstract

The invention relates to the crossing field of industrial control system (ICS) safety and artificial intelligence, and particularly discloses a PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement, which adopts a localized multi-agent collaborative architecture based on edge computing and is composed of a high-simulation equipment layer and an intelligent decision-making layer. The high-simulation equipment layer comprises a PLC dynamic mirror image, an HMI interface and a sensor data generator, and an active trapping environment is constructed through protocol fingerprint confusion and virtual and real data fusion technologies. The decision-making layer deploys a multi-agent task scheduling engine, integrates four kinds of agents including protocol analysis, behavior analysis, threat assessment and response generation, and realizes attack context perception and strategy dynamic generation based on a local RAG knowledge base. The load balancing agent dynamically allocates tasks according to equipment resources, and cooperates with offline knowledge update (USB flash disk encryption synchronization threat features) to form a closed-loop defense system, thereby ensuring physical isolation of an industrial network and realizing high-fidelity active defense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the cross - field of industrial control system (ICS) security and artificial intelligence, and more specifically, to a PLC high - interaction honeypot system based on multi - agent task splitting and RAG enhancement. Background Art

[0002] As a core component of the country's critical infrastructure, the industrial control system (ICS) is facing severe security threats and technical contradiction challenges. On the one hand, the attack means present a complex situation combining protocol - level penetration, AI - assisted attacks, and physical - level destruction. For example, 63% of PLC targeted attacks utilize the obfuscation vulnerability of the S7Comm protocol, and the passing rate of malicious ladder diagram code generated by GPT - 4 is as high as 68%. On the other hand, there are structural contradictions in the industrial security field: with the development of IIoT promoting the integration of OT and IT, 83% of PLCs still use plain - text communication; the 300 - ms - level delay introduced by traditional protection schemes far exceeds the 1 - ms real - time requirement of industrial control systems; a huge gap has formed between the 15 - year service cycle of equipment and the 47% annual growth rate of vulnerabilities. These contradictions make it difficult for traditional protection schemes such as firewalls and intrusion detection systems (IDS) to effectively cope with APT attacks and zero - day vulnerability threats, and there is an urgent need to build a new defense paradigm of virtual - physical isolation and active trapping.

[0003] Currently, the honeypot technology faces dilemmas in three aspects: simulation accuracy, economic cost, and intelligence level. Although virtual simulation solutions (such as HoneyPhy) have low costs, they have deficiencies in protocol fidelity and resource efficiency. For example, the Profinet timing error exceeds 35 μs, and a single instance requires 4.2 GB of memory. Physical device honeypots (such as GRFICS2.0) can highly restore hardware characteristics, but face high procurement costs (25,000 US dollars per device) and the risk of fingerprint leakage caused by firmware residues. Although the cutting - edge LLM - driven solutions (such as HoneyGPT - ICS) break through rule limitations, they also have problems such as high response latency (870 ms) and high protocol syntax error rate (41.2%). These three types of solutions all have obvious defects in key dimensions such as protocol depth, real - time guarantee, and domain knowledge adaptation, highlighting the core technical bottlenecks faced in building an industrial trapping environment.

[0004] Therefore, a PLC high - interaction honeypot system based on multi - agent task splitting and RAG enhancement is provided. Summary of the Invention

[0005] In order to solve the above - mentioned technical problems, this application is proposed.

[0006] Specifically, according to one aspect of the present application, there is provided a PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement, which includes:

[0007] An attack traffic recognition and processing module, configured to identify and filter potential malicious traffic through a high-fidelity device layer to obtain compliant traffic;

[0008] A protocol parsing agent module, configured to analyze the compliant traffic through an edge layer protocol splitter, identify the protocol type, and deeply extract protocol semantic features;

[0009] A threat assessment agent module, configured to perform threat assessment on detected traffic or behaviors through a rule engine and a lightweight AI model;

[0010] A task decomposition and multi-agent collaboration module, configured to efficiently detect attack behaviors and predict intentions through task decomposition and collaborative work among agents;

[0011] A RAG knowledge enhancement module, configured to obtain vulnerability information and attack characteristics related to the target device by retrieving a static device fingerprint library and a dynamic attack feature library;

[0012] A response generation agent module, configured to generate dynamic countermeasure strategies by combining the RAG knowledge base retrieval results with the real-time context, and coordinate the multi-agent data stream through a central scheduler to form a closed-loop analysis chain of "feature extraction - intention prediction - strategy generation";

[0013] A data closed-loop and system self-optimization module, configured to evaluate and optimize the adaptability of the defense strategy by collecting subsequent behavior data of the attacker, and update the knowledge base.

[0014] Preferably, the attack traffic recognition and processing module includes: an attacker sends malicious traffic to an exposed virtual programmable logic controller (PLC) IP address in a way of scanning or targeted attack; after the virtual programmable logic controller (PLC) at the terminal layer receives the data packet, it verifies the basic compliance of the protocol, where the verification content includes CRC check and transaction ID continuity; the compliant traffic is mirrored to the edge layer, and the original request data packet is continued to be temporarily stored at the terminal layer waiting for a response instruction.

[0015] Preferably, the threat assessment agent module includes: the rule engine layer matches the preset blacklist rules. If a blacklist rule is matched, the rule engine gives a rule engine score rule_score according to the severity of the rule; in the AI model inference layer, the temporal features of the traffic are analyzed by the TFLite lightweight LSTM model to generate an AI model score ai_score, where the range of the AI model score ai_score is from 0 to 1; the rule engine score rule_score and the AI model score ai_score are weighted and fused to obtain a threat score final_score, and the calculation formula is as follows: final_score = 0.6 * rule_score + 0.4 * ai_score; different response measures are taken according to the threat score final_score: if the threat score final_score < 0.7, trigger an edge autonomous response to mitigate potential threats; if the threat score final_score ≥ 0.7, upload the relevant metadata to the cloud layer and activate the multi-agent collaboration mechanism to further confirm and process high-threat events.

[0016] Preferably, the task decomposition and multi-agent collaboration module includes: multi-agent collaboration automatically generates a task directed acyclic graph DAG based on the attack complexity; the protocol parsing agent deeply analyzes the protocol data to extract semantic features; the behavior modeling agent constructs an attack temporal pattern based on the long short-term memory network and the attention mechanism model to predict the attack stage and potential targets.

[0017] Preferably, the RAG knowledge enhancement module includes: retrieving the static device fingerprint library, matching the model of the target programmable logic controller PLC, and obtaining known vulnerability information; through semantic retrieval technology, recalling attack features TTPs similar to the current attack behavior from the dynamic attack feature library.

[0018] Preferably, the response generation agent module includes: the response generation agent retrieves historical attack cases and vulnerability features through RAG technology, combines the real-time context, and generates a dynamic countermeasure strategy; the generated dynamic countermeasure strategy is sent to the terminal layer through the central scheduler; the terminal layer receives the response strategy sent by the cloud, constructs a protocol-compliant response packet; embeds a logical honey label in the response data, records the operation sequence of the attacker, and generates a timestamped attack chain log.

[0019] Among them, the terminal layer receives the response strategy sent by the cloud and constructs a protocol-compliant response packet. The specific construction steps are as follows: based on the request transaction ID of the attacker, add 1 to the transaction ID; dynamically calculate the CRC check code based on the response data.

[0020] Compared with the prior art, a PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement provided by this application is based on a multi-agent collaboration framework, decomposes PLC functions into atomic tasks and performs dynamic scheduling. At the same time, it uses RAG technology to enhance the context awareness and dynamic response capabilities of attack behaviors, and constructs an industrial control honeypot environment with high interactivity and high concealment. It can accurately simulate the protocol behaviors, logic control functions and network interaction characteristics of real industrial control devices. Through the closed-loop mechanism of multi-agent collaborative analysis (including protocol parsing, behavior modeling and response generation) and dynamic and static knowledge fusion retrieval (combining device fingerprint libraries and attack feature libraries), this system can effectively trap attackers and counter advanced threats in real time. This system provides an integrated "active trapping-dynamic interference-traceability and evidence collection" security protection capability for industrial control systems, significantly improving the security defense level in complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By describing the embodiments of the present application in more detail with reference to the accompanying drawings, the above and other objects, features and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 The overall system architecture diagram according to an embodiment of the present application is illustrated.

[0023] Figure 2 The key workflow diagram of the PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to an embodiment of the present application is illustrated.

[0024] Figure 3 The feature extraction flowchart of the protocol parsing agent module according to an embodiment of the present application is illustrated.

[0025] Figure 4 The workflow diagram of the task decomposition and multi-agent collaboration module according to an embodiment of the present application is illustrated.

[0026] Figure 5 The workflow diagram of the offline update mechanism according to an embodiment of the present application is illustrated. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0028] Embodiment:

[0029] Figure 1 The figure shows the overall system architecture diagram according to an embodiment of the present application. As Figure 1 shown, the system adopts a localized multi-agent collaborative architecture based on edge computing and consists of a high-fidelity device layer and an intelligent decision-making layer.

[0030] Among them, the high-fidelity device layer includes a PLC (Programmable Logic Controller) dynamic mirror, an HMI (Human Machine Interface), and a sensor data generator, and constructs an active trapping environment through protocol fingerprint obfuscation and virtual-real data fusion technologies. By constructing a high-interaction honeypot environment, dynamic interface baits, and sensor data simulation, it simulates a real industrial scenario, actively traps attackers and records their behaviors, and at the same time misleads their attack intentions, thereby protecting the real industrial system from threats. Its specific functions are as follows:

[0031] 1>. PLC dynamic mirror: The virtual PLC cluster module constructs a high-interaction honeypot environment that is indistinguishable from real industrial controllers through deep protocol simulation and dynamic behavior modeling. Based on reverse engineering, the firmware features of real PLCs are extracted (such as firmware hash, opcode distribution, and register mapping table of Siemens S7-1500), and multiple protocols (such as Modbus / TCP, S7Comm, EtherNet / IP, etc.) are supported. The dynamic response generation engine adjusts the returned data according to the attack context. For example, it returns real register values for requests of legal function codes, returns false success responses for abnormal write operations and records them in the isolated memory area. The trap injection mechanism inserts logical honey tags through a polymorphic strategy (static rules + AI dynamic generation) to mark attack traffic. The cluster is deployed in a containerized manner (Docker + Kubernetes), supports millisecond-level elastic scaling, and the end-to-end response latency ≤ 3ms, ensuring that 100% of attack behaviors are restricted in the simulation environment.

[0032] 2>. HMI (Human Machine Interface): The bait HMI clones the interaction logic and visual style of mainstream industrial monitoring software (such as WinCC, iFix) through dynamic interface generation technology. Automatically injects false controls (such as forged alarm pop-ups, trend charts) based on the attack context, and adds controllable deviations (±0.5% - 2%) to key parameters. At the same time, records the operation trajectories of attackers (such as mouse click coordinates, parameter modification sequences) for behavior analysis. By simulating the behavior patterns of real operators (such as periodic refreshing, parameter fine-tuning), it enhances the authenticity of the interaction and induces attackers to expose their operation intentions.

[0033] 3. Sensor Data Generator: The sensor data generator generates a time-series data stream based on a physical device model, incorporating Gaussian noise and periodic perturbations to simulate the signal characteristics of a real sensor. The data generator dynamically interacts with virtual PLC registers, such that pressure changes with temperature according to the ideal gas equation. Honeymarks are injected during the attack phase: data patterns with specific fingerprints (such as the 0xDEADBEEF byte sequence) are returned during the reconnaissance phase, triggering false threshold alarms (such as a sudden 10°C temperature rise) during the attack phase, misleading the attacker into judging the system status.

[0034] Specifically, the high-simulation device layer module uses a highly simulated PLC, HMI, and sensor environment to attract attackers and record their actions. It also dynamically responds to attackers' actions, tricking them into revealing more of their intentions. This fusion mechanism not only effectively identifies and records attack behavior but also confuses attackers through a highly interactive honeypot environment, enhancing the system's defense and threat detection capabilities.

[0035] Table 1 Description of agent related information

[0036]

[0037] like Figure 1 As shown in Table 1, the decision-making layer deploys a multi-agent task scheduling engine, integrating four types of agents: protocol parsing, behavior analysis, threat assessment, and response generation. Relying on the local RAG knowledge base, it realizes attack context perception and dynamic strategy generation. The threat assessment agent uses an LSTM-Attention hybrid model to score attacks. For high-threat attacks (score ≥ 0.7), it initiates in-depth analysis and injects false responses containing timing noise. For low-threat attacks (score < 0.7), it quickly returns a preset error code through template matching. At the same time, the load balancing agent dynamically allocates computing tasks according to the device resource status, and cooperates with the offline knowledge update mechanism (U disk encryption synchronizes threat characteristics) to form a closed-loop defense system, achieving high-fidelity active defense capabilities while ensuring the physical isolation of the industrial network.

[0038] Figure 2 The diagram shows the key workflow of the PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to the embodiment of the present application. Figure 2As shown, the system adopts a defense closed-loop that links a multi-agent decision center and a high-fidelity device layer. It specifically includes: an attack traffic identification and processing module, which is used to identify and screen potential malicious traffic through the high-fidelity device layer to obtain compliant traffic; a protocol parsing agent module, which is used to analyze the compliant traffic through an edge layer protocol splitter, identify the protocol type, and deeply extract protocol semantic features; a threat assessment agent module, which is used to assess the threat of detected traffic or behavior through a rule engine and a lightweight AI model; a task decomposition and multi-agent collaboration module, which is used to efficiently detect attack behaviors and predict intentions through task decomposition and collaborative work among agents; a RAG knowledge enhancement module, which is used to obtain vulnerability information and attack characteristics related to target devices by retrieving a static device fingerprint library and a dynamic attack feature library; a response generation agent module, which is used to generate dynamic countermeasures by combining the retrieval results of the RAG knowledge base and real-time context, and coordinate the multi-agent data stream through a central scheduler to form a closed-loop analysis chain of "feature extraction - intention prediction - policy generation"; a data closed-loop and system self-optimization module, which is used to evaluate and optimize the adaptability of defense strategies by collecting subsequent behavior data of attackers and update the knowledge base.

[0039] In the embodiment of the present application, the attack traffic identification and processing module is used to identify and screen potential malicious traffic through the high-fidelity device layer to obtain compliant traffic. Its specific implementation steps are as follows:

[0040] 1). The attacker sends malicious traffic, such as abnormal Modbus function codes and S7Comm exploit packages, to the exposed virtual programmable logic controller (PLC) IP address in the form of scanning or targeted attacks.

[0041] 2). After the virtual programmable logic controller (PLC) at the terminal layer receives the data packet, it verifies the basic compliance of the protocol. The verification content includes CRC check, transaction ID continuity, etc. For example: calculate the CRC16 checksum for a Modbus / TCP request. If it does not match the packet header, it is recorded as a protocol tampering attack.

[0042] 3). Mirror the compliant traffic to the edge layer, and continue to temporarily store the original request data packet at the terminal layer, waiting for a response instruction.

[0043] In the embodiment of the present application, the protocol parsing agent module is used to analyze the compliant traffic through an edge layer protocol splitter, identify the protocol type, and deeply extract protocol semantic features. Specifically, such as Figure 3As shown, the protocol parsing agent module identifies the protocol type from the original message through the edge layer protocol splitter. The protocol types include Modbus / TCP, OPC UA, etc. Then, it extracts the key fields through the Modbus parser and the S7Comm parser, and thus outputs the structured feature vector. Among them, these key fields include function codes, register addresses, operation values, etc.

[0044] Among them, the protocol parsing agent undertakes the tasks of real-time identification and preprocessing of industrial protocol traffic. It achieves accurate protocol classification through port matching (such as port 502 corresponding to Modbus / TCP, port 4840 corresponding to OPC UA) and deep packet detection technology (parsing feature fields such as function codes and transaction IDs in the protocol header), and synchronously extracts key metadata (source IP, register address, operation type, etc.) to generate a structured summary. It adopts zero-copy data mirroring and DPDK acceleration technology to ensure high throughput processing capacity (18,000 packets per second). While ensuring an identification accuracy of 99.3%, it provides lightweight protocol feature vectors for downstream modules, forming the data entry of the threat detection pipeline.

[0045] That is, the protocol parsing agent module significantly improves the real-time performance, accuracy, and overall performance of the system through high-precision protocol identification, efficient data processing, and lightweight feature extraction, providing a solid technical foundation for threat detection and defense in industrial environments.

[0046] In the embodiment of this application, the threat assessment agent module is used to perform threat assessment on the detected traffic or behavior through a rule engine and a lightweight AI model. The specific implementation process is as follows:

[0047] 1). The rule engine layer matches the preset blacklist rules (such as function code 90, unusal register range, etc.). If a blacklist rule is matched, the rule engine gives a rule engine score rule_score according to the severity of the rule.

[0048] Among them, the rule engine layer has built-in more than 200 industrial protocol feature rules, covering common attack patterns (such as function code abuse, register out-of-bounds access). The rule library supports dynamic updates and can receive incremental rule packages from the cloud (synchronized daily) to timely respond to new attack variants.

[0049] 2). In the AI model inference layer, the TFLite lightweight LSTM model analyzes the temporal features of the traffic (such as request interval, function code distribution, etc.) to generate an AI model score ai_score. Among them, the range of the AI model score ai_score is from 0 to 1.

[0050] Among them, the AI model layer deploys a lightweight LSTM network based on the TensorFlow Lite framework. The model is trained with a special industrial scenario dataset (including 120 million industrial control protocol samples) and has a higher detection rate for time-series attacks (such as slow scanning and low-frequency penetration). In addition, the model is updated quarterly through knowledge distillation technology to maintain the latest threat recognition ability.

[0051] 3), Weightedly fuse the rule engine score rule_score and the AI model score ai_score to obtain the threat score final_score. The calculation formula is as follows:

[0052] final_score = 0.6 * rule_score + 0.4 * ai_score;

[0053] 4), Take different response measures according to the threat score final_score: If the threat score final_score < 0.7, trigger edge autonomous response (such as returning normal data) to mitigate potential threats; If the threat score final_score ≥ 0.7, upload the relevant metadata to the cloud layer and activate the multi-agent collaboration mechanism to further confirm and process high-threat events.

[0054] That is, the threat assessment agent module realizes μs-level real-time threat analysis based on the local rule library and the quantized AI model, intercepts known malicious operations (such as Modbus abnormal function code 90) through the preset function code blacklist, and uses sliding window statistics (such as more than 100 accesses within 10 seconds trigger an alarm) to identify high-frequency scanning behaviors; Synchronously run an 8-layer TFLite quantized LSTM model, input feature vectors such as protocol type, register address, and time series interval, and output a threat score of 0-1 (such as a score of 0.82 is marked as high risk), achieving an average detection delay of 2.3 ms and an unknown attack detection rate of 72.5% on the edge device, providing an accurate judgment basis for the local decision center.

[0055] In the embodiment of the present application, the task decomposition and multi-agent collaboration module is used to efficiently detect attack behaviors and predict intentions through task decomposition and collaborative work among agents. As Figure 4 shown, the specific implementation process of the task decomposition and multi-agent collaboration module is as follows:

[0056] 1), Multi-agent collaboration automatically generates a task directed acyclic graph DAG based on the attack complexity.

[0057] 2), The protocol parsing agent deeply analyzes the protocol data to extract semantic features, such as the ROSCTR type of S7Comm.

[0058] 3) The behavior modeling agent constructs the attack timing pattern based on the Long Short-Term Memory with Attention (LSTM-Attention) model to predict the attack stage and potential targets. The attack stage includes reconnaissance, weaponization, infiltration, etc., and potential targets such as PLC shutdown instructions and HMI parameter tampering.

[0059] In the embodiment of the present application, the RAG knowledge enhancement module is used to obtain vulnerability information and attack characteristics related to the target device by retrieving the static device fingerprint library and the dynamic attack feature library. The specific implementation steps include: retrieving the static device fingerprint library, matching the model of the target programmable logic controller (PLC), and obtaining known vulnerability information such as CVE-2022-42475; recalling attack characteristics (TTPs) similar to the current attack behavior from the dynamic attack feature library through semantic retrieval technology, such as MITRE T0863.

[0060] In the embodiment of the present application, the response generation agent module is used to generate a dynamic countermeasure strategy by combining the RAG knowledge base retrieval result and the real-time context, and coordinate the multi-agent data stream through the central scheduler to form a closed-loop analysis chain of "feature extraction - intent prediction - strategy generation".

[0061] The specific implementation process of the response generation agent module is as follows:

[0062] 1) The response generation agent retrieves historical attack cases and vulnerability characteristics through RAG technology, and combines the real-time context (such as attack traffic characteristics, protocol types, etc.) to generate a dynamic countermeasure strategy. For example, for a specific vulnerability exploitation attempt, a strategy of "simulating firmware vulnerability - returning false success" is generated.

[0063] 2) The generated dynamic countermeasure strategy is sent to the terminal layer through the central scheduler.

[0064] 3) The terminal layer receives the response strategy sent from the cloud and constructs a protocol-compliant response packet. Specifically: inherit the request transaction ID, that is, based on the attacker's request transaction ID, add 1 to the transaction ID, such as 0x8A3F → 0x8A40, and dynamically calculate the CRC checksum to ensure that the response packet complies with the protocol specification.

[0065] 4) Embed a logical honey label in the response data, such as returning an abnormal value 128 for register address 40009; and record the attacker's operation sequence (such as "scan → vulnerability exploitation → lateral movement") to generate a timestamped attack chain log.

[0066] In the embodiments of the present application, the data closed-loop and system self-optimization module is used to evaluate and optimize the self-adaptability of the defense strategy and update the knowledge base by collecting the subsequent behavior data of the attacker. The specific implementation steps are as follows: collect the subsequent behavior data of the attacker (such as whether a false vulnerability is triggered) and feedback it to the cloud; dynamically adjust the policy weights according to the subsequent behavior data of the attacker through a reinforcement learning algorithm (such as DQN); in the offline update mechanism, adopt a dual-network isolation architecture to update the knowledge base offline.

[0067] Specifically, as Figure 5 shown, the offline update mechanism adopts a strict dual-network isolation architecture, which consists of an external network data collection end, a secure transmission carrier, and an internal network update execution end. Specifically, on the external network side, encrypted update packages (including threat features, device fingerprints, and response templates) are generated through mirror traffic capture and sensitive information filtering, and after being digitally signed (using the SM2 algorithm), they are written into a dedicated secure USB flash drive (hardware-level write protection + physical anti-tampering design); on the internal network side, the USB flash drive is authenticated and the update package is decrypted through a trusted channel, and the data integrity is ensured by combining version number verification (to prevent replay attacks) and hash tree verification. The local knowledge base is updated in an incremental synchronization manner (retaining 5 historical versions to support second-level rollback), and finally, service hot loading is triggered to achieve interruption-free service update. The overall process covers data collection, signature encryption, physical transmission, decryption verification, version control, and exception fusing mechanisms, realizing the secure synchronization of key threat intelligence and the dynamic evolution of defense strategies in a pure offline environment, meeting the stringent requirements of industrial control systems for data isolation and update reliability.

[0068] It is worth mentioning that the RAG knowledge base module constructs a multi-modal knowledge center in the field of industrial security, integrating two core data layers of static device fingerprints and dynamic attack intelligence. The static layer stores firmware hashes, register mapping tables, and process logic templates of more than 10,000 real PLC devices, and extracts unique device features through reverse engineering; the dynamic layer accesses MITRE ATT&CK ICS attack patterns, NVD vulnerability libraries, and honeypot interaction logs in real time, and updates global industrial control threat intelligence at a 15-second level. The two-level data is encoded into 768-dimensional semantic vectors through a unified vectorization engine (based on a Sentence-BERT fine-tuning model), and a Faiss vector index with a scale of billions is constructed, supporting hybrid queries of "exact retrieval + semantic retrieval", such as exactly matching device models and CVE vulnerabilities, or semantically searching for disposal strategies for similar attack chains.

[0069] Specifically, the knowledge base adopts a dual-channel update strategy: static data is synchronized offline quarterly (e.g., new device fingerprints are added to the database), while dynamic data relies on a streaming computing engine (Apache Flink) to process attack logs in real time, extract TTPs features, and incrementally update the index. During the retrieval phase, after entering attack context features, the system prioritizes SQL exact matching (e.g., function code + register address combination query). If no results are found, Faiss semantic search is initiated, recalling the top five similar cases from billions of vectors and combining it with a large language model (LLM) to generate context-adaptive response strategies. For example, when an abnormal Modbus function code is detected, RAG can correlate historical vulnerability exploitation records for the same device model and generate a dynamic response template containing false register values, enhancing the credibility and stealthiness of the countermeasures.

[0070] In summary, the PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to the embodiment of the present application is explained, which adopts a localized multi-agent collaborative architecture based on edge computing, and is composed of a high-simulation device layer and an intelligent decision-making layer. The high-simulation device layer includes a PLC dynamic image, an HMI human-machine interface, and a sensor data generator, and constructs an active trapping environment through protocol fingerprint confusion and virtual-real data fusion technology. The decision-making layer deploys a multi-agent task scheduling engine, integrates four types of agents: protocol parsing, behavior analysis, threat assessment, and response generation, and relies on the local RAG knowledge base to realize attack context perception and dynamic strategy generation. At the same time, the load balancing agent dynamically allocates computing tasks according to the status of device resources, and cooperates with the offline knowledge update mechanism (U disk encryption synchronization threat characteristics) to form a closed-loop defense system to ensure the physical isolation of the industrial network and achieve high-fidelity active defense capabilities.

[0071] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit of the technical solutions of the present invention.

Claims

1. A PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement, characterized by: include: Attack traffic identification and processing module, used to identify and filter potential malicious traffic through the high-simulation device layer to obtain compliant traffic; A protocol parsing agent module is used to analyze the compliant traffic through the edge layer protocol splitter, identify the protocol type, and deeply extract the protocol semantic features; The threat assessment agent module is used to perform threat assessment on detected traffic or behavior through a rule engine and a lightweight AI model; The task decomposition and multi-agent collaboration module is used to efficiently detect attack behaviors and predict intentions through task decomposition and collaboration between agents. The RAG knowledge enhancement module is used to obtain vulnerability information and attack signatures related to the target device by searching the static device fingerprint library and the dynamic attack signature library; The response generation agent module combines RAG knowledge base retrieval results with real-time context to generate dynamic countermeasures. It also coordinates multi-agent data streams through a central scheduler to form a closed-loop analysis chain of "feature extraction-intention prediction-strategy generation"; The data closed-loop and system self-optimization module is used to evaluate and optimize the adaptability of defense strategies by collecting subsequent attacker behavior data and updating the knowledge base; The task decomposition and multi-agent collaboration module includes: Multi-agent collaboration automatically generates a directed acyclic graph (DAG) of tasks based on attack complexity; The protocol parsing agent performs in-depth analysis of protocol data to extract semantic features; The behavior modeling agent constructs attack timing patterns based on the long short-term memory network and attention mechanism model to predict attack stages and potential targets.

2. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 1 is characterized in that: The attack traffic identification and processing module includes: The attacker sends malicious traffic to the exposed virtual programmable logic controller (PLC) IP address through scanning or targeted attacks. After receiving the data packet, the virtual programmable logic controller (PLC) at the terminal layer verifies the basic compliance of the protocol, including CRC check and transaction ID continuity. Mirror compliant traffic to the edge layer, and temporarily store the original request data packet at the terminal layer, waiting for response instructions.

3. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 2 is characterized in that: The threat assessment agent module includes: The rule engine layer matches the preset blacklist rules. If a blacklist rule is matched, the rule engine gives a rule engine score rule_score based on the severity of the rule. At the AI model inference layer, the TFLite lightweight LSTM model is used to analyze the time series characteristics of the traffic and generate an AI model score ai_score, where the AI model score ai_score ranges from 0 to 1. The rule engine score rule_score and the AI model score ai_score are weighted and fused to obtain the threat score final_score. The calculation formula is as follows: final_score=0.6*rule_score+0.4*ai_score; Different response measures are taken according to the threat score final_score: if the threat score final_score is less than 0.7, an autonomous edge response is triggered to mitigate potential threats; if the threat score final_score is greater than or equal to 0.7, relevant metadata is uploaded to the cloud layer, and the multi-agent collaboration mechanism is activated to further confirm and handle high-threat events.

4. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 3 is characterized in that: The RAG knowledge enhancement module includes: Search the static device fingerprint library, match the target programmable logic controller (PLC) model, and obtain known vulnerability information; Through semantic retrieval technology, attack signatures TTPs similar to the current attack behavior are recalled from the dynamic attack signature library.

5. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 4 is characterized in that: The response generation agent module includes: The response generation agent uses RAG technology to retrieve historical attack cases and vulnerability characteristics, and combines them with real-time context to generate dynamic countermeasures. The generated dynamic countermeasure strategy is sent to the terminal layer through the central scheduler; The terminal layer receives the response policy sent by the cloud and constructs a protocol-compliant response packet; Embed logical honeytokens in the response data and record the attacker's operation sequence to generate an attack chain log with a timestamp.

6. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 5 is characterized in that: The terminal layer receives the response policy sent by the cloud and constructs a protocol-compliant response packet. The specific construction steps are as follows: Based on the attacker's request transaction ID, increase the transaction ID by 1; The CRC checksum is dynamically calculated based on the response data.

7. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 6 is characterized in that: The data closed loop and system self-optimization module includes: Collect the attacker's subsequent behavior data and feed it back to the cloud; Through reinforcement learning algorithms, the strategy weights are dynamically adjusted based on the attacker's subsequent behavior data; In the offline update mechanism, a dual-network isolation architecture is used to update the knowledge base offline.

8. The PLC high-interaction honeypot system based on multi-agent task splitting and RAG enhancement according to claim 7 is characterized in that: The offline update mechanism includes an external network data collection terminal, a secure transmission carrier and an internal network update execution terminal.

Citation Information

Cited By

  • Code agent construction system, construction method thereof and code generation method

    CN121168498A

  • Multi-agent-based industrial process control system and method, agents and medium

    CN121209434A

  • AD graph honeypot optimization deployment method based on large model and multi-agent collaboration and related equipment

    CN121664468A

  • Asset simulation system for threat perception

    CN121902146A

  • An asset emulation system for threat awareness

    CN121902146B