Network security operation method and system based on adaptive multi-agent threat decoy

CN122513196BActive Publication Date: 2026-09-08QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610969300.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-08
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

目前主流SOAR方案主要通过集成预定义的静态剧本(Playbook)与专家规则引擎实现安全事件的自动化处置,然而面对具备高度对抗性、意图多变的现代攻击手段,现有自动化运维体系在智能决策、响应实时性以及主动诱捕能力方面仍存在根本性短板,无法满足网络实战化防护的核心需求

Benefits of technology

本发明通过Bi-GRU同步挖掘网络业务流量报文级微观时序特征与流级宏观统计特征,基于KL散度流量偏离度精准捕捉威胁上下文对象,解决传统单一特征难以全面表征异常流量的问题;搭建全局异步共享内存区,以快照缓存与状态管理实现威胁上下文可查询留存,规避异构模块异步调度造成的防御上下文丢失缺陷;结合大模型语义推理与RAG检索融合资产拓扑、威胁情报知识,完成碎片化告警聚合重构与攻击链路梳理,实现长链条攻击意图深度解析;通过认知-执行映射算子将防御导向转化为强化学习超参数向量,利用MAPPO算法在动态奖励约束下驱动诱捕环境自适应演进,借助攻防实时博弈动态更新诱导策略,有效增强网络未知攻击、链式攻击的识别研判能力,同时提升主动诱捕的适配性、实时性与长期留存防御效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513196B_ABST
    Figure CN122513196B_ABST
Patent Text Reader

Abstract

This invention proposes a network security operation method and system based on adaptive multi-agent threat trapping, belonging to the field of network security technology. It includes: extracting packet-level temporal features and flow-level statistical features of network service traffic to construct traffic temporal features; generating threat context objects based on traffic temporal features when current service behavior is abnormal; storing threat context objects using an indexed asynchronous buffer and constructing a threat context set; semantically enhancing the threat context set and calculating the probability distribution of defense strategies; constructing a policy-weight mapping operator to transform the probability distribution of defense strategies into a reinforcement learning reward function weight vector; and using the MAPPO algorithm based on the reward function weight vector to drive the agents to adaptively evolve under dynamic reward weight constraints, thereby achieving network security operation. This invention improves the accuracy of network attack identification, the completeness of context association, and the long-term effectiveness of proactive defense trapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a network security operation method and system based on adaptive multi-agent threat trapping. Background Technology

[0002] As enterprises deepen their digital transformation, network environments are becoming increasingly heterogeneous and complex, and cybersecurity has shifted from traditional "single-point perimeter defense" to "in-depth cyber warfare." In the field of enterprise-level cybersecurity operations and maintenance (SecOps), attackers, represented by APT groups, often employ advanced penetration techniques such as encrypted tunneling bypass, fileless attacks, and low-frequency lateral movement to breach boundaries and remain lurking in the core business areas of enterprises for extended periods, posing a serious threat to core assets.

[0003] To address increasingly complex security threats, Security Orchestration Automation and Response (SOAR) technology has emerged. Currently, mainstream SOAR solutions primarily automate the handling of security incidents by integrating predefined static playbooks with expert rule engines. However, facing highly adversarial and ever-changing modern attack methods, existing automated operation and maintenance systems still have fundamental shortcomings in intelligent decision-making, real-time response, and proactive decoy capabilities, failing to meet the core requirements of practical network protection. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a network security operation method and system based on adaptive multi-agent threat trapping, which improves the accuracy of network attack identification, the integrity of contextual association, and the long-term effectiveness of proactive defense trapping.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a network security operation method based on adaptive multi-agent threat trapping, comprising: Extract packet-level time-series features and flow-level statistical features from network service traffic to construct traffic time-series features; when current service behavior is abnormal, generate threat context objects based on traffic time-series features; The threat context objects are stored using an indexed asynchronous buffer, and a threat context set is constructed based on the time-aligned threat context objects; Semantic enhancement is performed on the threat context set, and the probability distribution of defense strategies is calculated; Construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function; Based on the reward function weight vector, the MAPPO algorithm is used to drive the agent's adaptive evolution under dynamic reward weight constraints, and achieve network security operation through real-time feedback game with the attacker.

[0006] Secondly, the present invention provides a network security operation system based on adaptive multi-agent threat trapping, comprising: The multi-granularity feature perception module is used to extract packet-level temporal features and flow-level statistical features of network service traffic to construct traffic temporal features; when the current service behavior is abnormal, a threat context object is generated based on the traffic temporal features. A shared blackboard state maintenance module is used to store the threat context objects using an indexed asynchronous buffer and to construct a threat context set based on the time-aligned threat context objects. The RAG-enhanced intent analysis module is used to semantically enhance the threat context set and calculate the probability distribution of defense strategies. The policy-weight dynamic mapping module is used to construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function. The adaptive interactive game module is used to drive the agent's adaptive evolution based on the reward function weight vector and the MAPPO algorithm under dynamic reward weight constraints. It achieves network security operation through real-time feedback game with the attacker.

[0007] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the network security operation method based on adaptive multi-agent threat trapping described in the first aspect.

[0008] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the network security operation method based on adaptive multi-agent threat trapping described in the first aspect.

[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes Bi-GRU to synchronously mine micro-temporal features at the packet level and macro-statistical features at the flow level of network service traffic. Based on KL divergence and traffic deviation, it accurately captures threat context objects, solving the problem that traditional single features cannot comprehensively represent abnormal traffic. A global asynchronous shared memory area is constructed, using snapshot caching and state management to achieve queryable retention of threat context, avoiding the loss of defense context caused by asynchronous scheduling of heterogeneous modules. Combining large-model semantic reasoning and RAG retrieval with asset topology and threat intelligence knowledge, fragmented alarm aggregation and reconstruction, and attack chain analysis are completed, enabling deep analysis of long-chain attack intentions. Through a cognitive-execution mapping operator, defense guidance is transformed into a reinforcement learning hyperparameter vector. The MAPPO algorithm drives the adaptive evolution of the trapping environment under dynamic reward constraints. By dynamically updating the induction strategy through real-time attack-defense game theory, the invention effectively enhances the identification and judgment capabilities of unknown network attacks and chain attacks, while improving the adaptability, real-time performance, and long-term retention of proactive trapping defenses.

[0010] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.

[0011] Figure 1 A main flowchart of a network security operation method based on adaptive multi-agent threat trapping provided in an embodiment of the present invention; Figure 2 This is an overall flowchart of a network security operation method based on adaptive multi-agent threat trapping, provided for an embodiment of the present invention. Detailed Implementation

[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0013] Example 1 like Figure 1 As shown in the figure, this embodiment discloses a network security operation method based on adaptive multi-agent threat trapping, including the following steps: S1: Extract packet-level temporal features and flow-level statistical features of network service traffic to construct traffic temporal features; when the current service behavior is abnormal, generate a threat context object based on the traffic temporal features; S2: Store the threat context object using an indexed asynchronous buffer, and construct a threat context set based on the time-aligned threat context object; S3: Perform semantic augmentation on the threat context set and calculate the probability distribution of defense strategies; S4: Construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function; S5: Based on the reward function weight vector, the MAPPO algorithm is used to drive the agent's adaptive evolution under dynamic reward weight constraints, and achieve network security operation through real-time feedback game with the attacker.

[0014] Next, combined Figure 2 This embodiment provides a detailed description of a network security operation method based on adaptive multi-agent threat trapping.

[0015] The method proposed in this embodiment is applicable to network security operations across the entire domain, and as an implementation method, it can be applied to enterprise intranet environments.

[0016] (I) Anomaly detection module based on Bi-GRU and multi-granularity temporal features This module, serving as the perception layer of the entire system, is responsible for receiving raw network traffic and accurately capturing concealed attack fingerprints within a massive traffic context. Taking a continuous packet sequence within a sliding window as input, this module extracts multi-scale features and calculates behavioral distribution deviations, ultimately outputting a high-dimensional feature snapshot (i.e., a threat context object). ) to the shared blackboard scheduling layer.

[0017] 1. Multi-scale feature construction and temporal dependency coding This module collects network traffic in units of a sliding window and converts the continuous packet sequence within the window into a time-series input. For each window... Let the total number of messages in the window be... Define message index Construct message-level feature sequences as message-level micro-temporal features: ; in, Indicates the first The local feature vector of a message can be represented as: ; in, Indicates the first The arrival interval of each message. Indicates the message payload length. Indicates the protocol type encoding.

[0018] Meanwhile, to characterize window-level flow behavior, statistical feature vectors are extracted as macroscopic statistical features at the flow level: ; in, Indicates the first Flow level statistical feature vectors for each time window; This represents the protocol distribution entropy within the window. Indicates connection frequency. This represents aggregate statistics such as total bytes or total packets.

[0019] To capture the bidirectional dependencies of attack behaviors over time, a Bi-GRU algorithm is used to analyze the sequences. Encode the sequence and take the first element within the sequence. Local feature vectors of individual messages As input: ; ; The forward and backward hidden states are then concatenated to obtain the window-level temporal representation vector: ; Subsequently, the time series representation vector Macroscopic statistical characteristics of window-level flow levels The data is then fused to form a joint behavioral representation of the current window, which serves as a temporal feature of traffic and a core input for subsequent anomaly detection and threat context construction. .

[0020] 2. Offset Quantization Determination and Threat Context Generation Based on Distribution Divergence This module performs anomaly offset determination on the joint behavior representation of the current window, based on the baseline of normal business traffic. The system first maintains the baseline distribution of normal business behavior. and jointly represent the current window Mapped to real-time distribution in the same behavior space Based on this, KL divergence is used to measure the degree of deviation of the current behavior from the baseline: ; in, Indicates the number of traffic flows within the normal baseline. Standard probability values ​​for network-like behavioral characteristics; Indicates the current window's combined representation After mapping, the first Real-time probability values ​​of network-like behavioral characteristics.

[0021] The window is considered to have abnormal behavior when the following conditions are met: Dynamic sensitivity threshold: ; Once an anomaly detection is triggered, the system immediately generates a threat context object. And asynchronously write it to the shared blackboard. The object is defined as: ; In the formula, The timestamp of the exception trigger (the original time, which does not change with the system processing). This indicates source information, including source IP, destination IP, source port, destination port, and protocol number; For asset identification, it represents a set of attributes of the affected asset, including asset type, business importance, network location, etc. The joint behavior of the current window represents the temporal characteristics of the traffic flow; Scoring for anomalous divergence; This represents the current status of the situation. The initial value is set to 0, indicating that the situation is pending further analysis.

[0022] As a structured threat context object, it uniformly encapsulates multi-dimensional key information of network anomalies, providing standardized input for subsequent threat assessment. The assessment process is based on the complete attributes of this object to accurately identify threat types and attack intentions, and match and schedule corresponding security response strategies.

[0023] Among them, asset identification This can be quickly obtained through a local asset mapping table or network topology information. By writing to the shared blackboard through a non-blocking communication mechanism, the perception layer and the subsequent judgment layer are decoupled, thus avoiding the blocking of network traffic processing during the anomaly detection process.

[0024] 3. Module Output Definition After the above processing, this perception module finally outputs a structured threat context object. This output contains a set of high-dimensional feature snapshots fused from message-level and stream-level data. It also includes quantized divergence anomaly scores and affected asset attributes.

[0025] This output serves as the direct input to the subsequent "asynchronous state maintenance module based on shared blackboard," providing it with a data source for non-blocking writes, thereby initiating the subsequent analysis and scheduling process for automated response.

[0026] (ii) Asynchronous state maintenance module based on shared blackboard This module is used to establish a unified asynchronous shared state area between the perception layer Agent_DL and the decision layer Agent_LLM, so as to solve the context inconsistency problem caused by the difference in computing latency between high-frequency perception and low-frequency inference.

[0027] This module uses the threat context object output by the perception layer. Instead of regenerating threat objects, the system receives, caches, indexes, maintains their state, and performs spatiotemporal alignment processing on them, ultimately outputting a set of structured threat contexts that can be directly invoked by the decision-making layer.

[0028] 1. Global asynchronous shared buffer and non-blocking scheduling The shared blackboard module maintains an indexed asynchronous buffer to store all threat context objects: ; in, A unique identifier for each threat object. For the first A threat context object.

[0029] When the perception layer generates new Subsequently, the system inserts the data into the buffer using a non-blocking write method and establishes a multi-dimensional index based on timestamps, asset identifiers, and connection characteristics to support subsequent rapid retrieval and correlation analysis.

[0030] This mechanism ensures that the perception layer can continue to operate without waiting for the processing results of the decision layer, thereby avoiding detection blockage or data loss caused by the inference delay of large models.

[0031] 2. Threat context state maintenance and lifecycle management exist After data is written to the shared blackboard, the system dynamically maintains its state. State variables are defined. Describes the threat handling phase, with values ​​as follows: To be assessed. Confirmed. Response has been issued.

[0032] State updates are driven by the output of the decision layer, and the state transition process is represented as follows: ; in, It indicates the assessment or response result at the current moment.

[0033] During the state transition process, the system synchronously records the corresponding judgment conclusions, response strategies, and execution results, and appends them to the system. This results in an enhanced threat context object with complete lifecycle information.

[0034] 3. Decision-making delay compensation and spatiotemporal alignment Because decision-making reasoning involves time delays, directly executing responses based on raw timestamps may lead to a mismatch between the strategy and the current attack state. Therefore, this module introduces a time alignment mechanism: ; in, This is the original exception trigger time. To account for the time delay between anomaly detection and policy generation, This is the compensation coefficient.

[0035] Based on the aligned timestamps, the system combines historical context information from the blackboard to estimate the evolution trend of attack behavior within the decision window, thereby generating a more forward-looking response reference.

[0036] 4. Module Output Definition After the above processing, the shared blackboard module outputs a set of structured threat contexts: ; Each of them For threat objects that have undergone the above processing: state variables have been attached. (Lifecycle stage), integrated assessment conclusions and response strategy information, and completed time alignment ( ), retain original behavioral characteristics and asset information.

[0037] This collection supports fast retrieval based on multidimensional indexes and serves as direct input to the decision layer (RAG enhancement and intent analysis module) for subsequent attack chain reconstruction and defense strategy generation.

[0038] (III) Deep Analysis Module Based on RAG Enhancement and Intent Judgment In this system, the decision-making layer (LLM commander) bears the core task of transforming structured threat contexts into global defense strategies. This module uses the set of structured threat contexts output by the shared blackboard module. As input, through knowledge enhancement and semantic reasoning, we can achieve in-depth analysis from fragmented alerts to attack intent, and output semantic defense strategy distribution. 1. RAG-based multidimensional security knowledge retrieval and semantic enrichment The decision-making layer first retrieves the threat context object to be processed from the shared blackboard output set: ; Based on the asset identification within it With behavioral characteristics (Right now (Joint behavioral representation of traffic time-series features) to construct query vectors And using RAG (Retrieval Enhanced Generation) technology, the top-K most relevant background knowledge points are matched in the vector knowledge base: ; It is the vector representation of the k-th knowledge / document in the knowledge base. The norm of a vector.

[0039] Search results This includes: asset topology information (network location, criticality), historical vulnerability records (CVE, etc.), and threat intelligence (attack groups, TTPs).

[0040] Based on this, semantic enhancement is performed on the original threat context to obtain: ; in, For knowledge association weight, for The search results.

[0041] This process expands the originally isolated detection results into an enhanced context with a global semantic background, eliminating the problem of "knowledge silos".

[0042] 2. Attack chain reconstruction and path probability assessment based on thought chain In obtaining enhanced context Subsequently, the LLM commander, based on the Chain of Thought (CoT) reasoning, models the current threat within a dynamic attack graph, constructing the dynamic attack graph: ; in, For asset identifier A t The set of asset state nodes obtained by mapping Let be the set of directed edges relating attack path transitions. by and Define the topology of the attack path to provide a structural framework for probabilistic modeling of the attack path for the alarm sequence below.

[0043] The set of directed edges is derived from the attack transfer relationships between assets extracted from the threat intelligence retrieved by RAG. .

[0044] Based on the structural constraints of G, the alarm sequence is defined as follows: The system determines that it belongs to a certain attack path (i.e. middle Node (The ordered sequence of states formed by the connection) By modeling the probability, we obtain the attack path probability evaluation results: ; in, Indicates the first step in the attack path The attack state actions are used to characterize the evolution of the attacker's behavior on the corresponding asset node; these attack state actions correspond to a dynamic attack graph. Middle node set The state transition process refers to the transfer behavior from one asset state node to another; for example, attack phase behaviors such as exploitation, lateral movement, and privilege escalation. Enhanced context supply. This indicates the RAG search results.

[0045] Through this probabilistic modeling, the system can embed current anomalous behavior into the complete attack chain, achieving a leap in reasoning from local features to global intent. The attack path probability assessment results are not directly used as policy output, but rather as risk posture features, which are integrated and updated in real time to reflect the current network security posture. Therefore, the probability distribution of subsequent defense strategies is as follows. What it depends on It has pre-embedded the likelihood information of attack path attribution, enabling attack intent inference and defense response decisions to be made through... To achieve causal coherence.

[0046] 3. Module Output Definition This module ultimately outputs a set of semantic defense strategy distributions: ; in, Indicates the first A semantically enhanced threat trapping target; This represents the structured threat context enhancement set, consisting of all semantically enhanced models. constitute.

[0047] Specifically, for each semantic threat trapping target Construct a corresponding set of defense strategies; let the total number of defense strategies in the set be... Define defense strategy index : ; in, This is a set of semantic defense strategies (such as "induced isolation", "intelligence trapping", and "continuous monitoring"). Indicates the first Several candidate semantic defense strategies.

[0048] definition: ; Indicating the current cybersecurity situation Below, targeting semantically enhanced threat trapping targets The generated probability distribution of defense strategies; Indicates the system at the 1st The global state vector at time step (k) is used to characterize the overall situational information of the network environment, including traffic characteristics, topology, and current defense deployment status. The probability distribution of the defense strategy is calculated as follows: ; ; in, This represents the semantic representation vector generated by the encoder's enhanced context. and These represent the mapping weight matrix and the bias term, respectively.

[0049] therefore: ; in, Representation strategy The activation probability satisfies: ; Right now, It is not a single strategy, but a set of candidate strategies. The probability distribution on the network is used to describe the dynamic priority and activation tendency of each semantic defense strategy under the current network situation.

[0050] Each strategy This corresponds to a semantic-level defense approach, including but not limited to: Deception, Isolation, Monitoring, Resource Exhaustion, Attribution, and Interference.

[0051] During the decision-making stage, the following can be adopted: ; To obtain the current optimal single defense strategy; alternatively, it can be based on probability distribution. Multiple strategies are weighted and combined to form a composite defense scheme.

[0052] The set of policy distributions output by the module This will serve as the input for the subsequent "policy-weight mapping module," which uses an automatic mapping mechanism from semantic policy to reward parameters to convert the policy probability distribution into reward function parameters, action biases, and policy weights in reinforcement learning, thereby achieving closed-loop linkage from the cognitive decision-making layer to the execution control layer.

[0053] (iv) Automated transformation module based on policy-weight mapping operator This module is used to distribute the semantic defense strategies output by the decision-making layer. It is automatically converted into reward function parameters that can be directly called by the reinforcement learning execution layer, realizing a closed-loop mapping from high-level cognitive decision-making to low-level execution control.

[0054] This module uses the strategy distribution output in Part 3. As input, the output is the hyperparameter vector of the reward function. .

[0055] 1. Semantic Vectorization Representation of Defense Strategy Distribution Semantic Enhancement Threat Decoy Targets Corresponding defense strategy probability distribution middle, Represents the set of candidate defense strategies. Indicates the first Several candidate semantic defense strategies.

[0056] First, each semantic strategy is vectorized and encoded to obtain a semantic embedding vector. : ; This represents a semantic encoding function. This represents the dimension of the semantic embedding vector.

[0057] Based on this, the policy distributions are weighted and fused to obtain the overall semantic representation: ; in, Let represent the probability that the j-th defense strategy is adopted. This vector represents the total number of strategies in the set of defense strategies. It characterizes the comprehensive semantic preferences of multi-strategy games under the current security situation.

[0058] 2. Construction of Policy-Weight Mapping Operator Construction Strategy - Weight Mapping Operator This is used to map semantic vectors to reward function parameters: ; in, For the first The reward weight vector corresponding to each threat target; Corresponding to different reward dimensions (such as intelligence gains, asset losses, interaction time, etc.). This is a transpose operation.

[0059] In practical implementation, the mapping operator The following methods can be used: a multilayer perceptron (MLP) or a linear mapping matrix based on expert rules, combined with a historical policy library for parameter calibration to ensure the consistency of the mapping from semantics to control parameters.

[0060] 3. Dynamic Reconstruction Mechanism of Reward Function Mapped weight vector It is sent to the execution layer in real time to construct the dynamic reward function: ; in, Indicates the first A basic reward component function is used to characterize the current cybersecurity situation. Execute attack state actions The resulting multidimensional benefits or costs include: intelligence gain, asset loss, and exposure level.

[0061] This mechanism enables reinforcement learning agents to dynamically adjust their behavioral strategies based on high-level defense intentions, achieving a shift from "fixed script execution" to "adaptive game control".

[0062] 4. Semantic consistency verification and policy recalibration To prevent offsets during semantic mapping, a consistency check operator is introduced: ; in, The standard weight distribution in the expert strategy library, This is a safety threshold. When the verification fails to meet the constraints, a policy recalibration mechanism is triggered on the mapping operator. Alternatively, semantic vectors can be input for correction to ensure the stability and security of the system during adversarial processes.

[0063] 5. Module Output Definition This module ultimately outputs the set of hyperparameters for the reward function: ; Each of them The reinforcement learning control parameters correspond to a threat object and are used to drive the multi-agent game strategy update in the execution layer.

[0064] This output serves as the direct input to the subsequent execution layer (MAPPO adaptive game module), enabling an automated closed-loop mapping from semantic policy to behavioral control.

[0065] (v) Adaptive Interactive Game Theory Module Based on MAPPO Algorithm This module, as the execution layer, is responsible for driving multi-agent trapping nodes to engage in real-time interactive game with attackers under dynamic reward constraints, thereby achieving adaptive evolution of defense strategies.

[0066] This module uses the reward parameter set output by the strategy-weight mapping module. As input, the output is the multi-agent joint policy and its interaction execution trajectory.

[0067] 1. Multi-agent game environment modeling (MAPOMDP) The system models the network trapping environment as a multi-agent partially observable Markov decision process: ; in, This refers to the global state space (attack behavior, system state, etc.). For the first The action space of an intelligent agent For the first Local observations of an agent This is the state transition function. For weight vector Parameterized reward function.

[0068] for The defense agents of the trapping nodes, whose joint policy is defined as: ; Among them, joint strategy This indicates the selection of a joint action given a global state s. The conditional probability distribution is used to clarify that the global joint action probability is obtained by multiplying the independent policies of each agent; For joint operations, This represents the local observation of the i-th agent.

[0069] 2. MAPPO-based policy optimization mechanism To achieve collaborative trapping and dynamic defense control in a multi-agent environment, this embodiment uses the MAPPO (Multi-Agent Proximal Policy Optimization) algorithm to jointly train the defense agents.

[0070] Each defense agent executes corresponding defense actions based on local network observation information, while the centralized Critic network evaluates the value of joint actions based on the global network state, thereby achieving a collaborative optimization mechanism of centralized training and distributed execution.

[0071] Its strategy optimization objective function is: ; in, This represents the probability ratio between the old and new strategies; Represents the dominance function; This indicates that the strategy updates the truncation coefficient; This represents the strategy pruning function.

[0072] During training, the system combines the defense policy weights output by the RAG enhanced intent analysis module and the policy-weight dynamic mapping module to process the semantic defense policy set. The corresponding multiple behavioral orientations (such as Deception, Isolation, Monitoring, Resource Exhaustion, Attribution, and Interference) are jointly optimized, and the MAPPO algorithm is used to enable multiple defense agents to form a collaborative response strategy based on the network attack situation.

[0073] After strategy optimization, the system generates a set of multi-agent joint defense strategies: ; in, Indicates the first Local observation state of an Agent; Indicates the corresponding defensive action; This represents the agent's action decision-making strategy. The total number of intelligent agents.

[0074] The generated joint defense strategy will serve as input to the adaptive interactive game module, driving each decoy node to execute specific defense actions, including honeypot deployment, access control adjustment, traffic redirection, dynamic isolation, and attack path induction, achieving adaptive and collaborative threat decoying in the network environment. In this module, multiple defense agents autonomously select actions based on local observation states and jointly optimize strategies using the MAPPO algorithm, enabling the decoy nodes to evolve collaboratively and form a dynamic, adaptive defense system.

[0075] 3. Semantic-driven dynamic reward function injection The reward function of the execution layer is the reward function hyperparameter vector output by the policy-weight dynamic mapping module. Dynamic parameterization. Among them, The semantic defense strategy is automatically generated using a policy-weight mapping operator, which adjusts the importance of different reward dimensions in the policy optimization process. To accurately guide attacker behavior, this system constructs a multi-dimensional reward function that includes retention time, asset loss, and intelligence value. ; The reward components include: Duration, which measures how long an attacker remains in the trapping environment, representing the current cybersecurity posture. Execute the action of the agent in attack state. The cumulative effect of the duration of the post-attack behavior, with the corresponding reward objective being to maximize this metric; Damage (asset loss) is used to measure the current cybersecurity situation. Execute attack state actions The metric for the degree of simulated asset loss caused by the trapping environment is used to characterize the inhibitory effect of defense strategies on attack behavior. Its optimization objective is to minimize this loss (therefore, a negative sign is used in the reward function for modeling); Evidence (intelligence acquisition) is used to measure the current cybersecurity posture. Execute attack state actions The richness of the acquired threat intelligence, including attack payloads and TTPs (tactics, techniques, and processes), is optimized to maximize the benefits of this information; different threat targets correspond to different reward weight vectors. This is used to implement differentiated strategy control, thereby constructing a targeted induction and defense strategy control mechanism. For example, for high-value APT (Advanced Persistent Threat) attack scenarios, the system will adaptively improve... The weighting is adjusted to enhance the ability to capture information about the attack chain and toolchain, thereby inducing attackers to expose a more complete behavioral trajectory.

[0076] 4. Actively induced adaptive adversarial evolution mechanism During continuous interaction, the system does not passively respond to attack behavior, but actively constructs an "induction space" based on reinforcement learning strategies to guide and manipulate the attacker's behavior path.

[0077] Specifically, the execution layer implements proactive defense through the following mechanisms: (1) Path Steering By dynamically adjusting the exposure strategies of each trapping node (such as faking vulnerabilities or opening specific services), attackers are proactively guided to migrate to pre-set high-value trapping areas, rather than entering real business systems.

[0078] (2) Behavior Manipulation Based on strategy By controlling the response rhythm, information feedback, and interaction details, attackers can continuously misjudge the authenticity of the environment, thereby extending their stay time and increasing the depth of their operations.

[0079] (3) Intelligence Maximization In the reward function Under the guidance of this, the intelligent agent is prioritized to induce attackers to expose their attack toolchains, exploitation methods, and lateral movement strategies, thereby achieving the collection of high-value threat intelligence.

[0080] Under the aforementioned mechanism, the strategy optimization objective can be expressed as: ; in, This represents the hyperparameter vector of the reward function. Under adjustment, the execution layer agent is at any time The total reward value obtained; The hyperparameter vector of the reward function is used to represent the importance weights of objectives such as retention time, asset loss, and intelligence value. To analyze the attack trajectory The intelligence value function extracted is used to quantify the threat intelligence gains generated by attack behavior; This is the intelligence weighting coefficient, used to adjust the degree of influence of intelligence value items on the overall reward.

[0081] This mechanism transforms the system from the traditional "detection-blocking" model to a proactive defense paradigm of "inducement-manipulation-exploitation".

[0082] 5. Module Output Definition Final output of this module: (1) Multi-agent joint strategy ; Used to guide the real-time response behavior of each trapping node. Indicates the first Each trapping node generates a probability distribution or deterministic mapping of the optimal action based on the current local observation state.

[0083] (2) Set of interactive execution trajectories The output contains a complete set of interaction trajectories of each agent during the adversarial process. Each trajectory segment A sequence of quadruplets:

[0084] in, This indicates the current cybersecurity situation; Indicates an action in an attack state; This represents an immediate reward scalar value, used to measure the current action in [the context of the action]. The defensive benefits generated below; This represents semantic context information, used to record semantic labels of the attack intent at the current moment, such as attack phase descriptions like "scanning," "privilege escalation," and "lateral movement." This semantic context information is then input as feedback to a deep parsing module based on RAG enhancement and intent analysis, used to correct subsequent attack intent identification results and threat assessment processes, thereby achieving a closed-loop collaborative optimization mechanism between the execution and decision-making layers.

[0085] This specific embodiment extracts traffic features from both packet and flow dimensions and constructs temporal representations. It then generates standardized threat context objects based on abnormal behavior, completes data caching and time alignment using an asynchronous buffer, and forms a complete threat context set. Threat information is parsed through semantic enhancement, and a probability distribution of defense strategies is output. This is then transformed into a reward weight vector using a policy-weight mapping operator. Simultaneously, the MAPPO algorithm is employed to drive autonomous iterative optimization by multiple agents under dynamic weight constraints, combining real-time attack-defense game theory to achieve threat trapping and proactive defense. This specific embodiment integrates the entire process of traffic perception, threat modeling, policy decision-making, and agent evolution. It can accurately identify network anomalies, dynamically adapt to attack-defense scenarios, effectively improve threat identification, analysis, and handling capabilities, achieve automated and adaptive network security operations, and strengthen the overall network protection level.

[0086] Example 2 This embodiment provides a network security operation system based on adaptive multi-agent threat trapping, including: The multi-granularity feature perception module is used to extract packet-level temporal features and flow-level statistical features of network service traffic to construct traffic temporal features; when the current service behavior is abnormal, a threat context object is generated based on the traffic temporal features. A shared blackboard state maintenance module is used to store the threat context objects using an indexed asynchronous buffer and to construct a threat context set based on the time-aligned threat context objects. The RAG-enhanced intent analysis module is used to semantically enhance the threat context set and calculate the probability distribution of defense strategies. The policy-weight dynamic mapping module is used to construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function. The adaptive interactive game module is used to drive the agent's adaptive evolution based on the reward function weight vector and the MAPPO algorithm under dynamic reward weight constraints. It achieves network security operation through real-time feedback game with the attacker.

[0087] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a network security operation method based on adaptive multi-agent threat trapping as described in Embodiment 1 above.

[0088] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the network security operation method based on adaptive multi-agent threat trapping as described in Embodiment 1 above.

[0089] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0090] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A network security operation method based on adaptive multi-agent threat decoy, characterized in that, include: Extract packet-level time-series features and flow-level statistical features from network service traffic to construct traffic time-series features; when current service behavior is abnormal, generate threat context objects based on traffic time-series features; The threat context objects are stored using an indexed asynchronous buffer, and a threat context set is constructed based on the time-aligned threat context objects; Semantic enhancement is performed on the threat context set, and the probability distribution of defense strategies is calculated; Construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function; Based on the reward function weight vector, the MAPPO algorithm is used to drive the agent's adaptive evolution under dynamic reward weight constraints, and achieve network security operation through real-time feedback game with the attacker.

2. The network security operation method based on adaptive multi-agent threat trapping as described in claim 1, characterized in that, The construction of traffic time-series features specifically includes: The message-level time-series features are encoded using a Bi-GRU model to obtain a window-level time-series representation vector. By fusing the window-level time-series representation vector with the flow-level statistical features, the flow time-series features are obtained.

3. The network security operation method based on adaptive multi-agent threat trapping as described in claim 1, characterized in that, The current process for determining abnormal business behavior is as follows: Obtain the baseline distribution of normal business behavior, and map the traffic time-series characteristics to a real-time distribution in the same behavior space; The traffic deviation is calculated based on the baseline distribution and the real-time distribution using KL divergence. If the traffic deviation is greater than the threshold, the current business behavior is determined to be abnormal.

4. A network security operation method based on adaptive multi-agent threat trapping as described in claim 1, characterized in that, The threat context object includes: traffic timing characteristics, anomaly trigger timestamp, source information, asset identifier, anomaly divergence score, and current handling status.

5. A network security operation method based on adaptive multi-agent threat trapping as described in claim 4, characterized in that, The semantic enhancement of the threat context set and the calculation of the probability distribution of defense strategies specifically include: Extract the threat context objects to be processed from the threat context set; Based on the traffic timing features and asset identifiers in the threat context object, a query vector is constructed, and the RAG technology is used to match the Top-K most relevant background knowledge in the vector knowledge base, including asset topology information, historical vulnerability records and threat intelligence. Based on the aforementioned background knowledge, semantic enhancement is performed on the threat context object to obtain an enhanced context object; A dynamic attack graph is constructed by using the asset identifiers in the enhanced context object as nodes and the attack path transfer relationships between assets extracted from the threat intelligence as edges; an alarm sequence is defined, and the probability of the alarm sequence belonging to a certain attack path is modeled based on the attacker's attack status action and the enhanced context object, so as to obtain the attack path probability evaluation result. The attack path probability assessment results are used as risk situation characteristics to update the current network security situation; The enhanced context object is mapped to a semantic representation vector, and based on the semantic representation vector and the updated current network security situation, the probability distribution of the defense strategy is calculated.

6. A network security operation method based on adaptive multi-agent threat trapping as described in claim 1, characterized in that, The constructed strategy-weight mapping operator transforms the probability distribution of the defense strategy into a hyperparameter vector of the reinforcement learning reward function, specifically including: The probability distribution of the defense strategy is vectorized to obtain an overall semantic representation; Construct a policy-weight mapping operator to map the overall semantic representation to a reward function weight vector; specifically: ; in, For the first The reward weight vector corresponding to each threat target; For policy-weight mapping operators; For overall semantic representation; Corresponding to different reward dimensions; This is a transpose operation.

7. A network security operation method based on adaptive multi-agent threat trapping as described in claim 1, characterized in that, The method, based on a reward function weight vector, employs the MAPPO algorithm to drive the agent's adaptive evolution under dynamic reward weight constraints. Through real-time feedback game with the attacker, it achieves network security operations, specifically including: Multiple trapping nodes in the network environment are modeled as defensive agents, and the entire network environment is modeled as a multi-agent partially observable Markov decision process. A joint policy of the defensive agents is defined. The reward function is obtained by parameterizing the reward function weight vector. The MAPPO algorithm is used to train a joint strategy for each defense agent. Specifically, the defense agents execute defense behaviors based on local observation states, including driving honeypot deployment, access control adjustment, traffic redirection, dynamic isolation, and attack path induction. After executing the actions, the environment state is updated, the current interaction reward is calculated according to the reward function, and the various behavioral orientations corresponding to the defense strategy are jointly optimized. After multiple rounds of training, a joint defense strategy that adapts to the evolution of the attack situation is formed. Based on a joint defense strategy, the dynamic evolution of the trapping environment is driven in real time, forming a closed-loop defense system that is continuously iterated and optimized, thereby achieving network security operations.

8. A network security operation system based on adaptive multi-agent threat trapping, characterized in that, include: The multi-granularity feature perception module is used to extract packet-level temporal features and flow-level statistical features of network service traffic to construct traffic temporal features; when the current service behavior is abnormal, a threat context object is generated based on the traffic temporal features. A shared blackboard state maintenance module is used to store the threat context objects using an indexed asynchronous buffer and to construct a threat context set based on the time-aligned threat context objects. The RAG-enhanced intent analysis module is used to semantically enhance the threat context set and calculate the probability distribution of defense strategies. The policy-weight dynamic mapping module is used to construct a policy-weight mapping operator to transform the probability distribution of the defense policy into a weight vector of the reinforcement learning reward function. The adaptive interactive game module is used to drive the agent's adaptive evolution based on the reward function weight vector and the MAPPO algorithm under dynamic reward weight constraints. It achieves network security operation through real-time feedback game with the attacker.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the network security operation method based on adaptive multi-agent threat trapping as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the network security operation method based on adaptive multi-agent threat trapping as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster control and navigation method based on MAPPO

    CN119248009A

  • Intelligent traffic vehicle behavior modeling method and system based on collaborative learning

    CN120096628A