Camouflage node deployment method and device, medium and electronic equipment
By employing a threat intelligence-driven method for deploying camouflaged nodes, combined with Bayesian attack graphs and generative AI optimization strategies, the limitations of existing camouflaged node configurations are addressed, achieving high realism and intelligent defense in complex attack scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-13
AI Technical Summary
In existing defense systems, the configuration of spoofed nodes relies on manual design or fixed templates, making them easy for attackers to detect. Furthermore, they are difficult to simulate real system behavior in long-term confrontations, resulting in insufficient credibility. Moreover, reinforcement learning does not generalize well in real-world scenarios and cannot adapt to changes in attacker behavior.
By acquiring threat intelligence and performing structured modeling, a Bayesian attack graph is constructed to identify key nodes and high-risk paths. Combined with attack and defense strategies, deployment intent is deduced to generate low-level fingerprint features and high-level interaction features. Generative AI and reinforcement learning are used to optimize the deployment of disguised nodes and form an adaptive defense strategy.
It enhances the realism and intelligence of the masquerading nodes, enabling them to quickly adapt to environmental changes in complex attack scenarios, improve the realism and diversity of the defense system, and effectively confuse attackers and acquire intelligence.
Smart Images

Figure CN121664477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a method, apparatus, medium and electronic device for deploying spoofed nodes. Background Technology
[0002] APT attacks are typically characterized by long-term latency, continuous confrontation, and multi-stage penetration. Their attack process often includes reconnaissance, exploitation, lateral movement, and data leakage. To improve defense capabilities, a common method is to deploy spoofed nodes or mimic resources in the network. These simulate normal system, service, or communication behavior to confuse attackers, thereby delaying their attack process and providing defenders with attack paths and strategy intelligence. In practical applications, the generation method and deployment strategy of spoofed nodes directly affect the authenticity and coverage of the deception system.
[0003] In existing defense systems, the configuration of masquerading nodes mainly relies on manual design or fixed templates. Typically, defenders pre-define a set of masquerading service characteristics, such as protocol responses, port openness, or system log styles. These nodes can attract attackers to some extent, but they often reveal limitations in long-term confrontations. Fixed configurations are easily detected by attackers and are difficult to simulate the behavioral characteristics of real systems under different times and conditions, resulting in insufficient credibility of masquerading nodes during deep interaction phases.
[0004] Furthermore, while reinforcement learning has been applied to optimize defense strategies in existing technologies, the training process typically relies on idealized simulation environments. Real-world network environments are highly non-stationary, with attacker behavior patterns constantly evolving over time, and even interactive features not covered during training. This leads to insufficient generalization of trained strategies in real-world scenarios. Moreover, the lag in policy updates makes it difficult to maintain stability and effectiveness in long-term adversarial situations.
[0005] Therefore, it is necessary to provide a method for deploying camouflaged nodes that can improve the realism and intelligence of the defense system in complex APT attack scenarios. Summary of the Invention
[0006] The purpose of this invention is to provide a method, apparatus, medium, and electronic device for deploying spoofed nodes, which can adaptively generate spoofed nodes driven by threat intelligence and optimize deployment strategies to improve the realism, diversity, and intelligence of the defense system in complex attack scenarios.
[0007] In a first aspect, the masquerading node deployment method provided by the present invention includes: acquiring threat intelligence; performing structured modeling based on the threat intelligence; constructing a Bayesian attack graph based on the threat intelligence modeling results; identifying a set of key nodes and a set of high-risk paths based on the structured modeling results and the Bayesian attack graph; constructing an attack and defense strategy based on the set of key nodes, the set of high-risk paths, and the defense deployment status; deducing the deployment intent based on the strategies of both the attackers and defenders to obtain a candidate deployment set; generating corresponding low-level fingerprint features and high-level interaction features based on the candidate deployment set; performing consistency verification on the low-level fingerprint features and high-level interaction features to obtain a masquerading node candidate set that meets the consistency requirements; and performing deployment training based on the masquerading node candidate set, including: deploying masquerading nodes based on the threat intelligence modeling results and the masquerading node deployment status; performing optimization training based on the attack and defense environment formed by the deployment to obtain an optimal strategy function; selecting the optimal action based on the optimal strategy function; and updating the masquerading node deployment based on the optimal action.
[0008] The beneficial effects of the camouflage node deployment method provided by this invention are as follows: by combining threat intelligence with game modeling, it solves the problem of the disconnect between existing deployment schemes and actual APT attacks; it generates camouflage nodes with low-level fingerprint characteristics and high-level interaction characteristics, so that the camouflage nodes have higher realism in both external scanning and deep interaction processes; and it designs a closed-loop mechanism of intelligence modeling—game inference—camouflage node generation—training optimization—deployment, which gives the defense system the ability to quickly adapt to environmental changes and improves the realism, diversity and intelligence of the defense system in complex APT attack scenarios.
[0009] In one possible embodiment, structured modeling based on threat intelligence includes constructing a network topology model, calculating node risk weights based on the network topology model, defining an IOC mapping table, and establishing a threat event table; constructing a Bayesian attack graph based on the structured modeling results, where nodes in the Bayesian attack graph represent system assets or attack states, given evidence. ,node The probability of being compromised satisfies the following formula: ,in, Indicates the node to be evaluated. Represents nodes in the Bayesian attack graph The parent node, Represents a node The set of preceding nodes, This represents the set of observed evidence of intrusion. Indicates in the evidence The parent node is known below. When the node is captured The conditional probability of being further compromised; the comprehensive risk value of a node is calculated based on the probability of it being compromised; and the set of critical nodes and the set of high-risk paths are identified based on the node risk weight and the comprehensive risk value.
[0010] In another possible embodiment, a candidate deployment set is obtained by inferring the deployment intent based on the strategies of both the attacker and the defender, including: constructing the defender's strategy based on the deployment of the defender's disguised nodes; constructing the attacker's strategy based on the set of high-risk paths; defining the attacker's profit function and the defender's profit and budget constraints; and calculating the profit equilibrium solution for both the attacker and the defender to obtain the candidate deployment set.
[0011] In other possible embodiments, generating corresponding low-level fingerprint features and high-level interaction features based on the candidate deployment set includes: constructing a node condition vector based on the candidate deployment set and the deployment status of masquerading nodes; generating low-level node features based on the node condition vector and performing a simulated attacker scan on the generated features; outputting low-level fingerprint features when the discrimination result meets the set criteria; constructing structured prompts based on the low-level fingerprint features; and applying the structured prompts to interact with the large language model to generate high-level interaction features.
[0012] Based on the threat intelligence modeling results and the deployment status of masquerading nodes, masquerading nodes are deployed using a candidate set. The resulting attack and defense environment is then optimized and trained to obtain the optimal policy function. This process includes: defining a state space, a masquerading node deployment action space, and a reward function for each deployment action; deploying nodes using the candidate set to obtain updated states; defining a policy network and a value network for optimization training; collecting masquerading node deployment trajectory samples, including deployment actions, rewards, and state changes before and after deployment; updating the policy network using the PPO algorithm; and iteratively training the network according to a preset optimization objective. The optimal policy function is obtained when the iterative training based on the PPO algorithm converges.
[0013] Deploying and updating dummy nodes based on the optimal action includes: when the optimal action is to add a new node, selecting nodes from the candidate node set that meet the following conditions for deployment: ,in, This represents the utility value of a node in the intelligence gathering dimension. This represents the threshold value of a node's utility in the intelligence gathering dimension. Indicates the deployment consistency score. This represents the stability threshold; when the optimal action is to cancel the node or tune the node, the following budget constraints must be met: ,in, Indicates the cost of node deployment. This indicates the upper limit of the resource budget.
[0014] The process of selecting the optimal action based on the optimal policy function and updating the masquerading node deployment based on the optimal action is continuously iterated. Each iteration includes: collecting and updating security evidence based on the updated masquerading node deployment status, evaluating the updated security evidence using a Bayesian attack graph, updating the system state based on the evaluation results of the Bayesian attack graph and the system status, and selecting a new optimal action for masquerading node deployment based on the updated system state using the optimal policy function.
[0015] Secondly, the present invention also provides a disguised node deployment device, comprising: a threat intelligence processing unit, used to acquire threat intelligence, perform structured modeling based on the threat intelligence, construct a Bayesian attack graph based on the threat intelligence modeling results, and identify a set of key nodes and a set of high-risk paths based on the structured modeling results and the Bayesian attack graph; The spoofing node generation unit is used to construct attack and defense strategies based on the set of key nodes, the set of high-risk paths and the defense deployment situation. Based on the strategies of both the attacker and the defender, the unit deduces the deployment intent to obtain a candidate deployment set. Based on the candidate deployment set, the unit generates corresponding low-level fingerprint features and high-level interaction features. The unit performs consistency verification on the low-level fingerprint features and high-level interaction features to obtain a spoofing node candidate set that meets the consistency requirements. The optimized training unit is used for deployment training based on the candidate set of disguised nodes, including: deploying disguised nodes based on the candidate set of disguised nodes according to the threat intelligence modeling results and the deployment status of disguised nodes, and performing optimized training based on the attack and defense environment formed by the deployment to obtain the optimal policy function; The deployment execution unit is used to select the optimal action based on the optimal strategy function, and then deploy and update the masquerading node based on the optimal action.
[0016] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for deploying spoofed nodes.
[0017] Fourthly, the present invention also provides an electronic device, comprising: a processor and a memory; the memory being used to store a computer program; the processor being used to execute the computer program stored in the memory, so that the electronic device performs the above-described spoofing node deployment method.
[0018] For the beneficial effects of the second to fourth aspects mentioned above, please refer to the description of the first aspect mentioned above. Attached Figure Description
[0019] Figure 1 A flowchart illustrating a method for deploying spoofed nodes provided in an embodiment of the present invention; Figure 2This is a schematic diagram of a camouflage node deployment device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.
[0021] This embodiment provides a method, apparatus, medium, and electronic device for deploying masquerading nodes.
[0022] See the instruction manual appendix Figure 1 Methods for deploying spoofed nodes include: S101: Obtain threat intelligence, perform structured modeling based on the threat intelligence, construct a Bayesian attack graph based on the threat intelligence modeling results, and identify the set of key nodes and the set of high-risk paths based on the structured modeling results and the Bayesian attack graph.
[0023] In one possible embodiment, structured modeling based on threat intelligence includes constructing a network topology model, calculating node risk weights based on the network topology model, defining an IOC mapping table, and establishing a threat event table; constructing a Bayesian attack graph based on the structured modeling results, where nodes in the Bayesian attack graph represent system assets or attack states, given evidence. ,node The probability of being compromised satisfies the following formula: ,in, Indicates the node to be evaluated. Represents nodes in the Bayesian attack graph The parent node, Represents a node The set of preceding nodes, This represents the set of observed evidence of intrusion. Indicates in the evidence The parent node is known below. When the node is captured The conditional probability of being further compromised; the comprehensive risk value of a node is calculated based on the probability of it being compromised; and the set of critical nodes and the set of high-risk paths are identified based on the node risk weight and the comprehensive risk value.
[0024] In a specific embodiment, APT-related threat intelligence is structured and modeled to form risk inputs that can be used for subsequent deployment and optimization. This structured modeling includes constructing a network topology model, calculating node risk weights based on the network topology model, defining an IOC mapping table, and establishing a threat event table. Specifically, constructing the network topology model includes: denoting the network as a directed graph. , where nodes Represents assets, edges This indicates the connection relationship. Edge weights are weighted by communication frequency or data traffic to reflect the degree of dependency and interaction strength between nodes.
[0025] For each node Taking into account three dimensions—vulnerability severity, asset value, and network reachability—node risk weights are defined as follows: ,in, Represents a node The overall severity score of vulnerabilities is calculated. The score is taken from the general vulnerability scoring system of the National Information Security Vulnerability Database or the National Vulnerability Database; when a node has multiple vulnerabilities, a weighted average is used. This represents the value of a node asset, calculated based on the node's criticality in the business process, data sensitivity, and business impact, with a value range of [0,1]. This value can be obtained from the system security policy library or through annotation by human experts. Represents a node The reachability reflects the probability that an attacker can reach the node from an external or compromised node; it is calculated based on the number of topological paths, path length, and edge weights. , To attack along the border The probability of successful penetration. For all arriving nodes The set of paths. Assign coefficients to risk weights to satisfy the following conditions: .
[0026] Defining the IOC mapping table includes: mapping intrusion indicators (IOCs, such as IP, port, protocol mode, etc.) to nodes and links in the network topology, and defining the mapping relationship: ,in, This represents the set of IOCs extracted from threat intelligence. For a set of nodes, Let it be a set of edges. Mapping function. Used to identify the affected node or communication link corresponding to each IOC, thereby establishing the association between threat indicators and topology entities.
[0027] Establish a threat event table to store the tactics, techniques, and processes (TTPs) of APT attacks, in order to characterize the evolution path of attack behavior in the reconnaissance, exploitation, lateral movement, and infiltration phases.
[0028] Based on the structured modeling of threat intelligence, a Bayesian Attack Graph (BAG) is constructed. In the BAG, each node represents a system asset or attack state, given evidence. ,node The probability of being captured satisfies ,in, This represents the network node to be evaluated, i.e., the asset whose probability of being compromised is currently being calculated. Indicates a node in BAG The direct preceding node (parent node), i.e., the node that may be involved in the attack path. The upstream node that has been compromised; For nodes The set of direct parent nodes represents the possible preceding nodes on the attack path; This represents the set of observed intrusion evidence (including IOC events identified by threat intelligence or detection systems, system log alerts, etc.). Indicates in the evidence Below, the parent node is known. When the node is compromised, The conditional probability of further breach. Given evidence. It has the following characteristics: it is derived from real-time threat intelligence and monitoring data, and is incomplete and dynamic. That is, the evidence set may be updated over time or contain noise, and the system needs to dynamically adjust the attack probability estimate through Bayesian inference.
[0029] Through Bayesian inference, the probability of attacking each node can be dynamically calculated. Specifically, this includes constructing the attack graph structure and the prior attack probability of each node based on threat intelligence modeling results. Prior probability can be determined based on a combination of indicators such as the node's historical risk record, vulnerability severity (CVSS score), and asset importance, reflecting the initial threat level of the node in the absence of external evidence. When the monitoring system or threat intelligence source provides new evidence of an attack... When (such as intrusion indications, log anomalies, traffic anomalies, etc.), the system's conditional probability of related nodes. Update to reflect the latest threat dynamics. For each node. Based on its set of parent nodes Conditional probability propagation is performed, and the conditional probability propagation and iterative calculation satisfy the following formula: By propagating and iteratively calculating conditional probabilities, the dynamic correction of the probability distribution of node compromise across the entire graph is achieved. Based on the updated node probabilities, the comprehensive risk value of each node is calculated (combining factors such as node value, reachability, and vulnerability score), and the cumulative risk value is obtained for all attack paths. ,in, Indicates the node risk value. This represents the cumulative value of path risk. This represents the node risk weight coefficient.
[0030] By comparing node risk values and path risk accumulation values, nodes are divided into different protection priority zones, thereby identifying a set of critical nodes and a set of high-risk paths. The set of critical nodes... For each attack path, the set of nodes with the highest risk level is selected, and nodes in the critical node set are given priority to enter the masquerade node candidate deployment area; the high-risk path set... The set of paths whose cumulative risk value exceeds a set risk threshold is defined as the high-risk path set, and the area covered by the high-risk path set is defined as the high-risk deployment candidate area. The identification of the critical node set and the high-risk path set is used to ensure that subsequent game theory and reinforcement learning optimization are carried out within the high-risk domain, thereby improving the efficiency of deployment resource utilization and the effectiveness of the overall protection strategy.
[0031] To further clarify the criteria for defining protection priority areas, these areas can be further divided into three categories: high-risk areas, medium-risk areas, and low-risk areas.
[0032] Among them, the high-risk zone refers to the area where the node risk value or the cumulative path risk value is greater than the first threshold, which usually corresponds to critical business assets, core communication links, or sensitive network segments that have been repeatedly intruded upon; the medium-risk zone is the area where the node risk value or the cumulative path risk value is between the first threshold and the second threshold, which corresponds to peripheral business nodes or intermediate hop nodes; the low-risk zone is the area where both the node risk value and the cumulative path risk value are lower than the second threshold, which is usually the edge or security redundant nodes.
[0033] The criteria for dividing protection priority zones are based on the probability of nodes being compromised in the Bayesian attack graph. Cumulative value of path risk The joint indicators are determined to satisfy the following relationship: if If it is, then it is classified as a high-risk area; if If so, it is classified as a medium-risk area; if Those areas are classified as low-risk areas. and To dynamically adjust thresholds, they can be adaptively set based on historical threat intelligence, attack frequency, and defense resource allocation strategies. This tiered classification mechanism ensures that masquerading nodes prioritize coverage of high-risk areas during deployment, thereby optimizing protection benefits and resource utilization efficiency.
[0034] S102: Construct an attack and defense strategy based on the set of key nodes, the set of high-risk paths, and the defense deployment. Based on the strategies of both the attackers and defenders, infer the deployment intent to obtain a candidate deployment set. Generate corresponding low-level fingerprint features and high-level interaction features based on the candidate deployment set. Perform consistency verification on the low-level fingerprint features and high-level interaction features to obtain a candidate set of disguised nodes that meet the consistency requirements.
[0035] In one possible embodiment, a candidate deployment set is obtained by inferring the deployment intent based on the strategies of both the attacker and the defender, including: constructing a defender's strategy based on the defender's masquerading node deployment; constructing an attacker's strategy based on a set of high-risk paths; defining the attacker's profit function and the defender's profit and budget constraints; and calculating the profit equilibrium solution for both the attacker and the defender to obtain the candidate deployment set.
[0036] In a specific embodiment, Stackelberg game theory is applied to deduce the deployment intent (which nodes to deploy spoofed nodes). Specifically, the defense strategy is constructed based on the defense deployment situation. ,in, Indicates at node Deploy spoofed nodes, Indicates at node No spoofed nodes are deployed. The attacker observes the deployment scheme. Then, select the attack path. .gather Obtained through BAG posterior inference and risk screening, specifically, it refers to the simple path from the intrusion entry point to the set of critical nodes that satisfies a preset length constraint. Top-K high-risk paths, collection The path risk calculation for the attack path in the middle satisfies the following formula , As node risk weight, The probability of conquering the set. Will follow the evidence It is dynamically adjusted based on updates.
[0037] The attacker's payoff function is defined as follows: .in, For path The expected loss is obtained by weighting the nodal asset value with the posterior probability: , The value of node assets (given / assessed by factors such as business importance and data sensitivity); For deployment Attack path The deterrence / interference effect is obtained by summing the comprehensive suppression coefficients of the nodes deployed along the attack path: , Reflected in The expected reduction in attack benefits after deploying camouflage (including misleading, delay, increased probability of exposure, etc.) is estimated from red team / blue team adversarial data or offline simulation.
[0038] The defender's payoff function and budget constraint are defined as follows: .in, The defensive benefits gained from the deployment include both loss reduction and intelligence gains. For deployment costs, For the node Unit cost of deployment It is the budget ceiling. The calculation satisfies the following formula: , The residual expected loss after deployment, Indicates that the defending side is in the path The intelligence gain obtained by deploying disguised nodes (such as honeypots, traps, and decoy servers). For weights.
[0039] The equilibrium solution for the gains of both the attacker and defender is: , This represents the optimal solution for the attack strategy. The optimal solution for the defense strategy is represented by the output candidate deployment set. .
[0040] In one possible embodiment, generating corresponding low-level fingerprint features and high-level interaction features based on the candidate deployment set includes: constructing a node condition vector based on the candidate deployment set and the deployment status of masquerading nodes; generating low-level node features based on the node condition vector and performing a simulated attacker scan on the generated features; outputting the low-level fingerprint features when the discrimination result meets the set criteria; constructing structured prompts based on the low-level fingerprint features; and applying the structured prompts to interact with a large language model to generate high-level interaction features.
[0041] In one specific embodiment, after obtaining the candidate deployment set, generative AI is used to generate dummy node configurations to ensure the nodes possess both authenticity and diversity. Generating dummy node configurations using generative AI requires constructing a node condition vector as input based on the candidate deployment set and the dummy node deployment details. Specifically, the node condition vector is: ,in, Indicates the node risk level. Indicates topological location, Indicates the role of the node. This indicates a deployment recommendation. Topology location. The logical location and structural characteristics of a node in the network topology, derived from... (Such as boundary / convergence / core, one-hop / two-hop adjacency, betweenness centrality, network segment / subnet, etc.), encoded in the form of one-hot + continuous feature vectors; The deployment suggestion is to find the optimal solution output by the attacking and defending sides in a Stackelberg game. With candidate set .
[0042] For example, a candidate set of fake nodes is generated by combining a conditional generative adversarial network (cGAN) with a large language model (LLM). The conditional generative adversarial network is used to generate low-level fingerprint features, and the large language model is used to generate high-level interaction features. Then, consistency verification is performed to obtain the candidate set of fake nodes.
[0043] Conditional generative adversarial networks (GANs) use generators Low-level features are generated based on the node condition vectors, satisfying the following: ,in, This is a random noise vector used to introduce randomness and variability in the node feature generation process. It follows a standard normal distribution or a uniform distribution. This represents the dimension of the node feature space, defines the length of the generator's output feature vector, and controls the complexity and distinguishability of the masquerading node features. Low-level features cover port and service fingerprints, protocol distribution, handshake / timing, traffic intensity, TCP / IP metrics, banner / fingerprint fragments, etc. Discriminator take over Used for determining the authenticity of data, among which, Represents real node features (from a historical asset / traffic sample library) or generated features Output a simulation score. Calculate. ,in, This represents the authenticity evaluation index of generated nodes, used to quantify the discriminator's score for the authenticity of generated features. The value ranges from [0,1], with a larger value indicating a higher similarity between the generated sample and the real sample. The Sigmoid function is used to normalize the discriminator output to the [0,1] interval in order to generate a node authenticity index. It has probabilistic significance. The adversarial objective between the generator and the discriminator is defined as: ,in, This represents the expectation under the true node feature distribution, used to measure the discriminator's ability to judge true samples. This represents the expectation within the random noise space of the generator input, used to calculate the average discrimination result of the generated samples. The termination / qualification criterion for adversarial training is set to satisfy the following condition on the validation set. And the conditional distribution distance ,in, This represents the threshold for evaluating the authenticity of generated nodes, used to determine whether the generated features meet the authenticity requirements; once the adversarial training between the generator and the discriminator reaches the criterion, it will... As a qualified low-level fingerprint feature output.
[0044] exist Under the constraints, structured cue words are constructed, and high-level semantic features of nodes are generated through interaction between the structured cue words and the Large Language Model (LLM). , Specifically, this includes system log styles, request-response pairs, error codes and exception distributions, and business field templates, used to ensure the authenticity and consistency of interactions with attackers. Structured hints must include at least: role tags. With the target TTP scenario; with A set of services / applications with consistent ports / protocols; OS and log family templates (time granularity, field schema, trigger frequency); request-response examples and state transitions of interaction protocols with simple session scripts; resource / rate limiting / latency constraints consistent with traffic intensity; and context with topology / network segment (such as upstream / downstream objects, session source / destination).
[0045] To ensure consistency between low-level fingerprint features and high-level interaction features, a consistency check is performed, which specifically includes defining a consistency score. ,in, Indicates passage and The protocol / port consistency obtained from the match is recorded as either 0 / 1 or a similarity score; This indicates the use of cosine / edit distance / BLEU, etc. The log template / fields obtained by comparing with the target template are consistent; Indicates to Event frequency and Consistent evaluation of throughput / handshake timing yields consistent traffic intensity and latency. As the weight. When the consistency check result meets the pass condition. If the consistency check fails, the system obtains the required low-level fingerprint features and high-level interaction features; otherwise, it rolls back / regenerates the fingerprint. .
[0046] Define the interaction capability levels of generated nodes to adapt to different deployment costs and deception depths. For example, generated nodes can be divided into three interaction levels: low, medium, and high. Low-level interaction nodes include static fingerprints and limited handshakes, with reachable responses, stateless operation, low cost, and moderate realism; mid-level interaction nodes include stateful protocol sessions and simple business logs, capable of supporting basic probing and shallow exploitation attempts; high-level interaction nodes include end-to-end multi-turn interactions (LLM-driven), cross-session consistency, and task response, capable of long-term enticement and generating high-value intelligence. Node interaction levels are categorized as follows: , The selection process takes into account resource constraints and influences deployment costs during subsequent optimization. With deterrence coefficient .
[0047] By combining the generated low-level fingerprint features and high-level interaction features that meet the requirements, a candidate set of spoofed nodes is generated: Each element in the set contains low-level features, semantic features, interaction level, and two quality metrics. .
[0048] S103: Deployment training based on the candidate set of disguised nodes, including: deploying disguised nodes based on the candidate set of disguised nodes according to the threat intelligence modeling results and the deployment status of disguised nodes, and optimizing the training based on the attack and defense environment formed by the deployment to obtain the optimal policy function.
[0049] In one possible embodiment, based on the threat intelligence modeling results and the deployment status of masquerading nodes, masquerading nodes are deployed using a candidate set of masquerading nodes. The resulting attack and defense environment is then optimized and trained to obtain the optimal policy function. This includes: defining a state space based on the threat intelligence modeling results and the deployment status of masquerading nodes; defining a masquerading node deployment action space and a reward function for node deployment actions; deploying nodes based on the candidate set of masquerading nodes to obtain an updated state; defining a policy network and a value network for optimization training; collecting masquerading node deployment trajectory samples, including masquerading node deployment actions, rewards, and state changes before and after deployment; updating the policy network based on the PPO algorithm; and iteratively training the network according to a preset optimization objective. The optimal policy function is obtained when the iterative training based on the PPO algorithm converges.
[0050] For example, the deployment training of spoofed nodes using the Proximal Policy Optimization (PPO) reinforcement learning algorithm specifically includes: Define state space Specifically, it consists of threat intelligence modeling results, game theory simulation outputs, and the current deployment status of spoofed nodes. The state vector is defined as: .in, Represents the node risk level vector It is obtained by "risk measurement, weight calculation" and updated with the latest evidence; Characteristic representation of the Top-K high-risk path set (e.g., risk score of each path) (e.g., path length and one-hot key nodes), derived from the current assessment after Bayesian recalculation; Indicates the current camouflage deployment vector (1 indicates at node) (Disguise has been deployed; 0 indicates it has not been deployed or has been withdrawn). This represents network load characteristics (link / host utilization, concurrent sessions, average RTT, etc.) used to constrain generation / interaction levels and migration costs. Evidence Changes (such as new alerts / traffic observations) will trigger BAG recalculation. Thus update .
[0051] Define action space For the candidate set of spoofed nodes Choose to generate, cancel, or adjust masquerading nodes, where generating masquerading nodes is based on the masquerading node candidate set. Select elements and deployment ;Revoking a disguise node means revoking an existing disguise. Adjusting the disguise node to allow adjusting the interaction level without reversing the change; Or strategy parameters (simulating lightweight reconfiguration in real-world operations and maintenance). Actions are multi-step combinations, uniformly denoted as... .
[0052] Define reward function Time step Instant rewards are defined as ,in, As weight; Indicates the defensive effect, i.e., execution. Subsequently, the expected loss for the Top-K high-risk paths decreases, and its calculation satisfies... This is used to reflect the amount of suppression on the critical path before and after deployment; This represents the cost of deploying or migrating a masquerading node, and its calculation satisfies , Indicates at node The unit cost of deploying or maintaining masquerading nodes. This indicates the overhead of reconfiguring the fake node (such as image distribution, script restart, migration latency); Represents the intelligence gain, the calculation of which satisfies , Represents a node The intelligence value coefficient (captured samples / behavioral sequences, etc.). This represents the set of nodes affected by the action. The generation quality, i.e., the average simulation score of the candidate samples used for the action, is calculated according to the following conditions: , This value comes from the discriminator output when generating low-level fingerprint features. If no new features are added in this step, this value is 0.
[0053] After the action is updated New observations Trigger BAG recalculation And then update and ,form .
[0054] Define policy network Used according to state Select Action That is, to generate, revoke, or adjust spoofed nodes, in a value network. Network parameters are used to estimate the long-term reward of the current state. Initialize in a random manner and optimize step by step during the iteration process.
[0055] In a deployment environment consisting of real or simulated attack environments, the agent acts according to the current strategy. Conduct multiple rounds of interaction to collect trajectory sample sets. ,in, This represents the immediate reward based on a comprehensive assessment of defense benefits, intelligence value, and resource costs. The trajectory sample set is stored in an experience pool for subsequent strategy optimization and value estimation.
[0056] To reduce variance and improve estimation stability, generalized advantage estimation (GAE) is introduced. ,in, Indicates the discount factor. This represents the GAE coefficient. Advantage estimation reflects the degree of advantage of the current action relative to its expected value, and is used to guide policy optimization. Calculate probability ratio This is used to constrain the range of changes between the old and new policies, preventing excessive policy updates from causing training instability. The optimization objective adopts a truncated loss form: , This represents the pruning function, used to constrain the policy update ratio. The range of changes is adjusted to prevent the strategy from undergoing excessive updates during optimization. Conservative updates are achieved by taking the smaller term, allowing the strategy to improve returns while maintaining similarity to the old strategy.
[0057] Define the overall optimization objective ,in, Indicates time The target value, i.e., the supervision signal for training the value network, This means that the intelligent agent considers not only the direct rewards it currently receives, but also potential future gains. This represents the strategy entropy term, used to enhance exploratory behavior; Represents the weight coefficients. Updated via gradient descent (or ascent). Iterative optimization is performed across multiple epochs and mini-batch samples. When KL divergence constraints are in place: When triggered, the update is terminated early and the next round of sampling begins.
[0058] Policy network during policy update based on PPO algorithm With value network Alternating optimization, training stops when either of the following conditions is met: (1) Verification reward convergence: most recent The average return increase per epoch is less than tol. (2) KL constraint: continuous Second trigger. (3) Stability: Policy entropy Reduced to the threshold and deployment change rate continuous less than one epoch The optimal policy parameters are obtained after training stops. With value function parameters And the convergence strategy can be obtained as The optimal policy function is .in, ,tol and In this embodiment of the invention, is a hyperparameter used to control the convergence and stability criteria of the strategy optimization process. epoch represents the sliding window length, used to count the most recent... The average reward change trend over training cycles, typically ranging from 5 to 10; tol represents the average reward improvement tolerance threshold, when continuous When the average return improvement over each epoch is less than this threshold, the strategy is considered to have converged. The value is typically [value missing]. ; This represents the threshold for the number of times the KL divergence constraint is triggered, used to prevent excessive policy updates. When the KL divergence is continuous... This exceeds the maximum allowed value. If training is terminated early, Typically, the value is 2 to 3. The values of the above parameters can be determined through validation set experiments or empirical adjustments based on different network environments and training stability requirements, in order to achieve a balance between training speed and convergence accuracy.
[0059] S104: Select the optimal action based on the optimal strategy function, and deploy and update the masquerading node based on the optimal action.
[0060] In one possible embodiment, updating the masquerading node deployment based on the optimal action includes: when the optimal action is to add a new node, selecting nodes from the candidate node set that meet the following conditions for deployment: ,in, This represents the utility value of a node in the intelligence gathering dimension. This represents the threshold value of a node's utility in the intelligence gathering dimension. Indicates the deployment consistency score. This represents the stability threshold; when the optimal action is to cancel the node or tune the node, the following budget constraints must be met: ,in, Indicates the cost of node deployment. This indicates the upper limit of the resource budget.
[0061] In one possible embodiment, the process of selecting the optimal action based on the optimal policy function and updating the masquerading node deployment based on the optimal action is continuously iterated. Each iteration includes: collecting and updating security evidence based on the updated masquerading node deployment status, evaluating the updated security evidence using a Bayesian attack graph, updating the system state based on the evaluation result of the Bayesian attack graph and the system status, and selecting a new optimal action for masquerading node deployment based on the updated system state using the optimal policy function.
[0062] In a specific embodiment, based on the optimal policy function Real-time network status Making decisions and dynamically deploying and executing optimal defensive actions specifically includes: During the system operation phase, the state at each moment... The optimal action is selected by the optimal policy function. Optimal action This indicates the optimal deployment operation to be performed in the current state (e.g., adding, removing, or adjusting masquerading nodes). After the action is executed, the network state is updated accordingly. This enables the mapping of strategy output to actual deployment behavior, forming a closed loop of interaction between strategy and environment.
[0063] To ensure the stability of dynamic deployment and optimal execution under resource constraints: for operations involving new nodes, only nodes meeting the specified conditions are selected from the candidate set of masquerading nodes. The nodes, among which, This represents the utility value of a node in the intelligence gathering dimension. This represents the threshold value of a node's utility in the intelligence gathering dimension. Indicates the deployment consistency score. This represents the stability threshold; operations to undo or adjust nodes must satisfy system budget constraints. ,in, Indicates the cost of node deployment. This indicates the upper limit of the resource budget. By setting the above conditions, the system can maintain the optimal strategy effect during dynamic deployment, while ensuring the smoothness of defense layout adjustments and the controllability of resource usage.
[0064] The system continuously iterates through a cycle of "evidence update → BAG re-estimation → state update → policy execution" when deploying spoofed nodes. In each iteration, new security evidence is generated. Collected; Bayesian Attack Graph (BAG) re-estimates the system's vulnerable state based on evidence; state Updated to Policy network based on Select and execute the optimal action. This iterative loop enables online adaptive adjustment of the strategy, meaning the system can automatically update the optimal response strategy when the attack environment or network status changes, achieving dynamic optimal deployment and continuous defense optimization.
[0065] The camouflaged node deployment method provided by this invention combines threat intelligence with game theory modeling, solving the problem of existing deployment schemes being out of sync with actual APT attacks. All deployment inputs are derived from dynamically updated intelligence data and attack path deduction results, thus ensuring that deployment decisions are consistent with the attack posture.
[0066] Generative AI is introduced as a node generation method. GAN generates low-level network fingerprint features and LLM generates high-level interaction patterns, so that the fake nodes can act as real nodes in scanning and detection, and maintain simulation in deep interaction, thus avoiding the problem that traditional static templates are easily identified.
[0067] By employing the PPO reinforcement learning algorithm, deployment strategies can be continuously optimized in dynamic environments. Through real-time updates of the state space (risk distribution, topological relationships, node status), the strategy is improved after each round of interaction, thereby maintaining stability and convergence in long-term adversarial situations.
[0068] In the design of the reward mechanism, not only the success rate of defense and the cost of deployment resources are considered, but also two dimensions are added: intelligence gathering value and the quality of AI-generated nodes. Through this improvement, the optimization goal of reinforcement learning is no longer limited to simple defense, but takes into account attack delay, intelligence acquisition and node simulation, thereby improving the overall system's comprehensive defense capabilities.
[0069] By designing a closed-loop mechanism encompassing intelligence modeling, game theory simulation, AI generation, PPO optimization, and dynamic deployment, the defense system is endowed with the ability to rapidly adapt to environmental changes. When attack paths are unknown or mutate, the system can recover to a near-optimal deployment state within a short convergence period, significantly shortening the response time from threat detection to defense strategy updates, and improving the overall agility and real-time performance of the system. Simultaneously, this closed-loop mechanism ensures the continuous convergence and dynamic evolution of deployment strategies, enabling the defense system to maintain stable robustness in long-term confrontations.
[0070] See the instruction manual appendix Figure 2 This embodiment also provides a masquerading node deployment device, which is used to implement the above method embodiment. The device includes: The threat intelligence processing unit 201 is used to acquire threat intelligence, perform structured modeling based on the threat intelligence, construct a Bayesian attack graph based on the threat intelligence modeling results, and identify a set of key nodes and a set of high-risk paths based on the structured modeling results and the Bayesian attack graph.
[0071] The masquerade node generation unit 202 is used to construct an attack and defense strategy based on the set of key nodes, the set of high-risk paths and the defense deployment situation, and to deduce the deployment intention based on the strategies of both the attacker and the defender to obtain a candidate deployment set. Based on the candidate deployment set, corresponding low-level fingerprint features and high-level interaction features are generated, and the consistency of the low-level fingerprint features and high-level interaction features is verified to obtain a masquerade node candidate set that meets the consistency requirements.
[0072] The optimization training unit 203 is used for deployment training based on the candidate set of disguised nodes, including: deploying disguised nodes based on the candidate set of disguised nodes according to the threat intelligence modeling results and the deployment status of disguised nodes, and performing optimization training based on the attack and defense environment formed by the deployment to obtain the optimal policy function.
[0073] The deployment execution unit 204 is used to select the optimal action according to the optimal strategy function and to perform masquerade node deployment and update according to the optimal action.
[0074] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0075] In other embodiments of this application, an electronic device is disclosed, such as... Figure 3As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304. These devices can be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions that can be used to perform actions such as... Figure 1 And the steps in the corresponding embodiments.
[0076] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0077] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0078] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0079] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A method for deploying masquerading nodes, characterized in that, include: Acquire threat intelligence, perform structured modeling based on the threat intelligence, construct a Bayesian attack graph based on the threat intelligence modeling results, and identify the set of key nodes and the set of high-risk paths based on the structured modeling results and the Bayesian attack graph. An attack and defense strategy is constructed based on the set of key nodes, the set of high-risk paths, and the defense deployment. Based on the strategies of both the attackers and defenders, the deployment intent is deduced to obtain a candidate deployment set. Based on the candidate deployment set, corresponding low-level fingerprint features and high-level interaction features are generated. The consistency of the low-level fingerprint features and the high-level interaction features is verified to obtain a candidate set of disguised nodes that meet the consistency requirements. Deployment training based on a candidate set of disguised nodes includes: deploying disguised nodes based on the candidate set of disguised nodes according to the threat intelligence modeling results and the deployment status of disguised nodes; and performing optimization training based on the attack and defense environment formed by the deployment to obtain the optimal strategy function. The optimal action is selected based on the optimal strategy function, and the masquerading node is deployed and updated based on the optimal action.
2. The method according to claim 1, characterized in that, Structured modeling based on threat intelligence includes building a network topology model, calculating node risk weights based on the network topology model, defining an IOC mapping table, and establishing a threat event table. A Bayesian attack graph is constructed based on the structured modeling results. Nodes in the Bayesian attack graph represent system assets or attack states, given evidence. ,node The probability of being compromised satisfies the following formula: ,in, Indicates the node to be evaluated. Represents nodes in the Bayesian attack graph The parent node, Represents a node The set of preceding nodes, This represents the set of observed evidence of intrusion. Indicates in the evidence The parent node is known below. When the node is captured The conditional probability of being further captured; The comprehensive risk value of a node is calculated based on the probability of the node being compromised. The set of critical nodes and the set of high-risk paths are then identified based on the node risk weight and the comprehensive risk value.
3. The method according to claim 1, characterized in that, Based on the strategies of both the attackers and defenders, a set of candidate deployments is derived by deducing their deployment intentions, including: Develop defense strategies based on the deployment of masquerading nodes by the defender; Constructing attack strategies based on high-risk path sets; Define the attacker's payoff function and the defender's payoff and budget constraints; The set of candidate deployments is obtained by calculating the equilibrium solution of the gains for both the attacker and defender.
4. The method according to claim 1, characterized in that, Based on the candidate deployment set, corresponding low-level fingerprint features and high-level interaction features are generated, including: Construct a node condition vector based on the candidate deployment set and the deployment status of masquerading nodes; Low-level node features are generated based on node condition vectors, and the generated features are used to detect and distinguish them in the form of a simulated attacker scan. When the discrimination result meets the set criteria, the low-level fingerprint features are output. Structured prompts are constructed based on the low-level fingerprint features, and then the structured prompts are used in conjunction with a large language model to generate high-level interaction features.
5. The method according to claim 1, characterized in that, Based on the threat intelligence modeling results and the deployment of camouflaged nodes, camouflaged nodes are deployed using the candidate set of camouflaged nodes. The resulting attack and defense environment is then optimized and trained to obtain the optimal policy function, including: Define the state space, the action space for masquerading node deployment, and the reward function for node deployment actions based on the threat intelligence modeling results and the masquerading node deployment status. The updated state is obtained by deploying based on the set of masquerading node candidates; Define a policy network and a value network for optimization training, and collect camouflage node deployment trajectory samples, which include camouflage node deployment actions, rewards, and state changes before and after deployment. The policy network is updated based on the PPO algorithm, and iterative training is performed according to the preset optimization objective. The optimal policy function is obtained after the iterative training based on the PPO algorithm converges.
6. The method according to claim 1, characterized in that, Deploying and updating masquerading nodes based on optimal actions includes: When the optimal action is to add a new node, select a node from the candidate node set that meets the following conditions for deployment: ,in, This represents the utility value of a node in the intelligence gathering dimension. This represents the threshold value of a node's utility in the intelligence gathering dimension. Indicates the deployment consistency score. Indicates the stability threshold; When the optimal action is to cancel the node or to tune the node, the following budget constraints must be met: ,in, Indicates the cost of node deployment. This indicates the upper limit of the resource budget.
7. The method according to claim 1, characterized in that, The optimal action is selected based on the optimal strategy function, and the process of deploying and updating masquerading nodes based on the optimal action is continuously iterated. Each iteration includes: collecting and updating security evidence based on the updated deployment of masquerading nodes; evaluating the updated security evidence using a Bayesian attack graph; updating the system state based on the evaluation results of the Bayesian attack graph and the system status; and selecting a new optimal action for masquerading node deployment using the optimal policy function based on the updated system state.
8. A device for deploying camouflaged nodes, characterized in that, The device includes: The threat intelligence processing unit is used to acquire threat intelligence, perform structured modeling based on the threat intelligence, construct a Bayesian attack graph based on the threat intelligence modeling results, and identify a set of key nodes and a set of high-risk paths based on the structured modeling results and the Bayesian attack graph. The spoofing node generation unit is used to construct an attack and defense strategy based on the set of key nodes, the set of high-risk paths and the defense deployment situation, and to deduce the deployment intention based on the strategies of both the attacker and the defender to obtain a candidate deployment set. Based on the candidate deployment set, it generates corresponding low-level fingerprint features and high-level interaction features, and performs consistency verification on the low-level fingerprint features and the high-level interaction features to obtain a spoofing node candidate set that meets the consistency requirements. An optimized training unit is used for deployment training based on a candidate set of disguised nodes, including: deploying disguised nodes based on the candidate set of disguised nodes according to the threat intelligence modeling results and the deployment status of disguised nodes, and performing optimized training based on the attack and defense environment formed by the deployment to obtain the optimal policy function; The deployment execution unit is used to select the optimal action based on the optimal strategy function, and then deploy and update the masquerading node based on the optimal action.
9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the masquerading node deployment method as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory to cause the electronic device to perform the masquerading node deployment method according to any one of claims 1 to 7.