Data honeypot confusion decision method and system based on policy gradient evolutionary game

CN122293425BActive Publication Date: 2026-08-07GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-05-11
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供基于策略梯度演化博弈的数据蜜点混淆决策方法及系统,改善了现有演化博弈方法中复制动态学习机制线性更新缺乏灵活性、策略探索能力受限及收敛速度慢的问题

Benefits of technology

[0010]本发明提供的一种基于策略梯度演化博弈的数据蜜点混淆决策方法的有益效果在于利用以学习机制为核心的演化博弈理论构建数据窃取与蜜点混淆的攻防演化博弈模型,在权衡蜜点混淆强度与通信成本的基础上建立攻防收益函数,并通过引入基于Soft-max策略梯度的演化博弈方法,利用归一化指数函数参数化策略空间并以随机梯度上升替代复制动态方程进行策略更新,改善了现有演化博弈方法中复制动态学习机制线性更新缺乏灵活性、策略探索能力受限及收敛速度慢的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293425B_ABST
    Figure CN122293425B_ABST
Patent Text Reader

Abstract

The application provides a data honeypot confusion decision method and system based on a policy gradient evolutionary game, and relates to the technical field of network security. The method provided by the application comprises the following steps: constructing an attack and defense evolutionary game model; performing parameterized representation on an attack strategy set and a honeypot confusion strategy set based on a normalized exponential function to obtain parameterized strategies; constructing an expected income function based on the parameterized strategies, and iteratively updating the strategy parameters by using a stochastic gradient ascent method; introducing a Jacobian matrix obtained by derivation of the normalized exponential function in the updating process, using diagonal elements of the Jacobian matrix to scale the gradient updating amplitude according to the size of the behavior selection probability, and using non-diagonal elements to make the updating directions of the behavior selection probabilities mutually exclusive; and when the strategy parameter updating amount is lower than a preset threshold, outputting a configuration vector with the maximum selection probability as an optimal honeypot confusion strategy. The evolutionary game method based on the Soft-max policy gradient is introduced, and effective optimization of the confusion strategy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a data honeypot confusion decision-making method and system based on policy gradient evolution game theory. Background Technology

[0002] With the development of information technology and the widespread application of data, data has become a core strategic resource. However, data theft is becoming increasingly serious, posing a severe threat to personal privacy, corporate assets, and even national security. Traditional passive security technologies based on perimeter defense and encryption have static characteristics, making them vulnerable to penetration and breach by attackers, and unable to cope with increasingly complex cybersecurity challenges. Therefore, data honeypot obfuscation technology, based on the concept of proactive defense, has emerged. It adds forged but realistic obfuscated data to real data, transmitting mixed data packets to interfere with attackers' perception of the real data, increasing their analysis costs and thus reducing the risk of data theft. However, the application of data honeypot obfuscation technology faces a core trade-off: if the proportion of obfuscated data is too small, too much real data is exposed, limiting the defense effect; if the proportion of obfuscated data is too large, a large amount of redundant data needs to be transmitted and processed, leading to increased communication latency and affecting system performance and normal business communication. Determining the optimal honeypot obfuscation strategy between data security and system communication performance has become an urgent problem to be solved.

[0003] Existing data honeypot obfuscation defense techniques mostly employ traditional game theory models or evolutionary game theory models for attack and defense decisions. Traditional game theory methods rely on static solutions to Nash equilibrium, failing to reflect the interactive learning and dynamic evolution of attack and defense strategies, making them ill-suited to complex and ever-changing network environments. While evolutionary game theory methods can depict the dynamic learning process of strategies, they generally employ a copy-based dynamic learning mechanism, which has three main drawbacks: First, the linear update mechanism lacks flexibility, relying on linear calculations of probability and payoff differences, resulting in low decision-making efficiency in high-dimensional policy spaces. Second, its strategy exploration capability is limited, focusing on utilizing existing game information driven by payoff differences, neglecting exploration of unknown strategies and easily getting trapped in local optima. Third, its convergence speed is relatively slow, requiring significant time for adjustments in dynamically changing and information-incomplete game scenarios, making it difficult to find the optimal honeypot obfuscation strategy in a timely manner, leading to severe loss of sensitive data. Therefore, there is an urgent need to develop a solution to address these problems. Summary of the Invention

[0004] The purpose of this invention is to provide a data honeypot confusion decision-making method and system based on policy gradient evolutionary game theory, which improves the problems of lack of flexibility in linear updates of the replication dynamic learning mechanism, limited policy exploration ability and slow convergence speed in existing evolutionary game theory methods.

[0005] The data honeypot confusion decision-making method and system based on policy gradient evolution game theory provided by this invention adopts the following technical solution:

[0006] Firstly, a data honeypot confusion decision-making method based on strategy gradient evolution game theory specifically includes:

[0007] An attack and defense evolution game model is constructed with attackers and defenders as participants. The attack strategy set is constructed based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is constructed based on the configuration vector of the proportion of obfuscated data configured for each port, and the defense payoff function is constructed based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port.

[0008] The attack strategy set and honeypot obfuscation strategy set are parameterized based on the normalized exponential function to obtain parameterized strategies;

[0009] The expected return function is constructed based on a parameterized strategy, and the strategy parameters are iteratively updated using the stochastic gradient ascent method. During the update process, a Jacobian matrix obtained by differentiating the normalized exponential function is introduced. The diagonal elements of the Jacobian matrix are used to scale the gradient update magnitude according to the magnitude of the behavior selection probability, and the off-diagonal elements are used to make the update directions of the behavior selection probabilities mutually exclusive. When the update amount of the strategy parameters is lower than a preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot confusion strategy.

[0010] The beneficial effect of the data honeypot confusion decision-making method based on policy gradient evolutionary game theory provided by this invention is that it uses evolutionary game theory with learning mechanism as the core to construct an attack and defense evolutionary game model of data theft and honeypot confusion. It establishes an attack and defense payoff function based on balancing the honeypot confusion intensity and communication cost. By introducing an evolutionary game method based on soft-max policy gradient, it uses a normalized exponential function to parameterize the policy space and uses stochastic gradient ascent to replace the copying dynamic equation for policy update. This improves the problems of lack of flexibility in linear update of copying dynamic learning mechanism, limited policy exploration ability and slow convergence speed in existing evolutionary game methods.

[0011] Optionally, the construction of the attack-defense evolutionary game model includes:

[0012] Define the game participants, which include the data-stealing end as the attacker and the data-protecting end as the defender;

[0013] The attack behavior set is defined as the set of port combinations that the attacker can choose to eavesdrop on, and the defense behavior set is the set of configuration vectors that allocate a certain proportion of obfuscated data to all ports of the defender.

[0014] The attack strategy set is defined as a probability distribution over the attack behavior set, and the honeypot obfuscation strategy set is defined as a probability distribution over the defense behavior set.

[0015] The attack payoff function is defined as the actual data value obtained by the attacker from each port in the selected port combination, while the defense payoff function is defined as the actual data value obtained by the attacker, the communication latency cost incurred by the attacker on the attacked port, and the communication latency cost incurred by the defender on the unattacked port.

[0016] Optionally, the attack strategy set and honeypot obfuscation strategy set are parameterized based on a normalized exponential function to obtain the parameterized strategy, which includes:

[0017] An attack strategy parameter vector is established, with each component corresponding one-to-one with each behavior of the attack strategy set. The normalized exponential function is used to map the attack strategy parameter vector into a parameterized attack strategy probability vector.

[0018] A defense strategy parameter vector is established, with each component corresponding one-to-one with each behavior of the honeypot obfuscation strategy set. The defense strategy parameter vector is then mapped to a parameterized honeypot obfuscation strategy probability vector using the normalized exponential function.

[0019] Optionally, when constructing the expected return function based on the parameterized strategy, the following is included:

[0020] The expected profit function includes an attack expected profit function and a defense expected profit function;

[0021] An attack expected payoff function is constructed based on the parameterized attack strategy probability vector, the attack payoff matrix, and the parameterized honeypot obfuscation strategy probability vector.

[0022] A defense expected benefit function is constructed based on the parameterized honeypot obfuscation strategy probability vector, the defense benefit matrix, and the parameterized attack strategy probability vector.

[0023] Optionally, the Jacobian matrix is ​​obtained by differentiating the normalized exponential function, and the values ​​of its diagonal elements are determined by multiplying the current action selection probability by its complement with respect to 1, while the values ​​of its off-diagonal elements are determined by multiplying the negative current action selection probability by the selection probabilities of other actions.

[0024] Optionally, when scaling the gradient update magnitude based on the probability of the behavior using the diagonal elements of the Jacobian matrix, the following steps are included:

[0025] By utilizing the diagonal elements of the Jacobian matrix, the gradient update magnitude corresponding to the action is reduced when the action selection probability approaches the upper limit, and the gradient update magnitude corresponding to the action is kept non-zero when the action selection probability approaches the lower limit.

[0026] Secondly, a data honeypot confusion decision-making system based on strategy gradient evolution game theory specifically includes:

[0027] The model building module is used to build an attack and defense evolution game model with attackers and defenders as participants. Among them, the attack strategy set is built based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is built based on the configuration vector of configuring the obfuscated data ratio for each port, and the defense benefit function is built based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port.

[0028] The strategy representation module is used to parameterize the attack strategy set and honeypot obfuscation strategy set based on the normalized exponential function to obtain parameterized strategies.

[0029] The policy solution module is used to construct the expected return function based on the parameterized policy and iteratively update the policy parameters using the stochastic gradient ascent method. During the update process, a Jacobian matrix obtained by differentiating the normalized exponential function is introduced. The diagonal elements of the Jacobian matrix are used to scale the gradient update magnitude according to the magnitude of the behavior selection probability, and the off-diagonal elements are used to make the update directions of the behavior selection probabilities mutually exclusive. When the policy update amount is lower than a preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot confusion policy.

[0030] The beneficial effects in the second aspect can be referred to the description in the first aspect.

[0031] Thirdly, the present invention also provides a storage medium that stores one or more programs that, when executed by a processor, implement the above-described data honeypot confusion decision-making method based on policy gradient evolution game.

[0032] Fourthly, the present invention also provides an electronic device, the electronic device comprising a memory and a processor, wherein:

[0033] The memory is used to store computer programs;

[0034] When the processor executes the computer program stored in the memory, it implements the above-mentioned data honeypot confusion decision-making method based on policy gradient evolution game. Attached Figure Description

[0035] Figure 1 A flowchart of a data honeypot confusion decision-making method based on policy gradient evolution game is provided in this embodiment of the invention;

[0036] Figure 2 This invention provides a structural diagram of a data honeypot confusion decision system based on policy gradient evolution game theory, as provided in an embodiment of the invention.

[0037] Figure 3 This is a flowchart of data theft and honeypot obfuscation countermeasure analysis provided in an embodiment of the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but does not exclude other elements or objects.

[0039] See Figure 1 This invention provides a data honeypot confusion decision-making method based on policy gradient evolution game, comprising the following steps:

[0040] S1. Construct an attack and defense evolution game model with attackers and defenders as participants; where the attack strategy set is constructed based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is constructed based on the configuration vector of configuring the proportion of obfuscated data for each port, and the defense payoff function is constructed based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port.

[0041] S2. Based on the normalized exponential function, the attack strategy set and the honeypot obfuscation strategy set are parameterized to obtain parameterized strategies.

[0042] S3. Construct the expected return function based on the parameterized strategy, and iteratively update the strategy parameters using the stochastic gradient ascent method. During the update process, introduce the Jacobian matrix obtained by differentiating the normalized exponential function. Use the diagonal elements of the Jacobian matrix to scale the gradient update magnitude according to the magnitude of the behavior selection probability, and use the off-diagonal elements to make the update directions of the behavior selection probabilities mutually exclusive. When the update amount of the strategy parameters is lower than the preset threshold, output the configuration vector with the highest selection probability as the optimal honeypot confusion strategy.

[0043] This invention provides a data honeypot obfuscation decision-making method based on strategy gradient evolution game theory, applied to server-side data security protection. Attackers exploit partially exposed ports on the server to launch attacks, establishing unauthorized access channels and attempting to hijack sensitive data forwarded by the server, such as business logs and core configurations. To counter this data theft, defenders employ data honeypot obfuscation technology for deception defense. The main idea is to add forged but realistic obfuscated data to the data packets used in communication between the user and the server. By controlling the proportion of obfuscated data in the forwarded data packets, the defender's honeypot obfuscation strategy space is formed. The server obfuscates the forwarded data through data honeypots, inserting obfuscated data into the real data. This creates a mixed data packet that is difficult to distinguish between real and fake data, transmitted through multiple ports. The output data stream may be intercepted by attackers; the data honeypot obfuscation technology reduces the loss of sensitive data. Alternatively, if not intercepted, the data flows to the user side, where users can verify the data through noise reduction to obtain normal data, supporting normal business operations while ensuring the security of sensitive data.

[0044] like Figure 3 This paper illustrates the process of countermeasures analysis against data theft and honeypot obfuscation. Attackers launch attacks using partially exposed ports of a server. During packet forwarding on the server's communication ports, they intercept data packets transmitted between the user and the server using passive sniffing or active hijacking, thereby obtaining sensitive information. To counter the attacker's data theft, defenders add varying proportions of obfuscated data to the data packets used in user-server communication. The obfuscated data in this invention refers to adding forged but realistic perturbation data to real data. By controlling the proportion of this obfuscation data, the defender's data honeypot obfuscation strategy space is formed. Ultimately, guided by the optimal data honeypot obfuscation strategy obtained through decision-making, the defender is assisted in transmitting data packets with a low probability of transmitting valid information on the attacked ports, thereby reducing the probability of the attacker obtaining valid user information.

[0045] In this offensive and defensive confrontation, attackers establish unauthorized access channels based on exposed server ports, attempting to steal forwarded sensitive data. Defenders, to disrupt the attackers' data theft and analysis, obfuscate the forwarded data using honeypots, inserting obfuscated data into the real data. The resulting mixed data stream may be intercepted by the attackers, but if not, it flows to the user. Users obtain normal data through noise reduction verification, thus supporting normal business operations. This interaction between the attackers and defenders constitutes the basic process of data theft and honeypot obfuscation confrontation, providing the technical scenario foundation for constructing the offensive-defense evolutionary game model of this invention.

[0046] In some embodiments, when constructing the attack-defense evolutionary game model with attackers and defenders as participants in step S1, the game participants are defined, including the data-stealing end as the attacker and the data-protecting end as the defender; the attack behavior set is defined as the set of port combinations that the attacker can choose to eavesdrop on, and the defense behavior set is the set of configuration vectors that allocate the proportion of obfuscated data to all ports of the defender; the attack strategy set is defined as the probability distribution on the attack behavior set, and the honeypot obfuscation strategy set is defined as the probability distribution on the defense behavior set; the attack payoff function is defined as constructed by the real data value obtained by the attacker from each port in the selected port combination, and the defense payoff function is constructed by the real data value obtained by the attacker, the communication latency cost borne by the attacker on the attacked port, and the communication latency cost borne by the defender on the unattacked port. Specifically, this includes:

[0047] Define game participants ,in This indicates an attacker, specifically the data-stealing entity. This indicates the defender, serving as the data protection endpoint;

[0048] Define attack behavior set , which is a set of port combinations that an attacker can choose to eavesdrop on. For the set of all possible attack ports, and .in, This indicates the total number of ports used by the server to forward data packets. This indicates the maximum number of ports that an attacker can exploit. Indicates from Select from the ports The total number of combinations of ports, The first one to be attacked Port combinations; define defense behavior sets This is a set of configuration vectors for the defender to configure the proportion of obfuscated data for all ports. One of these configuration vectors... , Indicates the first The percentage of obfuscated data in forwarded packets on each port; defining the behavioral sets of both attackers and defenders. ;

[0049] Define attack and defense strategy set , It is a set of behaviors The probability distribution on. This represents the set of attack strategies, which is a set of attack behaviors. The probability distribution on, where Indicates attacker's selection behavior The probability of; ( () represents the honeypot obfuscation strategy set, which is the defensive behavior set. The probability distribution on, where Indicates the defender's selection behavior The probability of;

[0050] Define attack and defense benefit functions . This represents the attack payoff function. This represents the defense payoff function, where the payoff for both sides is determined by the actions taken by the attacker. and defender selection behavior This occurs during a confrontation. The formula used for the attack and defense payoff function is as follows:

[0051] ,

[0052] in, Indicates the first The percentage of obfuscated data in forwarded data packets on each port; Indicates the first The percentage of real data in forwarded data packets on each port; Indicates the first The value of the data forwarded by each port; Indicates the first The unit communication latency cost caused by forwarding data through each port; This indicates that the attacker selected a combination of ports. And the defender selects the configuration vector The benefits of attacking at that time This indicates the defensive gains under the same confrontational circumstances;

[0053] The attack payoff matrix can be derived from the above payoff formula. From all possible offensive and defensive confrontation scenarios The matrix is ​​structured such that each row corresponds to one of the attacker's specific attack actions. The list corresponds to each defensive action of the defender. Defense Benefit Matrix Same reason Composition. The payout matrix is ​​represented as follows:

[0054] .

[0055] In some embodiments, in step S2, the attack strategy set and the honeypot obfuscation strategy set are parameterized based on the normalized exponential function to obtain the parameterized strategy. When obtaining the parameterized strategy, an attack strategy parameter vector is established, where each component corresponds one-to-one with each row of the attack strategy set. The normalized exponential function (Soft-Max function) is used to map the attack strategy parameter vector to a parameterized attack strategy probability vector. Similarly, a defense strategy parameter vector is established, where each component corresponds one-to-one with each row of the honeypot obfuscation strategy set. The normalized exponential function is used to map the defense strategy parameter vector to a parameterized honeypot obfuscation strategy probability vector. Specifically, this includes:

[0056] Establish attack strategy parameter vector Its components and attack behavior set Each attack behavior in the data corresponds one-to-one with each component. Indicates the corresponding attack behavior Parameter allocation values. Establish the defense strategy parameter vector. Each component and the honey spot confusion behavior set Each defensive behavior in the system corresponds one-to-one, and each component Indicates corresponding defensive behavior The parameter allocation value.

[0057] Using a normalized exponential function, the attack strategy parameter vector is... Mapped to a parameterized attack strategy probability vector Among them, the attack behavior The probability of selection is: ; the defense strategy parameter vector Mapped to a parameterized honeypot obfuscation strategy probability vector Among them, defensive behavior The probability of selection is: ;

[0058] in, express Attacker selects attack behavior at any time The probability of; express At any moment, the defender chooses a defensive action. The probability of; express Attack behavior at any time Parameterized probabilities; express Constant defensive behavior Parameterized probabilities; express Attack behavior at any time Corresponding attack strategy parameters; express Constant defensive behavior Corresponding defense strategy parameters; Indicates the total number of attack actions; Indicates the total number of defensive actions; This represents a temperature parameter used to balance strategy exploration and exploitation.

[0059] Through the above parameterized representation, the attack strategy set and the honeypot obfuscation strategy set are transformed into a continuous policy space described by finite-dimensional parameter vectors, where each possible attack or defense action is associated with an adjustable policy parameter. Updating the policy by optimizing the policy parameter allows for a better balance between exploration and exploitation during policy search: temperature parameter By controlling the sensitivity of the policy probability distribution to parameter changes, the policy probability adjusts smoothly rather than abruptly when parameters are updated, which helps to find the globally optimal policy and avoid premature convergence to a suboptimal policy. Simultaneously, transforming the policy space from a discrete form to a continuous form with finite parameters reduces the computational cost of optimizing policy parameters, and the stability of the policy function output is relatively good.

[0060] In some embodiments, in step S3, an expected return function is constructed based on a parameterized strategy, and the strategy parameters are iteratively updated using the stochastic gradient ascent method. During the update process, a Jacobian matrix obtained by differentiating the normalized exponential function is introduced. The diagonal elements of the Jacobian matrix are used to scale the gradient update magnitude according to the behavior selection probability, while the off-diagonal elements are used to make the update directions of each behavior selection probability mutually exclusive. When the policy parameter update amount is lower than a preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot obfuscation strategy. Specifically, this includes:

[0061] Based on the parameterized attack strategy probability vector obtained in step S2 And parameterized honeypot obfuscation strategy probability vector We construct the expected attack payoff function and the expected defense payoff function. The expected attack payoff function is determined by the parameterized attack strategy probability vector, the attack payoff matrix, and the parameterized honeypot obfuscation strategy probability vector. The expected defense payoff function is determined by the parameterized honeypot obfuscation strategy probability vector, the defense payoff matrix, and the parameterized attack strategy probability vector. The expected attack and defense payoff functions serve as the objective functions for updating the strategy parameters, characterizing the expected payoffs that both the attacker and defender can obtain under the current parameterized strategy. The formulas used are as follows:

[0062] ,

[0063] ,

[0064] in, express The expected profit of the attacker at any given moment; express The expected return of a constant defender; express The transpose of the parameterized attack strategy probability vector at time step; express The transpose of the parameterized honeypot confusion strategy probability vector at time step; express The parameterized attack strategy probability vector at time step; express The parameterized honeypot obfuscation strategy probability vector at time step.

[0065] To update the policy parameters, we calculate the gradients of the expected attack payoff with respect to the attack policy parameter vector, and the gradients of the expected defense payoff with respect to the defense policy parameter vector. Applying the chain rule, the formulas for calculating these gradients are as follows:

[0066] ,

[0067] in, express Expected return of a time-based attack with respect to the attack strategy parameter vector The gradient; express Expected return of defense at any time with respect to the defense strategy parameter vector The gradient; The Jacobian matrix represents the parameterized attack policy probability vector versus the attack policy parameter vector. The Jacobian matrix represents the ratio of the parameterized honeypot confusion strategy probability vector to the defense strategy parameter vector; Represents the attack payoff matrix; Represents the defense payoff matrix; This represents a parameterized attack strategy probability vector; This represents the probability vector of the parameterized honeypot obfuscation strategy. It completes the mapping from the strategy space to the parameter space, which is an endogenous structure for the autonomous evolution of offensive and defensive strategies.

[0068] Stochastic gradient ascent is used to update policy parameters. In the... At each update time, the attack policy parameter vector is updated along the positive direction of the gradient of the expected attack reward with a learning rate, and the defense policy parameter vector is updated along the positive direction of the gradient of the expected defense reward with the same learning rate. This dynamic change in parameters drives the adjustment of the policy probability. The learning rate is a preset positive real number used to control the magnitude of the policy parameter update in each iteration.

[0069] In the gradient calculation and parameter update process described above, the Jacobian matrix performs a nonlinear transformation on the gradient through its specific structure, which is manifested in the following two aspects:

[0070] First, there's the effect of gradient scaling. The diagonal elements of the Jacobian matrix are determined by multiplying the probability of selecting the current action by its complement with respect to 1, i.e. , When the probability of selecting a certain action approaches 1, the diagonal elements approach 0, significantly compressing the gradient component corresponding to that action. This minimizes the update magnitude of the policy probability, preventing the confused policy from converging prematurely and losing flexibility. Conversely, when the probability of selecting a certain action approaches 0, the diagonal elements also approach 0, but the gradient, though small, does not vanish, preserving non-zero update opportunities and maintaining the ability to explore potentially high-reward actions. This scaling factor... Sum and policy entropy The two strategies are directly proportional: when the strategy entropy is high, exploration and updates are faster, and the strategy tends to seek strategies with higher returns; when the strategy entropy is low, exploration and updates are slower, and the strategy focuses on utilizing strategies with higher current returns, thus achieving an autonomous balance between exploration and utilization.

[0071] Second, there is the inter-strategy coupling effect. The off-diagonal elements of the Jacobian matrix are determined by the product of the negative current action selection probability and the selection probabilities of other actions, i.e. When the probability of selecting a certain defensive action is increased through parameter updates, off-diagonal elements cause the probability of selecting other defensive actions to decrease accordingly. This inverse correlation mechanism balances the configuration of different strategies.

[0072] By combining the nonlinear transformations of the Jacobian matrix described above, the probability update of the attack and defense strategies is completed. The formula used is as follows:

[0073] ,

[0074] in, express The rate of change of the probability of selecting an attack behavior at any given moment; express The rate of change of the probability of choosing a defensive action at any given moment; Indicates the evolution rate parameter; Indicates temperature parameter; express The probability that an attacker will choose an attack behavior at any given moment; express The probability of the defender choosing a defensive action at any given time. Through the above update mechanism, the linear payoffs of attack and defense are transformed nonlinearly via the Jacobian matrix, enhancing the exploration capability of policy learning in the solution space.

[0075] After each parameter update, the policy parameter update amount is calculated, which is the change between the policy parameter vector at the current time and the policy parameter vector at the previous time. When the policy parameter update amount is lower than a preset threshold, the policy is determined to have converged. The configuration vector corresponding to the component with the highest selection probability in the current honeypot obfuscation policy probability vector is output as the optimal honeypot obfuscation policy. This optimal honeypot obfuscation policy provides defenders with a specific scheme for setting the proportion of obfuscated data on each data forwarding port.

[0076] See Figure 2 This invention provides a data honeypot confusion decision-making system based on strategy gradient evolution game, comprising the following steps:

[0077] The model building module 100 is used to build an attack and defense evolution game model with attackers and defenders as participants. Among them, the attack strategy set is built based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is built based on the configuration vector of configuring the obfuscated data ratio for each port, and the defense benefit function is built based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port.

[0078] The strategy representation module 200 is used to parameterize the attack strategy set and the honeypot obfuscation strategy set based on the normalized exponential function to obtain the parameterized strategy.

[0079] The policy solving module 300 is used to construct the expected return function based on the parameterized policy and iteratively update the policy parameters using the stochastic gradient ascent method. During the update process, the Jacobian matrix obtained by differentiating the normalized exponential function is introduced. The diagonal elements of the Jacobian matrix are used to scale the gradient update magnitude according to the magnitude of the behavior selection probability, and the off-diagonal elements are used to make the update directions of the behavior selection probabilities mutually exclusive. When the policy update amount is lower than a preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot confusion policy.

[0080] In another aspect, the present invention provides a readable storage medium storing a computer program that can be executed by a processor of the device in which the storage medium is located, to implement the data honeypot confusion decision-making method based on policy gradient evolution game as described in any of the above claims.

[0081] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0082] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0083] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0084] While embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations fall within the scope and spirit of the invention as set forth in the claims. Furthermore, the invention described herein may have other embodiments and can be implemented or carried out in various ways.

Claims

1. A data honeypot confusion decision-making method based on policy gradient evolution game, characterized in that, include: An attack and defense evolution game model is constructed with attackers and defenders as participants. The attack strategy set is constructed based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is constructed based on the configuration vector of the proportion of obfuscated data configured for each port, and the defense payoff function is constructed based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port. The attack strategy set and the honeypot obfuscation strategy set are parameterized based on the normalized exponential function to obtain parameterized strategies. This includes: establishing an attack strategy parameter vector, where each component corresponds one-to-one with each behavior of the attack strategy set; and using the normalized exponential function to map the attack strategy parameter vector into a parameterized attack strategy probability vector. The defense strategy parameter vector is also established, where each component corresponds one-to-one with each behavior of the honeypot obfuscation strategy set; and using the normalized exponential function to map the defense strategy parameter vector into a parameterized honeypot obfuscation strategy probability vector. The expected return function is constructed based on a parameterized strategy. The strategy parameters are iteratively updated using the stochastic gradient ascent method, including: calculating the gradient of the expected attack return with respect to the attack strategy parameter vector, and the gradient of the expected defense return with respect to the defense strategy parameter vector, respectively, applying the chain rule. The gradient calculation formula is as follows: ,in, express Expected return of a time-based attack with respect to the attack strategy parameter vector The gradient; express Expected return of defense at any time with respect to the defense strategy parameter vector The gradient; Represents the probability vector of parameterized attack strategy Attack strategy parameter vector Jacobian matrix; Represents the parameterized honeypot obfuscation strategy probability vector Defense strategy parameter vector Jacobian matrix; Represents the attack payoff matrix; Let the defense payoff matrix be represented. During the update process, a Jacobian matrix obtained by differentiating the normalized exponential function is introduced. Using the diagonal elements of the Jacobian matrix, the gradient update magnitude is scaled according to the behavior selection probability. The off-diagonal elements are used to ensure that the update directions of each behavior selection probability are mutually exclusive. This includes: using the diagonal elements of the Jacobian matrix to decrease the gradient update magnitude corresponding to the behavior when the behavior selection probability approaches the upper limit, and to retain the gradient update magnitude corresponding to the behavior as non-zero when the behavior selection probability approaches the lower limit. The update of the behavior selection probability is completed by combining the nonlinear transformation of the Jacobian matrix, as shown in the following formula: ,in, express The rate of change of the probability of selecting an attack behavior at any given moment; express The rate of change of the probability of choosing a defensive action at any given moment; Indicates the evolution rate parameter; Indicates temperature parameter; express The probability that an attacker will choose an attack behavior at any given moment; express The probability of the defender choosing a defensive action at any given time; when the policy parameter update amount is lower than the preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot obfuscation strategy.

2. The method as described in claim 1, characterized in that, The construction of the offensive and defensive evolutionary game model includes: Define the game participants, which include the data-stealing end as the attacker and the data-protecting end as the defender; The attack behavior set is defined as the set of port combinations that the attacker can choose to eavesdrop on, and the defense behavior set is the set of configuration vectors that allocate a certain proportion of obfuscated data to all ports of the defender. The attack strategy set is defined as a probability distribution over the attack behavior set, and the honeypot obfuscation strategy set is defined as a probability distribution over the defense behavior set. The attack payoff function is defined as the actual data value obtained by the attacker from each port in the selected port combination, while the defense payoff function is defined as the actual data value obtained by the attacker, the communication latency cost incurred by the attacker on the attacked port, and the communication latency cost incurred by the defender on the unattacked port.

3. The method as described in claim 1, characterized in that, When constructing the expected return function based on a parameterized strategy, the following are included: The expected profit function includes an attack expected profit function and a defense expected profit function; An attack expected payoff function is constructed based on the parameterized attack strategy probability vector, the attack payoff matrix, and the parameterized honeypot obfuscation strategy probability vector. A defense expected benefit function is constructed based on the parameterized honeypot obfuscation strategy probability vector, the defense benefit matrix, and the parameterized attack strategy probability vector.

4. The method as described in claim 1, characterized in that, The Jacobian matrix is ​​obtained by differentiating the normalized exponential function. The values ​​of its diagonal elements are determined by multiplying the current action selection probability by its complement with respect to 1, and the values ​​of its off-diagonal elements are determined by multiplying the negative current action selection probability by the selection probabilities of other actions.

5. A data honeypot confusion decision-making system based on strategy gradient evolution game, characterized in that, include: The model building module is used to build an attack and defense evolution game model with attackers and defenders as participants. Among them, the attack strategy set is built based on the combination of eavesdroppable ports, the honeypot obfuscation strategy set is built based on the configuration vector of configuring the obfuscated data ratio for each port, and the defense benefit function is built based on the value of the stolen real data, the communication latency cost of the attacked port, and the communication latency cost of the unattacked port. The strategy representation module is used to parameterize the attack strategy set and the honeypot obfuscation strategy set based on the normalized exponential function to obtain parameterized strategies. This includes: establishing an attack strategy parameter vector, where each component corresponds one-to-one with each behavior of the attack strategy set; and mapping the attack strategy parameter vector to a parameterized attack strategy probability vector using the normalized exponential function. It also includes establishing a defense strategy parameter vector, where each component corresponds one-to-one with each behavior of the honeypot obfuscation strategy set; and mapping the defense strategy parameter vector to a parameterized honeypot obfuscation strategy probability vector using the normalized exponential function. The strategy solving module is used to construct the expected reward function based on the parameterized strategy and iteratively update the strategy parameters using the stochastic gradient ascent method. This includes calculating the gradient of the expected attack reward with respect to the attack strategy parameter vector, and the gradient of the expected defense reward with respect to the defense strategy parameter vector, respectively, applying the chain rule. The gradient calculation formula is as follows: ,in, express Expected return of a time-based attack with respect to the attack strategy parameter vector The gradient; express Expected return of defense at any time with respect to the defense strategy parameter vector The gradient; Represents the probability vector of parameterized attack strategy Attack strategy parameter vector The Jacobian matrix; Represents the parameterized honeypot obfuscation strategy probability vector Defense strategy parameter vector The Jacobian matrix; Represents the attack payoff matrix; Let the defense payoff matrix be represented. During the update process, a Jacobian matrix obtained by differentiating the normalized exponential function is introduced. Using the diagonal elements of the Jacobian matrix, the gradient update magnitude is scaled according to the behavior selection probability. The off-diagonal elements are used to ensure that the update directions of each behavior selection probability are mutually exclusive. This includes: using the diagonal elements of the Jacobian matrix to decrease the gradient update magnitude corresponding to the behavior when the behavior selection probability approaches the upper limit, and to retain the gradient update magnitude corresponding to the behavior as non-zero when the behavior selection probability approaches the lower limit. The update of the behavior selection probability is completed by combining the nonlinear transformation of the Jacobian matrix, as shown in the following formula: ,in, express The rate of change of the probability of selecting an attack behavior at any given moment; express The rate of change of the probability of selecting defensive action at any given moment; Indicates the evolution rate parameter; Indicates temperature parameter; express The probability that an attacker will choose an attack behavior at any given moment; express The probability of the defender choosing a defensive action at any given time; when the policy update amount is lower than a preset threshold, the configuration vector with the highest selection probability is output as the optimal honeypot obfuscation strategy.

6. A storage medium, characterized in that, The storage medium stores one or more programs that, when executed by a processor, implement the data honeypot confusion decision-making method based on policy gradient evolution game as described in any one of claims 1-4.

7. An electronic device comprising a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the data honeypot confusion decision-making method based on policy gradient evolution game as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Dynamic honey spot placing method and device

    CN117176452A

  • Honeypot attack and defense confrontation strategy prediction method and system based on time-delay evolutionary game

    CN117579497A