Array graph jumping method and system for cloud native honey array
By adopting the array graph jump method based on repeated Bayesian game model and deep Q network algorithm in the cloud native honey array, the problem that honey array cannot achieve the minimum defense cost and optimal defense effect at the same time in an incomplete information environment is solved, and the dynamic optimization and effectiveness of honey array strategy are achieved.
Patent Information
- Application Number
- CN202510677121.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing honey wave graph jump technology cannot meet the minimum defense cost and optimal defense effect at the same time when facing an incomplete information environment.
A formation map jump method for cloud-native honey array is adopted. By obtaining the attacker's action set and constructing the array map configuration set, the array map jump strategy is optimized based on the repeated Bayesian game model and deep Q network algorithm to achieve dynamic adaptability and effectiveness of the honey array.
By optimizing the array map jump strategy, the expected effectiveness of the honey array is maximized, the calculation efficiency is improved, the dynamic adaptability and effectiveness of the array map strategy is ensured, and the optimal deception defense effect is achieved.
Smart Images

Figure CN120200856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a matrix diagram jump method and system for cloud-native honey arrays. Background Art
[0002] The four-honey system is an active defense mechanism composed of four parts: honey points, honey courts, honey holes, and honey arrays, and is applied to the field of network security. Among them, the honey array, as the central control unit, is responsible for overall scheduling of the entire honeynet system and dynamically adjusts according to network attack behaviors to achieve the purpose of deception defense. Cloud-native applications build multi-dimensional configuration capabilities for the honey array diagram by decomposing application programs into multiple independent and loosely coupled technical components and combining different container runtimes or database management systems. However, in actual network attack and defense scenarios, the honey array often cannot obtain all the information of network attackers. How to formulate the optimal honey array diagram strategy under incomplete information is a key challenge for the honey array jump technology. In addition, matrix diagram jumping consumes certain resources. How to achieve the optimal deception defense effect while minimizing the jump cost is also a problem that the honey array jump technology needs to solve.
[0003] Currently, existing research has proposed a dynamic heterogeneous scheduling method. This method improves the overall security of the container cloud environment by constructing a heterogeneous image resource pool and periodically selecting heterogeneous containers for dynamic rotation and cleaning operations. However, this method cannot meet the minimum defense cost while taking into account the optimal deception defense effect in the face of an incomplete information environment. Therefore, there is an urgent need to provide a solution to improve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a matrix diagram jump method and system for cloud-native honey arrays, which can improve the problem that the existing honey array diagram jump cannot simultaneously meet the minimum defense cost and the optimal defense effect in the face of an incomplete information environment.
[0005] In a first aspect, a matrix diagram jump method for cloud-native honey arrays provided by the present invention includes: Obtaining an attacker action set and constructing a matrix diagram configuration set according to technical components in the cloud-native technology stack, and respectively constructing an attack set and a defense set based on a repeated Bayesian game model; Generating a matrix diagram jump strategy based on the attack set and the defense set; Optimizing the matrix diagram jump strategy based on the deep Q-network algorithm.
[0006] A method for diagram hopping in a cloud-native honeypot array provided by the present invention obtains an attacker action set from the national vulnerability database and the Common Vulnerabilities and Exposures, constructs a diagram configuration set according to technical components, then models the attack and defense process based on the repeated Bayesian game model, generates a diagram hopping strategy based on the diagram configuration set, defines a payoff set based on the Common Vulnerability Scoring System and constructs a honeypot equation based on the repeated Bayesian game model, and finally optimizes the diagram hopping strategy based on the Deep Q-Network algorithm, realizes the optimization problem of diagram strategies in a large-scale configuration space, maximizes the expected utility of the honeypot array, improves the computing efficiency, and ensures the dynamic adaptability and effectiveness of the diagram strategies.
[0007] Optionally, the technical components include a container runtime, a cloud-native network, a database management system, a container image repository, and an orchestration and scheduling system.
[0008] Optionally, the attacker action set includes attack behaviors against the container runtime, attack behaviors against the container image repository, attack behaviors against the network protocol stack, attack behaviors against the database management system, and attack behaviors against the orchestration and scheduling system.
[0009] Optionally, when constructing the diagram configuration set according to the technical components, it includes: independently selecting at least one component from each of the three types of technical components, namely the container runtime, the cloud-native network, and the database management system, to form different diagram configurations, and different ones of the diagram configurations finally constitute the diagram configuration set.
[0010] Optionally, the diagram hopping strategy is a decision-making process for dynamically adjusting the transfer from the current diagram configuration to the next valid diagram configuration based on a probability distribution.
[0011] Optionally, when optimizing the diagram hopping strategy based on the Deep Q-Network algorithm, it includes: quantifying the payoffs of the attack set and the defense set based on the general scoring system to generate the payoff set, where the payoff set includes attack payoffs and defense payoffs, constructing a honeypot target equation by combining the payoff set and based on the repeated Bayesian game model, and solving for the optimal diagram hopping strategy based on the Deep Q-Network algorithm. The solving process is as follows: Step 1, initialize the parameters of the attack and defense environment model, the experience replay buffer, the evaluation network parameters, and the target network parameters. The parameters of the attack and defense environment model include the diagram configuration set, the attack strategy, the defense strategy, the attack probability of the attacker, and the payoff set; Step 2, in each training cycle, initialize the environment and the state, and extract the attacker and its attack action set according to the prior probability; Step 3, select a diagram action based on the ε-greedy strategy, where the diagram action is randomly selected with probability ε, otherwise the Q-value action considered optimal by the current evaluation network is selected; Step 4: Calculate the corresponding rewards based on the repeated Bayesian game model and the payoff set, and store the training data in the experience replay buffer; Step 5: Randomly sample training samples from the experience replay buffer, calculate the Q-value of the current state based on the evaluation network, and calculate the target Q-value of the next state based on the target network; Step 6: Calculate the loss function and update the parameters of the evaluation network; Repeat steps 2 to 6 until the array pattern hopping strategy converges, and output the optimal array pattern hopping strategy.
[0012] In a second aspect, the present invention further provides an array pattern hopping system for a cloud-native honey array, including: A model construction module, which obtains the attacker action set, constructs an array pattern configuration set according to the technical components in the cloud-native technology stack, and constructs an attack set and a defense set respectively based on the repeated Bayesian game model; A strategy generation module, which generates an array pattern hopping strategy based on the attack set and the defense set; A strategy optimization module, which optimizes the array pattern hopping strategy based on the deep Q-network algorithm. Description of the Drawings
[0013] Figure 1 It is a flowchart of an array pattern hopping method for a cloud-native honey array provided by an embodiment of the present invention.
[0014] Figure 2 It is a structural diagram of an array pattern hopping system for a cloud-native honey array provided by an embodiment of the present invention. Detailed Embodiments
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art in the field to which the present invention belongs.
[0016] See Figure 1 , the present invention provides an array pattern hopping method for a cloud-native honey array, including the following steps: S1. Obtain the attacker action set, construct an array pattern configuration set according to the technical components in the cloud-native technology stack, and construct an attack set and a defense set respectively based on the repeated Bayesian game model; S2. generating a battle formation jump strategy based on the attack set and the defense set; S3. Optimizing the array jump strategy based on a deep Q network algorithm.
[0017] In fact, the array diagram jumping method provided by the present invention obtains the attacker action set according to the common vulnerability disclosure data recorded in the national vulnerability database and constructs the array diagram configuration set according to the technical components, and then models the attack and defense process based on the repeated Bayesian game model. At the same time, the array diagram jumping strategy is generated based on the array diagram configuration set, and then the benefit set is defined based on the common vulnerability scoring system and the honey array equation is constructed based on the repeated Bayesian game model. Finally, the array diagram jumping strategy is optimized based on the deep Q network algorithm. Therefore, the honey array can adjust the array diagram jumping strategy according to the latest information, realizing the array diagram strategy optimization problem in a large-scale configuration space and maximizing the expected utility purpose of the honey array.
[0018] In some embodiments, in step S1, the technical components may be container runtime, cloud native network, database management system, container image library, and orchestration and scheduling system; the attacker action set may be attack behaviors against container runtime, attack behaviors against container image library, attack behaviors against network protocol stack, attack behaviors against database management system, and attack behaviors against orchestration and scheduling system, that is, any behavior that attempts to undermine the confidentiality, integrity, or availability of these technical components may be regarded as an attack behavior. In addition, as a comprehensive vulnerability knowledge base, the National Vulnerability Database records in detail the known vulnerabilities and corresponding common vulnerability disclosures of each technical component in the cloud native technology stack. These records cover a wide range of information from vulnerability details to attack techniques that attackers may exploit, providing attackers with potential system penetration paths. Therefore, common vulnerability disclosures are used to simulate the attack actions that attackers may take in a cloud native environment.
[0019] In some embodiments, in step S1, the array configuration set is constructed by three types of technical components: container runtime, cloud native network, and database management system, where the corresponding technical component sets are , , , at least one component is independently selected from each set of technical components to form a valid array configuration, so the valid array configuration can be , ,..., , these valid array configurations constitute the array configuration set , the number of valid array configurations ,when hour, The maximum value is 27.
[0020] In some embodiments, in step S1, the Bayesian game model is repeated to define the defense strategy and the attack strategy, which are the defense set and the attack set ( ). Among them, the repeated Bayesian game model definition is a finite set of matrix diagrams of the defense set , indicating the jump from the current matrix diagram configuration to the next valid matrix diagram configuration. For example, represents the action of the defender choosing to switch to another valid matrix diagram configuration in the current state; the repeated Bayesian game model definition represents a finite set of actions of different attacker types , and the specific actions therein are represented by the disclosure of common vulnerabilities existing in the technical components; in addition, represents the probability that the attacker launches an attack ( ), where represents the th attacker type, and the attacker type is represented by a triple: , where the first element represents the attacker type name, the second element is a two-dimensional value representing the attacker's technical expertise, and its value range is [0, 10], and the third element is the probability that the attacker
[0021] launches an attack. The set of attacker types has a corresponding discrete probability distribution. Since these attacker types and their related discrete probability distributions are specific to a certain technical component, they can be defined by security experts in the cloud native field and used as the input of the text model. In some embodiments, in step S1, for the attacker action set, as long as the attack action affects at least one type of technical component in the matrix diagram configuration , then there is . Specifically, the decision function for determining whether the attack is successful is defined according to the following two points: , one is whether the technical components affected by the attack action ( ) in a specific attack and defense stage match at least one of the technical components used in the current matrix diagram configuration; the other is to compare the professional ability of the attacker type with the exploitable score of the corresponding common vulnerability disclosure ( ).
[0022] In some embodiments, in step S2, the matrix diagram jump strategy is that the honey matrix selects the next valid matrix diagram configuration in the given current matrix diagram configuration under the selection of the next valid matrix diagram configuration of the decision-making process (where ). For example, if represents the probability of selection under a given deployment state under the selection of , then the jump strategy is where ∑ c ∈ C p c =1 ∀ p c ∈ [0,1] .
[0023] In some embodiments, in step S3, when optimizing the matrix diagram jump strategy based on the deep Q-network algorithm, the payoffs of both the attacker and the defender in the repeated Bayesian game model are quantified using the impact score (IS), exploitability score (ES), and base score (BS) of the vulnerabilities defined in the latest vulnerability scoring system v3.1. Among them, the impact score IS represents the impact on the confidentiality , integrity , and availability of the system when the attacker takes a specific attack, and the calculation formula is IS = 6.42 * [1 - (1 - C) * (1 - I) * (1 - A)] ; the exploitability score ES represents the complexity of the attack vectors , attack complexity , privilege requirement , and user interaction in four aspects when exploiting a specific vulnerability. The specific calculation formula is: ; the base score BS is jointly composed of the impact score IS and the exploitability score ES. Finally, after setting the corresponding attack sets for each type of attacker, the payoff sets of both sides of the game can be defined based on the vulnerability scores of the vulnerability scoring system as follows: For the attack payoff , there are two cases for whether the attack is successful: one is when the attack is successful, the attack payoff is represented by the base score of the vulnerability scoring system; the other is when the attack is unsuccessful, the attack payoff is represented by the exploitability score of the vulnerability scoring system, representing the time and cost required for the attacker to carry out the attack. At this time, the attack payoff is negative. Therefore, the attack payoff can be expressed as: , where Indicates the attack action performed under the array configuration and the obtained attack benefit; It is used to judge whether the attack is successful. When it is the case, it represents a successful attack. When it is the case, it represents an unsuccessful attack.
[0024] Regarding the defense benefit there are two cases for whether the defense is successful: one is when the defense is unsuccessful, the defense benefit is represented by the impact score of the vulnerability scoring system which represents the impact on the system after the vulnerability is successfully exploited. At this time, the defense benefit is negative; the other is when the defense is successful, the defense benefit is set to zero, indicating that the honey array does not obtain a positive reward due to the successful defense. Therefore, the defense benefit can be expressed as: , where represents the defense benefit obtained by performing the attack action under the array configuration ; It is used to judge whether the defense is successful. When it is the case, it represents an unsuccessful defense. When it is the case, it represents a successful defense.
[0025] In addition, when the attacker takes no action (i.e., no operation), the benefit is zero.
[0026] In some embodiments, in step S3, when optimizing the array jump strategy based on the deep Q-network algorithm, a honey array objective equation is constructed based on the benefit set and the repeated Bayesian game model to maximize the expected utility of the honey array. The equation is as follows: , where represents the prior probability that the honey array believes different types of attackers are; represents the probability that the attacker chooses the attack action ; represents the mixed strategy of the honey array, is the attack and defense benefit of the honey array, is the cost coefficient, represents the jump cost when the honey array jumps from the current array jump configuration to another array jump configuration ; represents the probability corresponding to jumping from the current array jump configuration to another array jump configuration .
[0027] Among them, the jump cost of the array configuration can be represented by a matrix, where rows (and columns) represent system configurations, and the cells represent the jump cost when the honey array converts from the current array configuration to another array configuration, ranging from [0, 10]. The greater the heterogeneity between configurations, the higher the jump cost.
[0028] In some embodiments, in step S3, the deep Q-network algorithm is a deep reinforcement learning algorithm that solves the dimensional explosion problem in the traditional Q-learning algorithm by combining experience replay and fixed Q-value targets. The deep Q-network algorithm uses two key components: an evaluation network and a target network. Among them, the evaluation network is responsible for generating Q-value predictions for all possible actions in the current state, while the target network is used to stabilize the training process and assist in updating the parameters of the evaluation network by generating target Q-values. In the implementation process of this algorithm, the attack and defense environment and network parameters are first initialized, and then in each training cycle, the agent interacts with the environment to collect training samples. These training samples are then used to update the network parameters, and the loss function is optimized by the gradient descent method. The algorithm continues to iterate until a predetermined number of training steps are reached or the convergence condition is met, specifically including the following steps: Step 1, initialize the parameters of the attack and defense environment model, the experience replay buffer ( ), the parameters of the evaluation network and the parameters of the target network , where the parameters of the attack and defense environment model include the array configuration set, the attack strategy, the defense strategy, the attack probability of the attacker, and the benefit set; Step 2, in each training cycle, initialize the environment and state , and draw an attacker according to the prior probability and randomly draw the attack action set of this type of attacker; Step 3, select the array action based on the ε-greedy strategy, where the array action is randomly selected with probability ε , otherwise select , where represents the action with the maximum Q-value, represents using the parameters of the evaluation network in the state and action Step 4, calculate the corresponding reward based on the repeated Bayesian game model and the payoff set and store the training data ( ) in the experience replay buffer , where represents the current state, represents the next state; Step 5, randomly sample training samples from the experience replay buffer , and calculate the Q value of the current state based on the evaluation network and the target Q value of the next state based on the target network. The formulas used are as follows: Q D = ∑ t=1 ∞ [ γ t - 1 r D ( c t , d t , a t )| c 1 = c; θ] , , where represents the cumulative defense reward, represents the discount factor, represents the immediate defense reward, represents the initial state as , represents the evaluation network parameters; represents the target Q value, represents the immediate reward, represents the discount factor, represents using the target network parameters in the next state , defense action and attack action to calculate the target Q value, represents the target network parameters, represents the number of training samples drawn from the experience replay buffer each time; Step 6, calculate the loss function and use gradient descent to backpropagate the loss to update the evaluation network parameters , when the update iteration times of the evaluation network parameters reach times, assign the evaluation network parameters to the target network parameters , where represents the interval steps for synchronizing the target network parameters. The loss function used is as follows: L(θ) = E y i - Q(c, d i , a i ; θ) ) 2 ] , Among them, represents the mean square loss of the temporal difference error, represents the target Q value, represents using the evaluation network parameters in the state , defense action and attack action to calculate the evaluation network Q value, represents the target Q value, represents the number of training samples extracted each time from the experience replay buffer ; Repeat steps 2 to 6 until the array diagram jump strategy converges, and output the optimal array diagram jump strategy , where the number of experience steps required for the algorithm to converge is 100k.
[0029] See Figure 2 , the present invention provides an array diagram jump system for cloud-native honey arrays, including the following steps: Model construction module 100, obtaining the attacker action set and constructing an array diagram configuration set according to the technical components in the cloud-native technology stack, and respectively constructing an attack set and a defense set based on the repeated Bayesian game model; Strategy generation module 200, generating an array diagram jump strategy based on the attack set and the defense set; Strategy optimization module 300, optimizing the array diagram jump strategy based on the deep Q network algorithm.
[0030] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes all fall within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein can have other embodiments, and can be implemented or realized in various ways.
Claims
1. A method for diagram jump of a cloud-native honey array, characterized in that Including: Obtain the attacker's action set and construct an array configuration set according to the technical components in the cloud native technology stack, and respectively construct an attack set and a defense set based on the repeated Bayesian game model; Generate an array jump strategy based on the attack set and the defense set; Optimize the array jump strategy based on the deep Q-network algorithm.
2. The method for matrix pattern jump according to claim 1, wherein The technical components include a container runtime, a cloud native network, a database management system, a container image library, and an orchestration and scheduling system.
3. The method for matrix diagram jump according to claim 2, wherein The attacker's action set includes attack behaviors against the container runtime, attack behaviors against the container image library, attack behaviors against the network protocol stack, attack behaviors against the database management system, and attack behaviors against the orchestration and scheduling system.
4. The method for matrix pattern jump according to claim 2, wherein When constructing the array configuration set according to the technical components in the cloud native technology stack, it includes: independently selecting at least one component from each of the three types of technical components, namely the container runtime, the cloud native network, and the database management system, to form different array configurations, and different array configurations ultimately constitute the array configuration set.
5. The method for matrix diagram jump according to claim 1, wherein The array jump strategy is a decision-making process for dynamically adjusting the transfer from the current array configuration to the next valid array configuration based on a probability distribution.
6. The method for matrix diagram jump according to claim 1, wherein When optimizing the array jump strategy based on the deep Q-network algorithm, it includes: quantifying the benefits of the attack set and the defense set based on a general scoring system to generate a benefit set, the benefit set including attack benefits and defense benefits, combining the benefit set and constructing a honey array objective equation based on the repeated Bayesian game model, and solving the optimal array jump strategy based on the deep Q-network algorithm. The specific solution process is as follows: Step 1, initialize the parameters of the attack and defense environment model, the experience replay buffer, the evaluation network parameters, and the target network parameters. The parameters of the attack and defense environment model include the array configuration set, the attack strategy, the defense strategy, the attacker's attack probability, and the benefit set; Step 2, in each training cycle, initialize the environment and state, and extract the attacker and its attack action set according to the prior probability; Step 3, select an array action based on the ε-greedy strategy, where the array action is randomly selected with probability ε, otherwise the Q-value action considered optimal by the current evaluation network is selected; Step 4, calculate the corresponding reward based on the repeated Bayesian game model and the benefit set and store the training data in the experience replay buffer; Step 5, randomly sample training samples from the experience replay buffer, calculate the Q-value of the current state based on the evaluation network, and calculate the target Q-value of the next state based on the target network; Step 6, calculate the loss function and update the evaluation network parameters; Repeat steps 2 to 6 until the array jump strategy converges, and output the optimal array jump strategy.
7. A battle diagram jump system for cloud-native honey arrays, characterized in that, Including: A model construction module that obtains the attacker's action set, constructs an array configuration set according to technical components, and constructs an attack set and a defense set based on the repeated Bayesian game model; A strategy generation module that generates an array jump strategy based on the attack set and the defense set; A strategy optimization module that optimizes the array jump strategy based on the deep Q-network algorithm.
Citation Information
Patent Citations
ATT&CK-based spoofing defense system, construction method and full-link defense implementation method
CN110430190A
Attack tree honeypot deployment defense method and device based on deep reinforcement learning
CN115580430A
CPPS optimal defense strategy game method for uncertain attacks
CN117439794A
Internet-of-things deception trapping strategy evaluation method and system based on DDPG safety game
CN118400169A
Honey array graph generation method and system based on confusion attack graph
CN118631579A