Honey point address mutation decision-making method and system and readable storage medium
By constructing a game model for offensive and defensive scenarios and using the two-person Q-learning algorithm, dynamically adjusting the honey point address, the problem of irrational strategy selection when facing rational attackers is solved, and a more efficient defense strategy is achieved.
Patent Information
- Application Number
- CN202510571201.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the face of rational attackers, the traditional static honey dot deployment scheme lacks effective policy choices, which leads to irrationality in the transformation of honey dot address and is unable to effectively deal with the attacker's intrusion strategy.
By constructing a game model of offensive and defense scenarios, defining the action set and return function of both offensive and defense sides, building a temporary utility function, and using the double-player Q-learning algorithm to train the state action matrix to obtain the defender's optimal strategy, thereby dynamically adjusting the honey point address.
The optimal honey point address transformation strategy in the case of rational attackers is realized, which improves the generalization and adaptability of the defense system, and ensures that the defense strategy can still get the optimal solution when facing the worst situation.
Smart Images

Figure CN120090879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security protection, and particularly to a honeypot address mutation decision method, system and readable storage medium. Background Art
[0002] A honeypot is a new type of deception defense technology different from traditional honeypots. It abandons the idea of actively attracting attackers to attack. It is mainly aimed at APT attackers and is designed and developed based on the core idea of "responding to the unknown with the unknown, responding to the hidden with the hidden, and responding to the threat with the threat". A honeypot can be dynamically adjusted from multiple dimensions such as the number of deployments, location, and simulation type through an intelligent decision algorithm, so as to disrupt the reconnaissance results of attackers.
[0003] An external network honeypot is a lightweight deception defense device designed specifically for the external network environment. It simulates the real services publicly available by enterprises on the external network and cooperates with honeypot domain names that have a certain degree of confusion but will not cause confusion to legitimate users.
[0004] In the traditional static honeypot defense system, in order to comprehensively protect the business system, the defender needs to deploy multiple honeypots in the network structure. The disadvantage of this defense method is that the defender needs to predict the hosts that the attacker may attack in advance and deploy a static honeypot network structure. In addition, a large number of honeypots need to be deployed in the network, which may increase the network load, and the defense system structures in different scenarios may vary, so the static honeypot deployment scheme is not the best method.
[0005] Traditional dynamic honeypot deployment schemes are usually applied to mobile target defense and achieve dynamic deployment by modifying the network location (such as the address) of the honeypot. However, in the changing honeypot deployment scheme, the strategy selection is crucial. At present, the dynamic transformation strategy based on random frequency lacks a rational strategy selection, while the transformation strategy based on unilateral optimization lacks consideration of the attacker's strategy. That is to say, when the attacker discovers the network system, it will select a specific target for intrusion. Usually, a rational attacker will choose a strategy that maximizes its own benefits and minimizes the defender's benefits to select the attack target. Summary of the Invention
[0006] The purpose of the present invention is to provide a honeypot address mutation decision method, system and readable storage medium, so as to propose a honeypot address active transformation strategy, enabling the defender to find the best honeypot address transformation strategy in the case of a rational attacker, thereby transforming the honeypot address to the best position.
[0007] In the first aspect, the present invention provides a honeypot address mutation decision method, and the method includes: Construct an attack-defense scenario network topology diagram and deploy honeypots to the attack-defense scenario network topology diagram; Construct a game model and define the action sets of both the attacker and the defender in the game model, so as to obtain the payoff function and cost function of both the attacker and the defender according to the action sets, and construct an instantaneous utility function according to the payoff function and the cost function; Construct and initialize a two-player Q-learning model according to the instantaneous utility function to obtain a state-action matrix, which is filled with multiple state-action values; Iteratively train the two-player Q-learning model to update the state-action matrix, and set the convergence index of the two-player Q-learning model, so as to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutate the selected honeypot address to the host address in the attack-defense scenario network topology diagram according to the optimal strategy.
[0008] Further, the steps of constructing the game model and defining the action sets of both the attacker and the defender in the game model include: Construct a game model according to the following formula: , where, represents the game model, represents the participants, represents the action set, represents the state set, represents the instantaneous utility; Define the action set of the defender according to the following formula: , Define the action set of the attacker according to the following: , Define the state sets of both the attacker and the defender according to the following formula: , where, means that in the current time slice, the defender mutates the honeypot address to the th host, means that the honeypot expands its own reconnaissance range in the current time slice, means that the attacker selects to intrude on the host in the current time slice.
[0009] Further, the steps of obtaining the payoff function and cost function of both the attacker and the defender according to the action sets, and constructing an instantaneous utility function according to the payoff function and the cost function include: Construct an instantaneous utility function according to the following formula: , where, Denote the instantaneous utility function of the defender at the state as, denote the revenue function of the defender as, denote the revenue function of the attacker
[0010] Furthermore, construct the revenue function of the defender according to the following formula: , Construct the revenue function of the attacker according to the following formula: , wherein, denote that the attacker selects the action as the attacker's strategy, and the defender selects the action as the state of the defender's strategy, and respectively denote the value of the host where the honeypot is located and the attacker's intrusion into host k, denote the security requirement coefficient of the current defense scenario, denote that at the state the number of attack records obtained by the honeypot, denote the attack ability of the attacker, denote the effective penetration information collected by the attacker at the current state ; Construct the cost function of the defender according to the following formula: , Construct the cost function of the attacker according to the following formula: , wherein, denote the communication overhead of the honeypot address mutation, denote the operation and maintenance cost of the honeypot under one mutation, denote the cost-related coefficient of the defender, denote the reconnaissance range of the defender's current honeypot, denote the cost-related coefficient when expanding the reconnaissance range of the honeypot, denote the attack cost-related coefficient of the attacker.
[0011] Furthermore, the steps of constructing and initializing the two-player Q-learning model according to the instantaneous utility function to obtain a state-action matrix filled with multiple state-action values include: Obtain the state-action values in the state-action matrix according to the instantaneous utility function: , wherein, represents the state-action value of the j-th row and k-th column at state ; Use each action in the defender's action set as the row name, each action in the attacker's action set as the column name, and fill all Q values into the state-action matrix according to the row name and the column name.
[0012] Furthermore, the step of iteratively training the double Q-learning model to update the state-action matrix includes: Define the training period and calculate the cumulative utility according to the instantaneous utility function: , wherein, represents the cumulative sum of instantaneous utilities from the starting state to the current state, L represents the state transition set from the current state to the starting state, D represents the set of defender strategies executed by the defender in the current training period, and A represents the set of attacker strategies executed by the attacker in the current training period; If the cumulative sum of instantaneous utilities of the current state is greater than or equal to the maximum cumulative utility threshold, or the cumulative sum of instantaneous utilities of the current state is less than or equal to the minimum cumulative utility threshold, then enter the termination state; With probability select a random action, and with probability select the current optimal action, satisfying , wherein, represents the change rate parameter of , and t represents the number of training iterations; Define a single double Q-learning training to obtain the target Q value corresponding to the current Q value in the current state, and the target Q value corresponds to the optimal action in the current state; Perform one Q value iteration update using the Bellman equation; When the training period reaches the termination state, enter the next training period to cyclically iterate the state-action matrix.
[0013] Furthermore, the step of setting the convergence index of the double Q-learning model to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutating the selected honeypot address to the host address in the attack-defense scenario network topology diagram includes: The convergence metrics include a preset number of training cycles and a preset minimum Q-value change threshold. If the number of training cycles reaches the preset number of training cycles, the training converges; Alternatively, in several training iterations, if the change in the Q-value is less than the preset minimum Q-value change threshold, the training converges; After training, the attacker selects the one with the lowest defender benefit , and the defender selects the one with the highest defender benefit according to the row where it is located , denotes the state the optimal strategy of the attacker under denotes the state the optimal strategy of the defender under; Mutate the honeypot address to the corresponding host address.
[0014] Furthermore, the defining of a single-time two-player Q-learning training to obtain the target Q-value in the current state corresponding to the current Q-value, the steps of the target Q-value corresponding to the optimal action in the current state include: Calculate the target Q-value according to the following formula: , wherein, denotes the new state entered after the attacker and defender select the original strategy in the current state, denotes the action selected by the attacker to minimize the defender's benefit, denotes the action selected by the defender to maximize the defender's benefit when the attacker selects .
[0015] In a second aspect, the present invention provides a honeypot address mutation decision system, the system comprising: A topology graph construction module, configured to construct an attack-defense scenario network topology graph and deploy honeypots to the attack-defense scenario network topology graph; A utility function construction module, configured to construct a game model, define the action sets of the attacker and defender in the game model, obtain the benefit function and cost function of the attacker and defender according to the action sets, and construct an instantaneous utility function according to the benefit function and the cost function; An action matrix generation module, configured to construct and initialize a two-player Q-learning model according to the instantaneous utility function to obtain a state-action matrix, and the state-action matrix is filled with a plurality of state-action values; The optimal strategy acquisition module is used to iteratively train the two-player Q-learning model to update the state-action matrix, and set the convergence index of the two-player Q-learning model, so as to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutate the selected honeypot address to the host address in the attack-defense scenario network topology map according to the optimal strategy.
[0016] In a third aspect, the present invention provides a readable storage medium storing one or more programs, which when executed by a processor implement the above-mentioned honeypot address mutation decision method.
[0017] In a fourth aspect, the present invention provides a computer device, which includes a memory and a processor, wherein: The memory is used to store a computer program; When the processor executes the computer program stored on the memory, it implements the above-mentioned honeypot address mutation decision method.
[0018] Compared with the prior art, the present invention has the following advantages: 1. By constructing a game model involving both the attacker and the defender, the present invention models the attack-defense problem of moving target defense and defines the game elements of dynamic honeypot address transformation. Through this model, complex network attack-defense scenarios can be effectively analyzed and processed, and it can also adapt to changes in the network environment and attacker strategies in real time, adjust the defender's strategy, and improve the generalization and adaptability of the defense system.
[0019] 2. By constructing a zero-sum game modeling utility function, the present invention comprehensively considers the benefits and costs of the defender and the attacker in the attack-defense confrontation scenario. This makes the defense strategy more reasonable economically, and can also link the opponent's benefit with its own benefit in the confrontation relationship, so as to more efficiently evaluate and optimize the defense strategy.
[0020] 3. The present invention establishes a Q-table (state-action matrix) through a three-dimensional table. The use of the three-dimensional table fully considers the strategies of both the defender and the attacker, enabling the expected discounted benefits of the attacker and the defender to be evaluated in detail in each state, and the three-dimensional table can more intuitively understand and analyze the effectiveness and update direction of the attack-defense strategy.
[0021] 4. The present invention uses the Minimax Q-learning algorithm to solve the optimal strategy of the defender in the game model. This algorithm ensures that the defender can obtain the optimal solution even in the worst case, thereby improving the robustness of the defense system in an environment with rapid changes. At the same time, the Minimax Q-learning method can accelerate the convergence speed of training the Q-table and make the training more efficient. Brief Description of the Drawings
[0022] Figure 1 It is a flowchart of the honeypot address mutation decision method proposed in an embodiment of the present invention; Figure 2 It is a schematic structural diagram of the honeypot address mutation decision system proposed in an embodiment of the present invention.
[0023] The following specific embodiments will further illustrate the present invention in conjunction with the above drawings. Specific Embodiments
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.
[0025] The applicant found that in mobile target defense, the network environment is dynamic and complex, and the strategies of attackers and defenders will change continuously. It is necessary to find the optimal defense strategy in a dynamic environment. The defender deploys honeypots to defend against the intrusion of attackers. To improve the efficiency of defense, the defender needs to dynamically modify the addresses of the honeypots to enhance the confusion of the honeypots to the attackers. The traditional address transformation strategies include random address transformation strategy and the address transformation strategy that only considers the defender unilaterally. However, the random address transformation strategy lacks rational consideration, and the address transformation strategy that only considers the defender unilaterally lacks the consideration of the attacker's strategy as a rational decision.
[0026] Based on this, the present invention proposes a honeypot address mutation decision method, enabling the defender to find the best honeypot address transformation strategy of the defender in the case of a rational attacker, so as to transform the honeypot address to the best position.
[0027] Please refer to Figure 1 , which shows a flowchart of a honeypot address mutation decision method provided by an embodiment of the present invention. The method includes steps S101 to S104, where: Step S101: Construct an attack-defense scenario network topology graph and deploy honeypots to the attack-defense scenario network topology graph; It should be noted that the network topology diagram of this attack and defense scenario is constructed by the hosts with independent addresses in the protected network system, that is , where represents the address set in the protected network.
[0028] In addition, it is also necessary to evaluate the value of the hosts participating in the construction of the topology diagram. Among them, the important data and key services owned by the host can measure the value of the host . Initially, the defender knows the value set of all hosts in the network , while the attacker needs to obtain the value information of the hosts through the intrusion process.
[0029] Specifically, during the process of deploying the honeypots, at initialization, randomly select the hosts in the topology diagram to deploy highly realistic honeypots, and observe the information collected by the honeypots within a time slice.
[0030] In addition, it is also necessary to determine the size of the time slice, that is, to determine the period of the honeypot address mutation. According to the performance and cost factors, determine the size of the time slice. In this embodiment, in each time slice, the defender will decide the address of the honeypot for address mutation.
[0031] Step S102: Construct a game model, and define the action sets of both the attacker and the defender in the game model, so as to obtain the revenue function and cost function of both the attacker and the defender according to the action sets, and construct an instantaneous utility function according to the revenue function and the cost function; In this step, the constructed game model is a Markov game model. In order to construct a Markov game model, it is also necessary to define a finite number of states composed of the network topology. The participants consist of a single defender and a single attacker. And it is necessary to define the defender's action set composed of address transformation and honeypot expansion and the attacker's action set composed of attacking hosts.
[0032] In addition, in order to construct the instantaneous utility function of the defender, the revenue and cost of the defender and the revenue and cost of the attacker involved in the confrontation scenario under the participation of both the attacker and the defender are comprehensively considered, so that when the defender selects the optimal strategy later, it is more reasonable economically and can evaluate and optimize the defense strategy more efficiently.
[0033] Step S103: Construct and initialize a two-player Q-learning model according to the instantaneous utility function to obtain a state-action matrix, and the state-action matrix is filled with multiple state-action values; It should be noted that in the process of establishing and initializing the two-player Q-learning model, a three-dimensional table needs to be used as the Q-table, that is, the state-action matrix is obtained, the Q-table is set, and the Q-table is initialized with the initial instantaneous utility.
[0034] Step S104: Iteratively train the two-player Q-learning model to update the state-action matrix, and set the convergence index of the two-player Q-learning model to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutate the selected honeypot address to the host address in the attack-defense scenario network topology map according to the optimal strategy.
[0035] It should be noted that in the iterative training process, it mainly involves defining the training cycle, using the algorithm for exploration and exploitation, selecting the value that minimizes the defender's gain and maximizes the defender's gain in the Q-table by the Minimax algorithm as the target Q-value, and cyclically updating the Q-table with the Bellman equation. After sufficient training, the Q-table converges, and the best strategy corresponding to each state can be known through the Q-table. Finally, the honeypot address mutation is performed through the optimal strategy.
[0036] In summary, according to the above-mentioned honeypot address mutation decision method, the optimal strategy of the defender can be found in the case of a rational attacker, so that the defender can determine the address that the honeypot needs to change according to the optimal strategy.
[0037] Furthermore, in some alternative embodiments, the specific steps of constructing the Markov game model include: (1) First, establish the current attack-defense problem as a Markov game , where represents the game model, represents the participants, represents the action set, represents the state set, represents the instantaneous utility; (2) Define the participants in the game model. The participants in the game include two agents, a single defender and a single attacker. In addition, it should be noted that in this embodiment, the attacker side is regarded as one agent, and the defender side is regarded as one agent. That is to say, even if the number of attackers increases, the attackers can be regarded as a single decision-making agent; (3) Define the action set in the game model. The action set of the game includes the action set of the defender and the action set of the attacker, the strategies of both sides adopt the discrete pure strategy form. Specifically, the action set of the defender is , and the action set of the attacker is , where represents that in the current time slice, the defender mutates the honeypot address to the th host, represents that the honeypot expands its own reconnaissance range in the current time slice, represents that the attacker selects to invade the host in the current time slice; (4) Define the state set . In the game, the state set depends on all strategy combinations selected by the attacker and the defender, and can be defined as . For a specific state, represents that the attacker selects the action as the attacker's strategy, and the defender selects the action as the defender's strategy.
[0038] Further, in some alternative embodiments, the process of constructing the instantaneous utility function based on this Markov game model is as follows: (1) For the defender, when it uses the honeypot to capture the attacker, the defender's benefit mainly comes from the value of the host and the obtained attack record information. Therefore, the defender's benefit function is constructed according to the following formula: , (2) For the attacker, its purpose is to invade the real system and avoid being captured by the honeypot as much as possible. The attacker prefers to penetrate into high-value hosts and collect as much effective penetration information as possible. Therefore, the attacker's benefit function is constructed according to the following formula: , where represents the state where the attacker selects the action as the attacker's strategy, and the defender selects the action as the defender's strategy, and respectively represent the values of the host where the honeypot is located and the host k invaded by the attacker, represents the security requirement coefficient of the current defense scenario, represents the number of attack records obtained by the honeypot in the state , represents the attack ability of the attacker, represents the effective penetration information collected by the attacker in the current state ; (3) In addition, for the defender, when the defender selects Mutate the honeypot address as the defender's strategy. However, the mutation of the honeypot address will increase the communication transmission cost and operation and maintenance cost. When the defender selects as its own strategy, it is necessary to consider the defense cost required to expand the reconnaissance scope, including value and expansion breadth. Therefore, according to the following formula, construct the cost function of the defender: , In addition, for the attacker, the attacker's cost mainly comes from its attack ability, that is, the attack means and the ability to identify honeypots. Therefore, according to the following formula, construct the cost function of the attacker: , where, represents the communication overhead of the honeypot address mutation, represents the operation and maintenance cost of the honeypot under one mutation, represents the cost correlation coefficient of the defender, represents the reconnaissance scope of the defender's current honeypot, represents the cost correlation coefficient when expanding the reconnaissance scope of the honeypot, represents the attack cost correlation coefficient of the attacker.
[0039] Furthermore, in some alternative embodiments, the specific process of obtaining the state-action matrix from the initial two-agent Q-learning model is as follows: (1) First, establish a three-dimensional Q-table for two-agent Q-learning. The traditional single-agent Q-learning algorithm, as a value-based reinforcement learning algorithm, trains a two-dimensional table Q-table with states as rows and actions as columns, and obtains the optimal strategy by comparing the state-action values in the Q-table. Different from the traditional single-agent-based Q-learning algorithm, for multi-agent Q-learning in the network attack and defense scenario, it is necessary to consider the state-action values after multiple agents make decisions simultaneously in a single state; (2) Considering multi-agent two-agent Q-learning, it is necessary to consider the state-action values brought by the actions of both the attacker and the defender in the state, that is, . Then, take each action in the defender's action set as the row name, take each action in the attacker's action set as the column name, and fill all the Q values into the state-action matrix according to the row name and the column name to obtain a two-dimensional table , that is, the state-action matrix. For details, please refer to Table 1 below: Table 1
[0040] , Among them, for any state there is a corresponding two-dimensional table , then there exists a three-dimensional Q-table: .
[0041] In addition, during the initialization of the Q-table, the Q-table needs to be initialized during the first training. In this game, the instantaneous utility function is used as the initial value of the two-dimensional table, that is .
[0042] Furthermore, in some alternative embodiments, the process of using Minimax Q-learning for two-player Q-learning training is as follows: (1) Define the training cycle. The starting state of the training cycle is defined as any non-terminal state, and the ending state of the training cycle is related to the cumulative utility of both the attacker and the defender. The cumulative utility refers to the sum of the instantaneous utilities from the starting state to the current state, that is , where L represents the set of state transitions from the current state to the starting state, D represents the set of defender strategies that the defender has executed during the current training cycle, A represents the set of attacker strategies that the attacker has executed during the current training cycle, represents the sum of the instantaneous utilities from the starting state to the current state.
[0043] (2) Judge the sum of the instantaneous utilities of the current state. If the sum of the instantaneous utilities of the current state is greater than or equal to the maximum cumulative utility threshold , or, the sum of the instantaneous utilities of the current state is less than or equal to the minimum cumulative utility threshold , then enter the terminal state. That is to say, for the same training cycle, it satisfies . If the cumulative utility of the defender is greater than or equal to the maximum cumulative utility threshold or less than or equal to the minimum cumulative utility threshold, that is or , then enter the terminal state, and start the next training cycle when the current training cycle ends; (3) Use algorithm for exploration and exploitation. During the training process, with a probability of , select a random action (exploration), and with a probability of , select the current optimal action (exploitation). The size of is related to the number of training iterations , and satisfies where represents the change rate parameter of (4) Define a single double-agent Q-learning training. When training with the double-agent Q-learning algorithm, it is necessary to find the target Q-value corresponding to the optimal action in the current state of the current Q-value, and use the Bellman equation for iterative update through the target Q-value. In a zero-sum game, both sides of the game will try to maximize the benefits of both sides as much as possible. Therefore, in this step, the Minimax algorithm can be used to obtain the optimal target Q-value of the defender in the current state, that is , represents the new state entered after the attacker and the defender choose the original strategy in the current state, where represents the action chosen by the attacker that minimizes the defender's benefit, represents the action chosen by the defender that maximizes the defender's benefit when the attacker chooses ; (5) Conduct one Q-value iteration. For one iterative update in the same training cycle, let the current state to be updated be . For the target Q-value, that is , then is the new state entered when chooses the original strategy , is the value of the row where the minimum value in the state-action matrix is located, then is the value of the column where the maximum value in the row of the state-action matrix is located. Obtain through the above operations, and use the Bellman equation for iterative update: , where represents the discount factor, which determines the importance of future rewards in the current decision. represents the learning rate, which represents the proportion of the target Q-value in the iterative current Q-value.
[0044] (6) Iterate the Q-table in a loop. Within a single training cycle, iterate the Q-value in a loop until the training cycle reaches the termination state. After one training cycle terminates, start a new training cycle for loop training.
[0045] Further, in some alternative embodiments, the specific steps to obtain the optimal strategy are as follows: (1) Converge the double-agent Q-learning after sufficient training. To accelerate the convergence speed, two convergence indicators are set in this step. The first indicator is that the number of training times reaches the preset number of training cycles. There is a maximum number of training cycles , when the number of training cycles reaches When this occurs, the training converges. The second metric is the change in the Q-value. There is a preset minimum Q-value change threshold , and when, in a number of training iterations, the change in the Q-value is less than , then the training converges.
[0046] (2) Obtain the optimal strategy based on the trained three-dimensional Q-table. For the three-dimensional Q-table after training convergence: Any state corresponds to a state-action matrix . For the possible states that may occur during the attack and defense process, the optimal strategy for the defender is to, as an attacker, select the one with the lowest defender benefit in , and the defender, according to in the row where it is located, selects the one with the maximum defender income of . Then, for the defender, is the best strategy for the defender in state .
[0047] (3) The defender executes the optimal strategy. According to the obtained optimal strategy , the defender will mutate the honeypot address to the host address corresponding to through the QUIC protocol. In moving target defense, the address can be regarded as the unique identifier of the defender's host to a certain extent, and the attacker will also target the address for attack behavior. Traditional methods of modifying the address may cause service interruption, thus affecting the normal operation of the business. For the problem of service continuity, the address transformation method using the QUIC transport protocol can solve this problem. Based on the characteristic of quickly establishing a connection in QUIC, it can enable normal users to maintain business continuity when accessing the host with the transformed address.
[0048] It should also be noted that the honeypot address to be mutated can be randomly selected from a large number of pre-set honeypot addresses. In addition, the honeypot address generally includes the honeypot address and the honeypot CID.
[0049] In summary, according to the above-mentioned honeypot address mutation decision method, it has the following advantages: 1. By constructing a game model involving both the attacker and the defender, this invention models the attack and defense problems of moving target defense and defines the game elements of dynamic honeypot address transformation. Through this model, complex network attack and defense scenarios can be effectively analyzed and processed, and it can also adapt to changes in the network environment and attacker strategies in real time, adjust the defender's strategy, and improve the generalization and adaptability of the defense system.
[0050] 2. The present invention constructs a zero-sum game modeling utility function to comprehensively consider the benefits and costs of defenders and attackers in an attack and defense confrontation scenario. This makes the defense strategy more economically reasonable and can also link the opponent's benefits with its own benefits in the confrontation relationship, thereby more efficiently evaluating and optimizing the defense strategy.
[0051] 3. The present invention establishes a Q-table (state-action matrix) through a three-dimensional table. The use of the three-dimensional table fully considers the strategies of both the defender and the attacker, enabling a detailed evaluation of the expected discounted benefits of the attacker and the defender in each state, and making it more intuitive to understand and analyze the effectiveness and update direction of the attack and defense strategies.
[0052] 4. The present invention uses the Minimax Q-learning algorithm to solve the optimal strategy of the defender in the game model. This algorithm ensures that the defender can obtain the optimal solution even in the worst-case scenario, thereby improving the robustness of the defense system in an environment with rapid changes. At the same time, the Minimax Q-learning method can accelerate the convergence speed of training the Q-table, making the training more efficient.
[0053] Please refer to Figure 2 , which shows a schematic structural diagram of a honeypot address mutation decision system in an embodiment of the present invention. The system includes: A topology graph construction module 10, configured to construct an attack and defense scenario network topology graph and deploy honeypots to the attack and defense scenario network topology graph; A utility function construction module 20, configured to construct a game model and define the action sets of both the attacker and the defender in the game model, so as to obtain the benefit functions and cost functions of both the attacker and the defender according to the action sets, and construct an instantaneous utility function according to the benefit functions and the cost functions; An action matrix generation module 30, configured to construct and initialize a two-player Q-learning model according to the instantaneous utility function to obtain a state-action matrix, and the state-action matrix is filled with multiple state-action values; An optimal strategy obtaining module 40, configured to perform iterative training on the two-player Q-learning model to update the state-action matrix, and set the convergence index of the two-player Q-learning model, so as to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutate the selected honeypot address to the host address in the attack and defense scenario network topology graph according to the optimal strategy.
[0054] On the other hand, the present invention also proposes a readable storage medium, on which one or more programs are stored, and when the program is executed by a processor, the above-mentioned honeypot address mutation decision method is implemented.
[0055] On the other hand, the present invention also provides a computer device, including a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned honeypot address mutation decision method.
[0056] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0057] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0058] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in the memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0059] Although the embodiments of the present invention have been described in detail above, it will be obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present invention as described in the claims. Moreover, the present invention as described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A method for making a decision on a honey spot address mutation, characterized in that: The method comprises: Constructing a network topology diagram for an attack and defense scenario, and deploying honey points on the network topology diagram for the attack and defense scenario; Constructing a game model and defining an action set of the attacker and defender in the game model, so as to obtain a benefit function and a cost function of the attacker and defender according to the action set, and constructing an instantaneous utility function according to the benefit function and the cost function; Constructing and initializing a two-person Q-learning model according to the instantaneous utility function to obtain a state-action matrix, wherein the state-action matrix is filled with a plurality of state-action values; The two-person Q-learning model is iteratively trained to update the state-action matrix, and a convergence indicator of the two-person Q-learning model is set to obtain the optimal strategy of the defender according to the updated state-action matrix, and the selected honey spot address is mutated to the host address in the attack and defense scenario network topology diagram according to the optimal strategy.
2. The method for making a decision on a honey spot address mutation according to claim 1, characterized in that: The steps of constructing a game model and defining an action set of the attacker and defender in the game model include: The game model is constructed according to the following formula: , in, represents the game model, Indicates participants, Represents an action set, Represents a state set, Indicates instantaneous utility; The defender's action set is defined according to the following formula: , The attacker's action set is defined as follows: , The state sets of the attacker and defender are defined according to the following formula: , in, Indicates that in the current time slice, the defender mutates the honeypoint address to Hosts, Indicates that the honey point expands its reconnaissance range in the current time slice. Indicates that the attacker chooses to attack the host in the current time slice Invasion.
3. The method for making a decision on a honey spot address mutation according to claim 2, characterized in that: The step of obtaining the benefit function and cost function of both the attacker and the defender according to the action set, and constructing the instantaneous utility function according to the benefit function and the cost function comprises: The instantaneous utility function is constructed according to the following formula: , in, Indicates in status The defender’s instantaneous utility function is represents the defender’s payoff function, represents the attacker’s payoff function, represents the defender’s cost function, Denotes the attacker’s cost function.
4. The method for making a decision on a honey spot address mutation according to claim 3, characterized in that: The defender's profit function is constructed according to the following formula: , The attacker's profit function is constructed according to the following formula: , in, Indicates the attacker chooses an action As an attacker strategy, the defender chooses actions As a state of defender strategy, and Respectively represent the host where the honeypot is located and attackers hack into the host The value of Indicates the security requirement coefficient of the current defense scenario, Indicates that the defender is in state The number of attack records obtained by the honey spot at that time, Indicates the attacker's attack capability. Indicates the attacker is in the current state Effective penetration information collected during the process; The defender's cost function is constructed according to the following formula: , The attacker's cost function is constructed according to the following formula: , in, represents the communication cost of the honeypoint address mutation, represents the operation and maintenance cost of the honey point under a mutation, represents the cost correlation coefficient of the defender, Indicates the detection range of the defender's current honey spot. represents the cost correlation coefficient when expanding the honey spot reconnaissance range, Represents the attack cost correlation coefficient of the attacker.
5. The method for making a decision on a honey spot address mutation according to claim 4, characterized in that: The step of constructing and initializing a two-person Q-learning model according to the instantaneous utility function to obtain a state-action matrix, wherein the state-action matrix is filled with a plurality of state-action values, comprises: According to the instantaneous utility function, the state-action value in the state-action matrix is obtained: , in, Indicates in status The state action value of row j and column k at the time; Each action in the defender's action set is used as a row name, each action in the attacker's action set is used as a column name, and all Q values are filled into the state-action matrix according to the row name and the column name.
6. The method for making a decision on a honey spot address mutation according to claim 5, characterized in that: The step of iteratively training the two-person Q-learning model to update the state-action matrix includes: Define the training cycle, and calculate the cumulative utility based on the instantaneous utility function: , in, represents the instantaneous utility accumulation from the starting state to the current state, L represents the state transition set from the current state to the starting state, D represents the defender strategy set executed by the defender in the current training cycle, and A represents the attacker strategy set executed by the attacker in the current training cycle; If the instantaneous utility accumulation sum of the current state is greater than or equal to the maximum accumulated utility threshold, or the instantaneous utility accumulation sum of the current state is less than or equal to the minimum accumulated utility threshold, then enter the termination state; by The probability of selecting a random action is The probability of selecting the current optimal action satisfies ,in, express The rate of change parameter, t represents the number of training iterations; Define a single two-person Q-learning training to obtain a target Q value in a current state corresponding to a current Q value, wherein the target Q value corresponds to an optimal action in the current state; Use the Bellman equation to iterate and update the Q value once; When the training cycle reaches the terminal state, it enters the next training cycle to cyclically iterate the state-action matrix.
7. The method for determining the mutation of a honey spot address according to claim 6, characterized in that: The step of setting the convergence index of the two-person Q-learning model to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutating the selected honey spot address to the host address in the attack and defense scenario network topology diagram according to the optimal strategy includes: The convergence index includes a preset number of training cycles and a preset minimum Q value change threshold. If the number of training cycles reaches the preset number of training cycles, the training converges; Alternatively, if the change in Q value is less than the preset minimum Q value change threshold in several training iterations, the training converges; After training, the attacker chooses the one with the lowest defender benefit. , the defender based on The line where the defender has the highest profit is selected , Indicates status The attacker’s optimal strategy is Indicates status The optimal strategy for the defender. Mutate the honeypoint address to The corresponding host address.
8. The method for making a decision on a honey spot address mutation according to claim 6, characterized in that: The step of defining a single two-person Q-learning training to obtain a target Q value in a current state corresponding to a current Q value, wherein the target Q value corresponds to an optimal action in the current state includes: The target Q value is calculated according to the following formula: , in, It indicates the new state that the attacker and defender enter after choosing the original strategy in the current state. represents the action chosen by the attacker to minimize the defender's payoff, Indicates that the defender chooses The action that maximizes the defender's payoff under the circumstances.
9. A honey spot address mutation decision system, characterized in that: The system comprises: A topology construction module is used to construct a network topology map of an attack and defense scenario and deploy honey points on the network topology map of the attack and defense scenario; A utility function construction module is used to construct a game model and define an action set of the attacker and defender in the game model, so as to obtain the benefit function and cost function of the attacker and defender according to the action set, and construct an instantaneous utility function according to the benefit function and the cost function; An action matrix generation module, used to construct and initialize a two-person Q-learning model according to the instantaneous utility function to obtain a state-action matrix, wherein the state-action matrix is filled with a plurality of state-action values; An optimal strategy acquisition module is used to iteratively train the two-person Q-learning model to update the state-action matrix, and set the convergence index of the two-person Q-learning model to obtain the optimal strategy of the defender according to the updated state-action matrix, and mutate the selected honey spot address to the host address in the attack and defense scenario network topology diagram according to the optimal strategy.
10. A readable storage medium, characterized in that: The readable storage medium stores one or more programs, which, when executed by a processor, implement the honey spot address mutation decision method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Preknown honey spot deployment method and system based on intelligent time-delay differential game, and server
CN118041645A
Game model-based network defense method and system, and readable storage medium
CN119011243A