Host scanning and defense strategy collaborative optimization method based on random game
By optimizing host scanning and defense strategies based on random game methods, the local optimization problem of scanning and detection technology in network security is solved, and the dynamic balance of offense and defense strategies and the improvement of detection accuracy are achieved.
Patent Information
- Application Number
- CN202510444390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, horizontal scanning and defense strategies have efficiency and concealment conflicts in network security. The group scanning strategies face challenges in traffic aggregation effects and environmental dynamic characteristics. The threshold detection mechanism on the defense side cannot adapt to the time-varying characteristics of traffic, resulting in scanning and detection technology falling into suboptimal balance of local optimization.
The coordinated optimization method of host scanning and defense strategy based on random game is adopted to build the action space of attackers and defenders, and solve Nash equilibrium is used to solve Nash equilibrium by using Lemke-Howson algorithm and Nash Q-Learning to dynamically optimize offensive and defense strategies to improve response efficiency and detection accuracy.
Real-time dynamic optimization of offense and defense strategies has been achieved, the dynamic balance between efficiency, security and availability of the network security system has been improved, and scanning efficiency and detection accuracy have been improved.
Smart Images

Figure CN120389879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a collaborative optimization method for host scanning and defense strategies based on stochastic games. Background Art
[0002] As a fundamental means for network space asset mapping and security assessment, the core contradiction of horizontal scanning technology lies in the fundamental conflict between efficiency and concealment. The traditional serial scanning mechanism is limited by the linearly increasing detection time-consuming and is difficult to meet the rapid mapping requirements of large-scale target networks. The grouped scanning strategy introduces a parallel architecture, divides the target set into subgroups and implements concurrent detection within the groups. Under ideal conditions, the scanning efficiency can be increased to an order of magnitude linearly related to the number of groups. However, this optimization path faces double constraints: First, although the expansion of the group size can reduce the inter-group switching overhead, it will trigger the spatio-temporal aggregation effect of detection traffic, significantly enhancing the abnormal salience of traffic characteristics. Second, the dynamic characteristics of the network environment, including service load fluctuations, service access mode changes, and link quality changes, make it difficult for any fixed grouping strategy to maintain the global optimum. Attackers have to make a static trade-off between scanning speed and exposure risk, and this rigid decision-making mode is essentially inconsistent with the elastic characteristics of the network environment, resulting in severely limited adaptability of existing scanning tools in complex scenarios.
[0003] On the defense side, the anomaly detection mechanism based on threshold determination has long faced fundamental defects: its core logic is established on the assumption of the steady-state distribution of traffic characteristics and rigidly responds to dynamic confrontation through static rules. The intrusion detection system (IDS) intercepts by counting the connection request frequencies of a single source address, but the setting of a fixed threshold forces the defender to make an either-or choice between false positives and false negatives. This one-dimensional decision-making framework cannot adapt to the time-varying characteristics of network traffic. For example, the sudden access of legitimate services and the slow penetration of attack behaviors may present similar characteristics at the traffic level, and the static threshold can neither distinguish the essential differences between the two nor respond to the strategy evolution of attackers. More seriously, attackers can use the publicity or detectability of threshold parameters to accurately determine the defense boundary through exploratory scanning, and then design sub-threshold attack patterns to avoid detection. The defense system thus falls into a vicious cycle of passive response: reducing the threshold increases the attack capture rate at the cost of sacrificing service availability; increasing the threshold reduces false positives but opens a penetration window for advanced persistent threats.
[0004] The current technical system regards scanning and detection as isolated technical actions rather than strategic games in dynamic confrontation, resulting in both the attacker and the defender falling into a sub-optimal equilibrium of local optimization. The grouping strategy of the attack side and the threshold rule of the defense side are both solidified in the form of preset parameters, lacking both real-time perception of the environmental state and feedback adjustment of the confrontation results. This fragmented design causes two negative effects: on the one hand, the fixed grouping pattern of the attacker generates periodic traffic fingerprints, providing stable recognition features for the machine learning detection model; on the other hand, the static threshold of the defender cannot distinguish between normal business fluctuations and the gradual change of malicious behavior, resulting in a continuous mismatch between the security policy and the real threat. The deeper problem is that the existing technical framework fails to establish an interactive mechanism model for attack and defense strategies, unable to quantify the impact of the adjustment of the attacker's scanning strategy on the detection efficiency, nor to evaluate the change in the attack cost caused by the change of the defense rule. The lack of this systematic modeling ability makes the evolution of scanning and detection technologies fall into the quagmire of zero-sum games and unable to achieve the dynamic balance among efficiency, security, and usability in the network security system. Therefore, there is an urgent need to provide a solution to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for collaborative optimization of host scanning and defense strategies based on stochastic games, which improves the problems of insufficient response efficiency and detection accuracy of traditional fixed strategies.
[0006] A method for collaborative optimization of host scanning and defense strategies based on stochastic games provided by the present invention adopts the following technical solutions:
[0007] Based on the grouped scanning method, formulate the attack action space of the attacker at each time step, perform attack actions based on the attack action space until the scanning of the target network is completed, construct an action strategy based on the attack actions at all time steps, and construct an attack strategy space based on the action strategy;
[0008] Based on the detection threshold of malicious scanning behavior, formulate the defense action space of the defender at each time step, perform defense actions based on the defense action space until the attack action ends, construct a defense strategy based on the defense actions at all time steps, and construct a defender strategy space based on the defense strategy;
[0009] Construct the game state at the current moment based on the number of successfully scanned hosts, the number of false alarms, and the alarm trigger status, determine the game state at the next moment based on the attack action and the defense action, and calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state;
[0010] Initialize the Q-tables of both the attacker and the defender. Transform the game into a linear complementarity problem based on the Lemke-Howson algorithm, solve for the single-step Nash equilibrium through pivot operations, determine the optimal actions of both the attacker and the defender and execute them, update the rewards of the execution results to the Q-tables, and iterate until the hosts of the target network are fully scanned;
[0011] Calculate the total attack revenue and total defense revenue based on the instantaneous rewards of the attacker and the defender at each time step. Select the optimal strategy based on the total attack revenue and total defense revenue, solve for the optimal action at each time step based on Nash Q-Learning, generate a dynamic action path covering the entire game process, and obtain the equilibrium strategy.
[0012] Optionally, in the process of formulating the attacker's action space based on the grouped scanning method, it includes: the attacker adopts the grouped scanning method, specifies a port number, and only sends probe packets in parallel to the specified port of a group of hosts at each time step. Take the number of hosts included in each time step as the attacker's action, and dynamically reduce the action space of the target hosts based on the attacker's action.
[0013] Optionally, when the alarm trigger status is 0, it means that the attacker's scanning group size does not exceed the threshold set by the defender, that is, no alarm has been triggered during the previous attack action; when the alarm trigger status is 1, it means that the attacker's scanning group size first exceeds the threshold set by the defender, and the defender detects an abnormal scanning behavior at this time, and it happens to be at the moment when the state changes from never triggered to triggering an alarm; when the alarm trigger status is 2, it means that the alarm has been triggered, and the system has been in the state where the defense measures are activated before.
[0014] Optionally, in the process of determining the next moment's game state based on the attack action and the defense action, it includes: when the alarm trigger status is 0, update the number of successfully scanned hosts and the number of false alarms in the state space based on the attack action and the defense action; when the alarm trigger status is 1, update the number of false alarms in the state space based on the attack action and the defense action, and update the alarm trigger status; when the alarm trigger status is 2, update the number of false alarms in the state space based on the attack action and the defense action.
[0015] Optionally, in the process of updating the number of false alarms in the state space based on the attack action and the defense action, when the defense action misjudges a legitimate request as a malicious scan, add the number of legitimate requests to the number of false alarms.
[0016] Optionally, calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state, including: calculate the attacker's instantaneous reward based on the time cost of initiating the scan, the number of successfully scanned hosts, and the alarm trigger penalty of the attack action; calculate the defender's instantaneous reward based on the number of false alarms, the successful detection reward, and the undetected penalty of the defense action.
[0017] Optionally, in the process of initializing the Q-tables of both the attacker and the defender, it includes: constructing a three-dimensional Q-table for both the attacker and the defender, with dimensions of the number of states, the number of its own actions, and the number of actions of other agents. The initial Q-values are sampled from a uniform distribution U(-ε, ε), where ε is a preset minimum value, such as 0.01.
[0018] Optionally, in the structure of the three-dimensional Q-table, the two-dimensional Q-table corresponding to each game state is in matrix form, with rows representing its own actions and columns representing the actions of the other party.
[0019] Optionally, in the process of transforming the game into a linear complementary problem based on the Lemke-Howson algorithm, it includes: constructing a payoff matrix M for both the attacker and the defender, introducing artificial variables to transform the game into a linear complementary problem, and finding non-negative vectors x and y in the two-dimensional Q-table corresponding to the game state such that y = Mx + q and x T y = 0, where q is the artificial variable.
[0020] Optionally, in the process of calculating the total attack payoff and the total defense payoff based on the instantaneous rewards of each time step of the attacker and the defender, the calculation method is to take the expectation sum of the instantaneous rewards of each time step.
[0021] A method for collaborative optimization of host scanning and defense strategies based on stochastic games provided by the present invention has the beneficial effects as follows:
[0022] 1. The present invention constructs a three-dimensional state space with the number of successfully scanned hosts, the number of false alarms, and the alarm trigger status, which can comprehensively reflect the real-time dynamics of both the attacker and the defender, and optimize the strategy selection of both the attacker and the defender in real time according to the dynamics, thereby improving the response efficiency and detection accuracy;
[0023] 2. In the equilibrium strategy of the present invention, the strategies of the attacker and the defender are optimal responses to each other. Any party's individual deviation from the equilibrium strategy will lead to a decrease in its own payoff, thus ensuring the long-term stability of the strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the method for collaborative optimization of host scanning and defense strategies provided by an embodiment of the present invention;
[0025] Figure 2 It is a schematic diagram of an attack action provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described clearly and completely below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present invention pertains. The words such as "including" used herein are intended to mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.
[0027] The embodiments of the present invention provide a method for collaborative optimization of host scanning and defense strategies based on stochastic games. Refer to Figure 1 , including:
[0028] S1. Formulate the attack action space of the attacker at each time step based on the grouped scanning method, perform attack actions based on the attack action space until the scanning of the target network is completed, construct an action strategy based on the attack actions of all time steps, and construct an attack strategy space based on the action strategy;
[0029] S2. Formulate the defense action space of the defender at each time step based on the malicious scanning behavior detection threshold, perform defense actions based on the defense action space until the attack actions end, construct a defense strategy based on the defense actions of all time steps, and construct a defender strategy space based on the defense strategy;
[0030] S3. Construct the game state at the current moment based on the number of successfully scanned hosts, the number of false alarms, and the alarm trigger status, determine the game state at the next moment based on the attack actions and defense actions, and calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state;
[0031] S4. Initialize the Q-tables of both the attacker and the defender, transform the game into a linear complementary problem based on the Lemke-Howson algorithm, solve the single-step Nash equilibrium through pivot operations, determine the optimal actions of both the attacker and the defender and execute them, update the rewards of the execution results to the Q-tables, and iterate until all the hosts in the target network are completely scanned;
[0032] S5. Calculate the total attack benefit and the total defense benefit based on the instantaneous rewards of the attacker and the defender at each time step, select the optimal strategy based on the total attack benefit and the total defense benefit, solve for the optimal action at each time step based on Nash Q-Learning, generate a dynamic action path covering the entire game process, and obtain the equilibrium strategy.
[0033] In some embodiments, during the execution of step S1, it includes:
[0034] S1.1. Develop the attacker's attack action space at each time step based on the grouped scanning method;
[0035] S1.2. Construct an action strategy based on the attack actions at all time steps.
[0036] Specifically, during the execution of step S1.1, refer to Figure 2 , including: The attacker uses the grouped scanning method, specifies a port number, and only sends probe packets to the specified port in a group of hosts in parallel at each time step. The number of hosts included in each time step is used as the attacker's action, and the attacker's action space is dynamically reduced based on the attack actions before the current moment.
[0037] Actually, the attack action space is expressed as is the attack action, N is the number of hosts in the target network, and the attack action space A att is dynamically reduced by the cumulative value of historical attack actions to ensure that the remaining number of hosts to be scanned is non - negative.
[0038] Furthermore, execute step S1.2, construct an action strategy based on the attack actions at all time steps, and construct an attack strategy space from all possible action strategies.
[0039] In some embodiments, during the execution of step S2, it includes:
[0040] S2.1. Develop the defender's defense action space at each time step based on the malicious scanning behavior detection threshold;
[0041] S2.2. Construct an action strategy based on the attack actions at all time steps.
[0042] Specifically, during the execution of step S2.1, it includes: The defender uses a threshold - based malicious scanning behavior detection method, sets a threshold for the number of homologous request access packets received from different target addresses in a time period. When the scanning probe exceeds this threshold, it is determined as abnormal scanning traffic. The defender's action is the threshold selected based on the current game state, and the defense action space is expressed as A def = {0, 1,..., N}, where N is the number of hosts in the target network.
[0043] Furthermore, the actions selected by the defender at all time steps from the start of scanning to the completion of scanning form a strategy, and all possible strategies form a strategy space.
[0044] In some embodiments, during the execution of step S3, it includes:
[0045] S3.1. Define the game state;
[0046] S3.2. Game state transition;
[0047] S3.3. Calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state.
[0048] Specifically, in the process of executing step S3.1 to define the game state, it includes: constructing the game state at the current moment based on the number of successfully scanned hosts, the number of false positives, and the alarm trigger status, determining the game state at the next moment based on the attack action and the defense action, and calculating the instantaneous rewards of the attacker and the defender at each time step based on the game state.
[0049] Further, the game state is represented as a triple s = (n, m, d), where: n ≤ N represents the number of hosts that have been successfully scanned currently, which can reflect the scanning progress of the attacker, m represents the number of false positives of the IDS for the current legitimate source address, and d represents a boolean value, and its different values have specific status indication meanings.
[0050] Specifically, when the alarm trigger status d is 0, it means that the scale of the attacker's scanning packets does not exceed the threshold set by the defender, that is, no alarm has been triggered during the previous attack action; when the alarm trigger status d is 1, it means that the scale of the attacker's scanning packets first exceeds the threshold set by the defender, and the defender detects an abnormal scanning behavior at this time, and it is exactly at the moment when the state changes from never triggered to triggering an alarm; when the alarm trigger status d is 2, it means that an alarm has been triggered and the system has been in the state where the defense measures have been activated before.
[0051] Further, in the process of executing step S3.2 to determine the game state at the next moment based on the attack action and the defense action, when the alarm trigger status is 0, update the number of successfully scanned hosts and the number of false positives in the state space based on the attack action and the defense action; when the alarm trigger status is 1, update the number of false positives in the state space based on the attack action and the defense action, and update the alarm trigger status; when the alarm trigger status is 2, update the number of false positives in the state space based on the attack action and the defense action.
[0052] Further, in the process of updating the number of false positives in the state space based on the attack action and the defense action, when the defense action misjudges a legitimate request as a malicious scan, add the number of legitimate requests to the number of false positives.
[0053] Specifically, in the process of executing step S3.3 to calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state, it includes: calculating the instantaneous reward of the attacker based on the time cost of initiating a scan for an attack action, the number of successfully scanned hosts, and the penalty for triggering an alarm; calculating the instantaneous reward of the defender based on the number of false alarms for a defense action, the reward for successful detection, and the penalty for non-detection.
[0054] Furthermore, the expressions of the reward functions of the attacker and the defender at time step t are as follows:
[0055]
[0056] Where is the instantaneous reward of the attacker at time step t. When d k = 0, the attacker does not trigger an alarm, and the attack instantaneous reward consists of the time cost and the reward for successfully scanning hosts . When d k = 1, the attacker triggers an alarm for the first time, and the attack instantaneous reward consists of the time cost and the penalty for triggering an alarm . When d k = 2, the defense system has entered the continuous alert mode, and no matter how the attacker acts, no additional benefits can be obtained. The attack instantaneous reward consists of the time cost ; is the instantaneous reward of the defender at time step t. When d k = 0, the defense instantaneous reward consists of the false alarm penalty and the non-detection penalty . When d k = 1, the defense instantaneous reward consists of the false alarm penalty and the reward for successful detection . When d k = 2, the defense consists of the false alarm penalty ;
[0057] In some embodiments, in the process of executing step S4, it includes:
[0058] S4.1. Initialize the Q-tables of both the attacker and the defender;
[0059] S4.2. Solve for the optimal actions under the single-step equilibrium condition.
[0060] Specifically, in the process of executing step S4.1, it includes: constructing a three-dimensional Q-table for both the attacker and the defender, with dimensions of the number of states, the number of its own actions, and the number of actions of other agents. The initial Q-values are sampled from a uniform distribution U(-ε, ε), where ε is a preset minimum value, such as 0.01.
[0061] Further, in the three-dimensional structure of the Q-table, the Q value corresponding to each state is in matrix form, where the rows represent its own actions and the columns represent the actions of the other party.
[0062] Specifically, in the process of executing step S4.2, the game is transformed into a linear complementary problem based on the Lemke-Howson algorithm, the single-step Nash equilibrium is solved through pivot operations, the optimal actions of both the attacker and the defender are determined and executed, and the rewards of the execution results are updated to the Q-table, and the loop iteration is performed until the host of the target network is completely scanned;
[0063] Further, in the process of transforming the game into a linear complementary problem based on the Lemke-Howson algorithm, it includes: constructing the payoff matrix M of both the attacker and the defender, introducing artificial variables to transform the game into a linear complementary problem, and finding non-negative vectors x and y in the two-dimensional Q-table corresponding to the game states, such that y = Mx + q and x T y = 0, where q is the artificial variable.
[0064] In some embodiments, in the process of executing step S5, it includes:
[0065] S5.1. Calculate the total attack revenue and the total defense revenue based on the instantaneous rewards of the attacker and the defender at each time step;
[0066] S5.2. Define the optimization strategy;
[0067] S5.3. Generate the equilibrium strategy.
[0068] Specifically, in the process of executing step S5.1, when calculating the total attack revenue and the total defense revenue based on the instantaneous rewards of the attacker and the defender at each time step, the calculation method is to take the expectation sum of the instantaneous rewards at each time step.
[0069] Further, the calculation formula for the total attack revenue is:
[0070]
[0071] where, U att is the total attack revenue, π att is the attack strategy, π def is the defense strategy, is the attack instantaneous reward at time step k, and T is the sum of all time steps;
[0072] The calculation formula for the total defense revenue is:
[0073]
[0074] where, U def is the total defense revenue, π att is the attack strategy, π defAs a defense strategy, is the defense instantaneous reward at time step k, and T is the sum of all time steps.
[0075] Specifically, during the execution of step S5.2, the purpose of both the attacker and the defender is to select the optimal strategy according to their respective reward functions to maximize their own reward functions.
[0076] Furthermore, execute step S5.3, solve for the optimal action at each time step based on Nash Q-Learning, generate a dynamic action path covering the entire game process, and obtain an equilibrium strategy.
[0077] Furthermore, the equilibrium strategy satisfies the following conditions:
[0078]
[0079] Among them, is the attack equilibrium strategy, is the defense equilibrium strategy, and π att is the attack strategy space ∏ att except for any attack strategy other than, and π def is the defense strategy space Π def any defense strategy other than.
[0080] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A collaborative optimization method for host scanning and defense strategies based on stochastic games, characterized in that, Including the following steps: Formulate the attack action space of the attacker at each time step based on the grouped scanning method, perform attack actions based on the attack action space until the scanning of the target network is completed, construct an action strategy based on the attack actions at all time steps, and construct an attack strategy space based on the action strategy; Formulate the defense action space of the defender at each time step based on the malicious scanning behavior detection threshold, perform defense actions based on the defense action space until the attack actions end, construct a defense strategy based on the defense actions at all time steps, and construct a defender strategy space based on the defense strategy; Construct the game state at the current moment based on the number of successfully scanned hosts, the number of false positives, and the alarm trigger status, determine the game state at the next moment based on the attack actions and defense actions, and calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state; Initialize the Q-tables of both the attacker and the defender, transform the game into a linear complementary problem based on the Lemke-Howson algorithm, solve the single-step Nash equilibrium through pivot operations, determine the optimal actions of both the attacker and the defender and execute them, update the rewards of the execution results to the Q-tables, and iterate until all the hosts in the target network are completely scanned; Calculate the total attack benefit and the total defense benefit based on the instantaneous rewards of the attacker and the defender at each time step, select the optimal strategy based on the total attack benefit and the total defense benefit, solve for the optimal action at each time step based on Nash Q-Learning, generate a dynamic action path covering the entire game process, and obtain the equilibrium strategy.
2. The collaborative optimization method for host scanning and defense strategies based on stochastic game according to claim 1, wherein In the process of formulating the attacker's action space based on the grouped scanning method, it includes: The attacker adopts the grouped scanning method, specifies a port number, and only sends probe packets in parallel to the specified port of a group of hosts at each time step. Take the number of hosts included in each time step as the action of the attacker, and dynamically reduce the action space of the target hosts based on the action of the attacker.
3. A collaborative optimization method for host scanning and defense strategies based on stochastic games according to claim 1, characterized in that Including: When the alarm trigger status is 0, it means that the attacker's scanning group size does not exceed the threshold set by the defender, that is, no alarm has been triggered during the previous attack action process; When the alarm trigger status is 1, it means that the attacker's scanning group size first exceeds the threshold set by the defender. The defender detects abnormal scanning behavior at this time, and it is exactly at the moment when the state changes from never triggered to triggering an alarm; When the alarm trigger status is 2, it means that an alarm has been triggered, and the system has been in the state where the defense measures are activated before.
4. A collaborative optimization method for host scanning and defense strategies based on stochastic game according to claim 3, characterized in that, In the process of determining the game state at the next moment based on the attack actions and defense actions, it includes: When the alarm trigger status is 0, update the number of successfully scanned hosts and the number of false positives in the state space based on the attack actions and defense actions; When the alarm trigger status is 1, update the number of false positives in the state space based on the attack actions and defense actions, and update the alarm trigger status; When the alarm trigger status is 2, update the number of false positives in the state space based on the attack actions and defense actions.
5. A collaborative optimization method for host scanning and defense strategies based on stochastic game according to claim 4, characterized in that, In the process of updating the number of false positives in the state space based on the attack actions and defense actions, when the defense action misjudges a legitimate request as a malicious scan, add the number of legitimate requests to the number of false positives.
6. The collaborative optimization method for host scanning and defense strategy based on stochastic game according to claim 1, characterized in that Calculate the instantaneous rewards of the attacker and the defender at each time step based on the game state, including: Calculate the instantaneous reward of the attacker based on the time cost of initiating a scan by the attack action, the number of hosts successfully scanned, and the penalty for triggering an alarm; Calculate the instantaneous reward of the defender based on the number of false alarms of the defense action, the reward for successful detection, and the penalty for undetected.
7. A collaborative optimization method for host scanning and defense strategies based on stochastic games according to claim 1, characterized in that, In the process of initializing the Q-tables of both the attacker and the defender, including: Construct a three-dimensional Q-table for both the attacker and the defender, with dimensions of the number of states, the number of its own actions, and the number of actions of other agents. The initial Q-values are sampled from a uniform distribution U(-ε, ε), where ε is a preset minimum value, such as 0.
01.
8. A collaborative optimization method for host scanning and defense strategies based on stochastic games according to claim 7, characterized in that In the structure of the three-dimensional Q-table, the two-dimensional Q-table corresponding to each game state is in matrix form, with rows representing its own actions and columns representing the actions of the other party.
9. A collaborative optimization method for host scanning and defense strategies based on stochastic game according to claim 8, characterized in that In the process of transforming the game into a linear complementarity problem based on the Lemke-Howson algorithm, including: Construct the payoff matrix \(M\) for both the attacker and the defender, introduce artificial variables to transform the game into a linear complementarity problem, and find non-negative vectors \(x\) and \(y\) in the two-dimensional Q-table corresponding to the game state such that \(y = Mx+q\) and \(x\) T \(y = 0\), where \(q\) is the artificial variable.
10. A method for collaborative optimization of host scanning and defense strategies based on stochastic games, characterized in that, In the process of calculating the total attack benefit and the total defense benefit based on the instantaneous rewards of the attacker and the defender at each time step, the calculation method is to take the expectation sum of the instantaneous rewards at each time step.
Citation Information
Cited By
Dynamic defense switching period optimization method and device, medium and electronic equipment
CN120979835A
Dynamic red team detection method for safe alignment of large language model
CN121770795A