Defence strategy model training method, defence strategy determination method and apparatus

By acquiring defense status data from sample servers and adjusting parameters using defense strategy model training methods, the problem of defense schemes relying on predefined identification standards in existing technologies is solved, enabling more reasonable defense strategy deployment and effective defense against intelligent attacks.

CN116707870BActive Publication Date: 2026-01-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310557573.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2026-01-06
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

Existing network attack defense solutions rely too heavily on predefined identification criteria, resulting in poor defense effectiveness and difficulty in dealing with intelligent attacks.

Method used

By acquiring defense status data from sample servers and using defense strategy model training methods, the parameters of the defense strategy model are adjusted to determine the probability of each server selecting to deploy a defense strategy. Based on the characteristics of each server, corresponding strategies are deployed to improve the rationality and efficiency of the defense strategy.

Benefits of technology

It enables a more rational deployment of defense strategies, improving defense efficiency and defense performance against smart attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116707870B_ABST
    Figure CN116707870B_ABST
Patent Text Reader

Abstract

The application provides a defense strategy model training method, a defense strategy determination method and equipment. The method comprises the following steps: obtaining sample defense state data of a plurality of sample servers in a first time period; processing the sample defense state data in the first time period according to a defense strategy model to obtain a plurality of sets of action selection probabilities corresponding to a plurality of time periods respectively and a plurality of state values corresponding to the plurality of time periods respectively; and adjusting parameters of the defense strategy model according to the plurality of sets of action selection probabilities corresponding to the plurality of time periods respectively and the plurality of state values corresponding to the plurality of time periods respectively to obtain a trained defense strategy model. The defense strategy model can be used to determine the probability of each sample server selecting and deploying a defense strategy, so that the corresponding defense strategy can be more reasonably deployed in combination with the characteristics of the server itself to prevent network attacks, the efficiency and rationality of the deployment of the defense strategy are improved, and good defense performance against intelligent attacks can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method for training a defense strategy model, a method for determining a defense strategy, and a device. Background Technology

[0002] With the rapid development of cutting-edge communication technologies such as 5G high-speed networks and ubiquitous mobile communications, network traffic has seen a significant increase in both quantity and complexity. At the same time, traffic security is facing a constant stream of emerging malicious attack patterns.

[0003] Therefore, it is necessary to take appropriate measures to defend against cyberattacks. Currently, the main defense measures against cyberattacks are based on predefined identification standards to identify and defend against them. These measures have high detection accuracy and efficiency, but over-reliance on predefined identification standards can only mitigate the harm of malicious traffic.

[0004] In summary, the above-mentioned defense solutions are ineffective, and there is an urgent need to provide a more effective network attack defense solution. Summary of the Invention

[0005] This application provides a method for training a defense strategy model, a method for determining a defense strategy, and an apparatus to address the problem of poor defense effectiveness against network attacks in existing technologies.

[0006] In a first aspect, this application provides a method for training a defense strategy model, applied to a first device, the method comprising:

[0007] The sample defense status data of multiple sample servers in the first time period is obtained. For each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in the preset defense strategy set during the first time period.

[0008] According to the defense strategy model, the sample defense status data of the first time period is processed to obtain the action selection probability set and the status value corresponding to the multiple time periods respectively. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability of the sample server selecting to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0009] Based on the action selection probability set corresponding to each of the multiple time periods and the state value corresponding to each of the multiple time periods, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

[0010] In one possible implementation, the step of processing the sample defense state data of the first time period according to the defense strategy model to obtain action selection probability sets corresponding to multiple time periods and state values ​​corresponding to multiple time periods includes:

[0011] The sample defense state data of the t-th time period is input into the defense strategy model to obtain the action selection probability set corresponding to the t-th time period and the state value corresponding to the t-th time period; the initial value of t is 1, and t is a positive integer greater than or equal to 1;

[0012] Based on the action selection probability set corresponding to the t-th time period, determine the sample defense status data of the multiple sample servers in the (t+1)-th time period;

[0013] Update t to t+1 and repeat the above operation until t is greater than or equal to T, to obtain the action selection probability set and state value corresponding to T time periods respectively, where T is a positive integer.

[0014] In one possible implementation, the defense strategy model includes a sub-action selection network and a state value network; the step of inputting the sample defense state data of the t-th time period into the defense strategy model to obtain the action selection probability set corresponding to the t-th time period and the state value corresponding to the t-th time period includes:

[0015] The sample defense status data of the t-th time period is input into the sub-action selection network to obtain the sub-action selection probability of each sample server for each defense strategy in the t-th time period;

[0016] Based on the sub-action selection probability of each sample server for each defense strategy in the t-th time period, the action selection probability set corresponding to the t-th time period is obtained;

[0017] The sample defense state data of the t-th time period is input into the state value network to obtain the state value corresponding to the t-th time period.

[0018] In one possible implementation, obtaining the action selection probability set corresponding to the t-th time period based on the sub-action selection probability of each of the sample servers for each of the defense strategies in the t-th time period includes:

[0019] In the t-th time period, for any sample server, the total probability corresponding to the sample server is obtained based on the sub-action selection probability of the sample server for each of the defense strategies.

[0020] Based on the total probability, the sub-action selection probability of the sample server for each of the defense strategies is normalized to obtain the normalized probability of the sample server selecting sub-actions for each of the defense strategies.

[0021] Based on the normalized probabilities of sub-action selection for each of the aforementioned sample servers in response to each of the aforementioned defense strategies, the set of action selection probabilities corresponding to the t-th time period is obtained.

[0022] In one possible implementation, determining the sample defense status data of the plurality of sample servers in the (t+1)th time period based on the action selection probability set corresponding to the t-th time period includes:

[0023] Based on the action selection probability set corresponding to the t-th time period, determine the sub-actions that the multiple sample servers will choose to execute for each of the defense strategies in the (t+1)-th time period;

[0024] Based on the sub-actions selected by the multiple sample servers for each of the defense strategies during the (t+1)th time period, the multiple sample servers are instructed to update the deployed defense strategies.

[0025] Based on the updated defense strategies deployed by the multiple edge proxy servers, the sample defense status data of the multiple sample servers in the (t+1)th time period is determined.

[0026] In one possible implementation, determining the sub-actions selected by the plurality of sample servers for each defense strategy in the (t+1)th time period based on the action selection probability set corresponding to the t-th time period includes:

[0027] Determine the maximum number of defense strategies that can be deployed on each of the sample servers;

[0028] Based on the action selection probability set corresponding to the t-th time period and the maximum number of defense strategies that can be deployed on each of the sample servers, the sub-actions to be executed by the multiple sample servers for each of the defense strategies in the (t+1)-th time period are determined.

[0029] In one possible implementation, adjusting the parameters of the defense strategy model based on the action selection probability set corresponding to each of the multiple time periods and the state value corresponding to each of the multiple time periods to obtain a trained defense strategy model includes:

[0030] For any given time period, based on the sample defense status data for that time period, obtain the intrinsic security reward parameters of the defense strategy corresponding to that time period;

[0031] Based on the sample defense status data for each time period, the corresponding intrinsic security reward parameters of the defense strategy, the corresponding action selection probability set, the corresponding state value, and the sub-actions selected and executed by the multiple sample servers for each defense strategy, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

[0032] Secondly, this application provides a method for determining a defense strategy, applied to a second device, the method comprising:

[0033] Obtain the defense status data of multiple servers in the current time period. For each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in the preset defense strategy set in the current time period.

[0034] The defense status data of the current time period is input into the defense strategy model to obtain the action selection probability set corresponding to the current time period, wherein the defense strategy model is a model trained according to the method described in any one of the first aspects;

[0035] Based on the action selection probability set corresponding to the current time period, determine the sub-actions that the multiple servers will choose to execute for each of the defense strategies in the next time period. The sub-actions include deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0036] For any of the plurality of servers, according to the sub-actions that the server selects to execute for each of the defense strategies in the next time period, an instruction message is sent to the server, the instruction message being used to instruct the server to execute the sub-actions selected for each of the defense strategies.

[0037] Thirdly, this application provides a defense strategy model training device, comprising:

[0038] The acquisition module is used to acquire sample defense status data of multiple sample servers in the first time period. For each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in the preset defense strategy set during the first time period.

[0039] The processing module is used to process the sample defense status data of the first time period according to the defense strategy model to obtain the action selection probability set corresponding to multiple time periods and the status value corresponding to multiple time periods respectively. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability of the sample server selecting to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0040] The training module is used to adjust the parameters of the defense strategy model based on the action selection probability set corresponding to the multiple time periods and the state value corresponding to the multiple time periods, so as to obtain a trained defense strategy model.

[0041] In one possible implementation, the processing module is specifically used for:

[0042] The sample defense state data of the t-th time period is input into the defense strategy model to obtain the action selection probability set corresponding to the t-th time period and the state value corresponding to the t-th time period; the initial value of t is 1, and t is a positive integer greater than or equal to 1;

[0043] Based on the action selection probability set corresponding to the t-th time period, determine the sample defense status data of the multiple sample servers in the (t+1)-th time period;

[0044] Update t to t+1 and repeat the above operation until t is greater than or equal to T, to obtain the action selection probability set and state value corresponding to T time periods respectively, where T is a positive integer.

[0045] In one possible implementation, the defense strategy model includes a sub-action selection network and a state value network; the processing module is specifically used for:

[0046] The sample defense status data of the t-th time period is input into the sub-action selection network to obtain the sub-action selection probability of each sample server for each defense strategy in the t-th time period;

[0047] Based on the sub-action selection probability of each sample server for each defense strategy in the t-th time period, the action selection probability set corresponding to the t-th time period is obtained;

[0048] The sample defense state data of the t-th time period is input into the state value network to obtain the state value corresponding to the t-th time period.

[0049] In one possible implementation, the processing module is specifically used for:

[0050] In the t-th time period, for any sample server, the total probability corresponding to the sample server is obtained based on the sub-action selection probability of the sample server for each of the defense strategies.

[0051] Based on the total probability, the sub-action selection probability of the sample server for each of the defense strategies is normalized to obtain the normalized probability of the sample server selecting sub-actions for each of the defense strategies.

[0052] Based on the normalized probabilities of sub-action selection for each of the aforementioned sample servers in response to each of the aforementioned defense strategies, the set of action selection probabilities corresponding to the t-th time period is obtained.

[0053] In one possible implementation, the processing module is specifically used for:

[0054] Based on the action selection probability set corresponding to the t-th time period, determine the sub-actions that the multiple sample servers will choose to execute for each of the defense strategies in the (t+1)-th time period;

[0055] Based on the sub-actions selected by the multiple sample servers for each of the defense strategies during the (t+1)th time period, the multiple sample servers are instructed to update the deployed defense strategies.

[0056] Based on the updated defense strategies deployed by the multiple edge proxy servers, the sample defense status data of the multiple sample servers in the (t+1)th time period is determined.

[0057] In one possible implementation, the processing module is specifically used for:

[0058] Determine the maximum number of defense strategies that can be deployed on each of the sample servers;

[0059] Based on the action selection probability set corresponding to the t-th time period and the maximum number of defense strategies that can be deployed on each of the sample servers, the sub-actions to be executed by the multiple sample servers for each of the defense strategies in the (t+1)-th time period are determined.

[0060] In one possible implementation, the training module is specifically used for:

[0061] For any given time period, based on the sample defense status data for that time period, obtain the intrinsic security reward parameters of the defense strategy corresponding to that time period;

[0062] Based on the sample defense status data for each time period, the corresponding intrinsic security reward parameters of the defense strategy, the corresponding action selection probability set, the corresponding state value, and the sub-actions selected and executed by the multiple sample servers for each defense strategy, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

[0063] Fourthly, this application provides a defense strategy model training device, comprising:

[0064] The acquisition module is used to acquire sample defense status data of multiple sample servers in the first time period. For each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in the preset defense strategy set during the first time period.

[0065] The processing module is used to process the sample defense status data of the first time period according to the defense strategy model to obtain the action selection probability set corresponding to multiple time periods and the status value corresponding to multiple time periods respectively. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability of the sample server selecting to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0066] The training module is used to adjust the parameters of the defense strategy model based on the action selection probability set corresponding to the multiple time periods and the state value corresponding to the multiple time periods, so as to obtain a trained defense strategy model.

[0067] Fifthly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the defense strategy model training method as described in any of the first aspects, or the processor implements the defense strategy determination method as described in the second aspect.

[0068] Sixthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the defense strategy model training method as described in any of the first aspects, or, when the computer program is executed by a processor, it implements the defense strategy determination method as described in the second aspect.

[0069] The defense strategy model training method, defense strategy determination method, and device provided in this application first acquire sample defense status data of multiple sample servers in a first time period. For each sample server, the sample defense status data in the first time period indicates whether the sample server has deployed each defense strategy in the preset defense strategy set within the first time period. Then, according to the defense strategy model, the sample defense status data of the first time period is processed to obtain action selection probability sets and state values ​​corresponding to multiple time periods. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to that time period indicates the probability of the sample server selecting to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy. Based on the action selection probability sets and state values ​​corresponding to multiple time periods, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model. The defense strategy model can determine the probability of each sample server selecting to deploy a defense strategy, thereby more rationally deploying corresponding defense strategies based on the characteristics of the server itself to prevent network attacks, improving the efficiency and rationality of defense strategy deployment, and achieving good defense performance against intelligent attack erosion. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;

[0072] Figure 2 A flowchart illustrating the defense strategy model training method provided in this application embodiment;

[0073] Figure 3 This is a schematic diagram of the internal processing of the defense strategy model provided in the embodiments of this application;

[0074] Figure 4 A schematic diagram illustrating the processing flow of the defense strategy model provided in this application embodiment;

[0075] Figure 5 A flowchart illustrating the defense strategy determination method provided in this application embodiment;

[0076] Figure 6 This is a schematic diagram of the structure of the defense strategy model training device provided in the embodiments of this application;

[0077] Figure 7 This is a schematic diagram of the defense strategy determination device provided in the embodiments of this application;

[0078] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0079] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0080] With the rapid development of the latest communication technologies such as 5G high-speed networks and ubiquitous mobile communications, network traffic has seen a significant increase in both quantity and complexity. At the same time, traffic security is being threatened by a constant stream of malicious attack patterns (for example, the number of distributed denial-of-service attacks worldwide is projected to double to 15.4 million by 2023).

[0081] A traffic-oriented attack can be roughly divided into three steps:

[0082] (1) Information gathering.

[0083] The information gathering process refers to the process of collecting data or information about the server to be attacked. For example, it can be used to obtain the server's network data, IP address, location information, data on interactions with other devices, and so on.

[0084] (2) Generate a report.

[0085] In step (1), the relevant information of the server to be attacked has been collected. Based on the collected information, the server to be attacked can be parsed and analyzed to generate a corresponding report. The report may include the information collected in step (1), as well as the results of the analysis of the information, such as the attributes and weaknesses of the server to be attacked.

[0086] (3) Generate attack traffic.

[0087] After the report is generated, attack traffic can be generated based on it. The report includes information such as the attributes and weaknesses of the server to be attacked, so corresponding attack traffic can be generated based on this information in the report.

[0088] The generated attack traffic can be sent to the target server to launch a cyberattack. For example, if the report includes vulnerabilities in the target server, such as its weak defenses against a certain type of cyberattack, then attack traffic corresponding to that type of cyberattack can be generated to launch such a cyberattack against the target server.

[0089] Currently, defense measures against cyberattacks include static defense methods. Static defense methods primarily rely on predefined identification criteria to identify and defend against cyberattacks. While they offer high detection accuracy and efficiency, their over-reliance on these criteria only mitigates the harm caused by malicious traffic. Therefore, the attack capabilities of a persistent attacker (i.e., the device initiating the cyberattack) in the aforementioned three steps are fully countered by reactive, passive defense mechanisms.

[0090] Addressing the shortcomings of static defense methods, network mobile target defense was proposed. Network mobile target defense proactively hinders the information gathering process from the outset, further preventing traffic-oriented attacks.

[0091] Specifically, mobile target defense is based on attack surface theory. It doesn't merely maintain improved malicious matching rules for traffic, but rather moves the attack surface of the target system, thus theoretically rendering attacks ineffective. More importantly, the defender periodically transfers users between proxies and eliminates stealth spies based on long-term malicious assessments of each user.

[0092] However, the emergence of sophisticated attacks has exposed other shortcomings in network mobility defense, which are summarized as inherent security flaws. Most existing network mobility defense schemes rely on fixed or optimal malicious assessment formulas to ensure satisfactory defense performance. Sophisticated attackers can derive assessment criteria by integrating intelligent technologies such as deep learning. Then, malicious attack traffic may be altered during its generation, degrading or even disabling the entire network mobility defense system. In short, due to a lack of consideration for its own security, existing network mobility defense methods are vulnerable to sophisticated attacks. Therefore, strengthening the inherent security of network mobility defense mechanisms is essential.

[0093] Based on this, embodiments of this application provide a defense scheme against network attacks to improve the inherent security of network mobile target defense mechanisms. Firstly, in conjunction with... Figure 1 This application describes an applicable scenario based on an embodiment of the present application.

[0094] Figure 1 This is a schematic diagram of an application scenario provided in the embodiments of this application, such as... Figure 1As shown, it includes a main server, a security center server, multiple edge proxy servers, multiple clients, and attack devices.

[0095] The master server is used to select multiple edge proxy servers to act as proxy servers. Taking a specific business service provided by the master server as an example, to reduce the risk of network attacks, the master server can deploy this business service across these multiple edge proxy servers. The business service deployed on these edge proxy servers is identical. A client may access this business service from any of these edge proxy servers.

[0096] The attacking device collects relevant information about the edge proxy server by accessing business services deployed on the edge proxy server through multiple clients, generates reports, and then uses these reports to generate attack traffic to attack the edge proxy server. The security center server can deploy corresponding defense strategies for these multiple edge proxy servers to reduce the risk of being attacked by the network.

[0097] The following is based on Figure 1 The application scenarios are illustrated to introduce the solutions of the embodiments of this application.

[0098] Figure 2 This is a flowchart illustrating the defense strategy model training method provided in an embodiment of this application. The method is applied to a first device, such as... Figure 2 As shown, it includes:

[0099] S21, obtain sample defense status data of multiple sample servers in the first time period. For each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in the preset defense strategy set in the first time period.

[0100] The multiple sample servers in this application embodiment are servers that may be subject to network attacks, for example, they can be... Figure 1 The example shown is an edge proxy server, or other devices that may be vulnerable to cyberattacks.

[0101] For any given sample server, one or more defense strategies may be deployed on that sample server. These one or more defense strategies belong to the set of preset defense strategies.

[0102] In this embodiment, the defense strategies included in the preset defense strategy set can be obtained based on relevant data from historical network attacks. Specifically, for a particular attacked device, when the device is subjected to a network attack, it will take corresponding defensive actions against the network attack. During this process, the device will generate corresponding data, and the defense strategy against the network attack can be obtained by collecting the generated data.

[0103] Specifically, the first step is to acquire a dataset representing real-world network attack traffic. This dataset has a suitable size, potentially including multiple rows and columns. Each row corresponds to a single network request, i.e., the network response traffic. Each column contains data related to the network traffic within that request, such as the source IP address, destination IP address, source port information, destination port information, data request protocol, packet size, etc. For each network request, a corresponding classification label is also included. This label indicates whether the network request is legitimate or an attack request.

[0104] The defense strategy refers to a subset of data in the dataset and its corresponding classification labels. This subset of data can be used as samples to train a neural network model. By training the model with this subset of data and its corresponding classification labels, the model can determine whether a network request is a legitimate request or a network attack request.

[0105] For example, this dataset contains 100,000 rows, corresponding to 100,000 network requests, and 30 columns, each corresponding to different types of data in the network requests. Selecting 10,000 rows as features to train the neural network model is one defense strategy. Training the neural network model based on different features in the dataset will result in different trained neural network models, and therefore different corresponding defense strategies.

[0106] In this embodiment of the application, a publicly available dataset can be used as the aforementioned dataset, or the aforementioned dataset can be obtained by capturing packets from real network traffic data.

[0107] Before introducing the solutions of the embodiments of this application, relevant concepts involved in the embodiments of this application will be introduced for ease of understanding.

[0108] To defend against attacks in a network in real time, time can be divided into equal time periods t∈[1,T], where t represents the t-th time period, and there are a total of T time periods, where T is a positive integer. Within each time period, a new defense strategy is deployed based on the current sample defense state data, the benefit of the new strategy is calculated, and finally, the process transitions to the next sample defense state. To capture the defense strategy that changes over time, the above process can be formulated as a Markov decision process.

[0109] First, let's introduce the state space.

[0110] The defense strategies deployed on different sample servers represent the defense state, denoted as a multidimensional vector S. t ={s 1,1 ,…,s N,M}. s i,j =1 indicates that feature j is deployed as a defense strategy on the i-th sample server, s i,j =0 indicates that feature j is not deployed as a defense strategy on the i-th sample server.

[0111] In this embodiment of the application, the network state space is represented as It consists of all possible defense states. However, due to the limited computing resources of the sample server, too many features cannot be deployed simultaneously. In this embodiment, C is used. i ∈[1,M] represents the maximum number of features that can be deployed on the i-th sample server, which is also the maximum number of defense strategies that can be deployed on the i-th sample server, thus representing the computing power of the sample server. In this embodiment, M represents the number of defense strategies included in the preset defense strategy set, and M is a positive integer.

[0112] The motion space will be introduced below.

[0113] An action is defined as performing an add, keep, or remove operation for each feature (i.e., defense strategy) of the sample server in the t-th time period, represented as a multi-dimensional vector A. t ={a 1,1 ,…,a N,M}. Where a i,j ∈{-1,0,1} (1≤i≤N,1≤j≤M) represents the action for feature j on the i-th sample server. i,j =-1 indicates that feature j, a is removed from the i-th sample server. i,j =0 indicates that feature j, a is preserved on the i-th sample server. i,j =1 indicates that feature j is added to the i-th sample server.

[0114] Representing the action space as It consists of all the actions. The action space size is 3. N×M This growth is exponential, increasing with the number of edge proxy servers and features. Besides its massive scale, the current action space also contains a large number of illegal actions, which are actions that do not meet the computational resource constraints of the sample servers or conflict with the defense status of the previous time period.

[0115] For the first time period among multiple time periods, the sample defense status data for the first time period is obtained based on whether the sample servers have deployed each defense strategy in the preset defense strategy set during the first time period.

[0116] Let there be a total of T time periods, and let S1 be the sample defense status data for the first time period, where S1 = {s 1,1 ,...,s N,M}. Where any element s i,j This indicates whether the i-th sample server has deployed feature j as a defense strategy, where i∈[1,N], j∈[1,M], M and N are both positive integers, M is the number of defense strategies included in the preset defense strategy set, and N is the number of sample servers. If s i,j =1, which means that the i-th sample server has deployed feature j as a defense strategy; if s i,j =0, which means that feature j was not deployed as a defense strategy for the i-th sample server.

[0117] S22, Based on the defense strategy model, the sample defense state data of the first time period is processed to obtain the action selection probability set corresponding to multiple time periods and the state value corresponding to multiple time periods. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability of the sample server choosing to execute sub-actions for each defense strategy. Sub-actions include deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0118] Figure 3 This is a schematic diagram of the internal processing of the defense strategy model provided in the embodiments of this application, such as... Figure 3 As shown, in one possible implementation, the sample defense state data of the t-th time period is input into the defense strategy model to obtain the action selection probability set and the state value corresponding to the t-th time period. Here, t is initially set to 1, and t is a positive integer greater than or equal to 1.

[0119] Then, based on the action selection probability set corresponding to the t-th time period, the sample defense status data of multiple sample servers in the (t+1)-th time period are determined.

[0120] If t is less than T, update t to t+1 and repeat the above operation. Until t is greater than or equal to T, obtain the action selection probability set and state value corresponding to T time periods, where T is a positive integer.

[0121] like Figure 3 As shown, the defense strategy model includes a sub-action selection network and a state value network. For any time period t, the sample defense state data for time period t are input into the sub-action selection network and the state value network for processing, respectively. The following section combines... Figure 4 This process will be described.

[0122] Figure 4 This is a schematic diagram of the processing flow of the defense strategy model provided in the embodiments of this application, as shown below. Figure 4 As shown, it includes:

[0123] S41, input the sample defense status data of the t-th time period into the sub-action selection network to obtain the sub-action selection probability of each sample server for each defense strategy in the t-th time period.

[0124] For any time period t, the first device first inputs the sample defense status data of time period t into the sub-action selection network. The sub-action selection network processes the sample defense status data of time period t and outputs the sub-action selection probability of each sample server for each defense strategy in time period t.

[0125] The probability of choosing the sub-action for any i-th sample server and selecting the j-th defense strategy can be expressed as: The row index represents the sample server's ID, and the column index represents the feature ID, which is also the ID of the defense strategy.

[0126] S42, based on the sub-action selection probability of each sample server for each defense strategy in the t-th time period, obtain the action selection probability set corresponding to the t-th time period.

[0127] Specifically, in the t-th time period, for any sample server, the probability of selecting sub-actions for each defense strategy is determined based on the sample server's sub-action selection probability. Obtain the total probability corresponding to the sample server

[0128] Then, based on the total probability The sub-action selection probabilities of the sample server for each defense strategy are normalized to obtain the normalized probabilities of the sample server's sub-action selection for each defense strategy. in:

[0129]

[0130] After obtaining the normalized probabilities of sub-action selection for each defense strategy on each sample server. Then, normalized probabilities are selected based on the sub-actions of each sample server for each defense strategy. Obtain the action selection probability set P corresponding to the t-th time period. t The elements included are

[0131] S43, input the sample defense state data of the t-th time period into the state value network to obtain the state value corresponding to the t-th time period.

[0132] For any time period t, the first device first inputs the sample defense status data of time period t into the state value network. The state value network processes the sample defense status data of time period t and outputs the state value corresponding to time period t. The state value corresponding to time period t is used to evaluate the sample defense status data of time period t and evaluate its rationality.

[0133] After obtaining the action selection probability set corresponding to the t-th time period, the sample defense status data of multiple sample servers in the (t+1)-th time period can be determined based on the action selection probability set corresponding to the t-th time period.

[0134] Specifically, firstly, based on the action selection probability set corresponding to the t-th time period, the sub-actions to be executed by multiple sample servers for each defense strategy in the (t+1)-th time period are determined.

[0135] In one possible implementation, the maximum number of defense strategies that can be deployed on each sample server is first determined. Since the computing resources of each sample server are fixed, the maximum number of defense strategies that can be deployed on each sample server is also fixed. In this embodiment, for any i-th sample server, C is used. i C represents the maximum number of defense strategies that can be deployed on the i-th sample server. i ∈[1,M].

[0136] After determining the maximum number of defense strategies that can be deployed on each sample server, the probability set P is selected based on the action corresponding to the t-th time period. t And the maximum number C of defense strategies that can be deployed on each sample server. i Determine the sub-actions that multiple sample servers will choose to execute in response to each defense strategy during the (t+1)th time period. This indicates that in the (t+1)th time period, the i-th sample server selects to execute a sub-action targeting the j-th defense strategy. This indicates that during the (t+1)th time period, the i-th sample server chooses to execute the sub-action for the j-th defense strategy, which is to deploy the j-th defense strategy; This indicates that during the (t+1)th time period, the i-th sample server chooses to execute the sub-action for the j-th defense strategy as not deploying the j-th defense strategy.

[0137] Specifically, after obtaining the action selection probability set P corresponding to the t-th time period... t Then, for the i-th sample server, select the probability set P based on the action corresponding to the t-th time period. t, Randomly select C from M defense strategies. i Each defense strategy is used as the defense strategy for the i-th sample server, until all sample servers have been selected. After all sample servers have been selected, the set of sub-actions A for the (t+1)-th time period can be obtained based on the sub-actions executed by each sample server for each defense strategy. t+1 .

[0138] S23. Based on the action selection probability set corresponding to multiple time periods and the state value corresponding to multiple time periods, the parameters of the defense strategy model are adjusted to obtain the trained defense strategy model.

[0139] For any time period t, based on the sample defense status data of time period t, the intrinsic security reward parameter of the defense strategy corresponding to time period t can be obtained.

[0140] Specifically, since the defense strategies deployed on the sample server in this embodiment exist in the form of features, the significance summation of the defense strategies (i.e., features) deployed on the sample server can yield the significance parameter corresponding to the t-th time period. Furthermore, calculating the Euclidean distance of the defense strategies (i.e., features) deployed on the sample server can yield the dissimilarity parameter corresponding to the t-th time period.

[0141] Then, based on the significance parameter and the difference parameter corresponding to the t-th time period, the intrinsic security reward parameter of the defense strategy corresponding to the t-th time period can be obtained. The specific calculation process can be found in the following formula (2):

[0142]

[0143] Among them, R t Let be the intrinsic security reward parameter for the defense strategy corresponding to the t-th time period, and λ be a parameter that adjusts intrinsic security and detection performance, which is a preset value. Let be the significance parameter corresponding to the t-th time period. Let be the difference parameter corresponding to the t-th time period.

[0144] The solution in this application takes into account both malicious traffic detection and intrinsic security. In order to correctly evaluate the performance of feature combination (defense strategy), an enhanced intrinsic security criterion is proposed and incorporated into the reward function.

[0145] Based on the proposed proactive defense architecture with inherent security enhancement requirements, and taking into account both salient and differential targets, the optimization problem of network mobile target defense strategy is transformed into a reinforcement learning problem.

[0146] The proximal policy optimization algorithm was modified by proposing an action space reduction method through abstraction of the original action space. This effectively reduces the size of the action space and accelerates the algorithm's convergence speed. Furthermore, an innovative multi-sub-action selection algorithm is proposed, which better aligns with real-world system models.

[0147] After obtaining the intrinsic security reward parameter R of the defense strategy for each time period t Then, based on the sample defense status data of each time period, the corresponding intrinsic security reward parameters of the defense strategy, the corresponding action selection probability set, the corresponding state value, and the sub-actions selected and executed by multiple sample servers for each defense strategy, the parameters of the defense strategy model are adjusted to obtain the trained defense strategy model.

[0148] In this embodiment of the application, it is necessary to update the sub-action. The parameters, but the output A of the selected action. t It is a set of sub-actions. In the classic proximal policy optimization algorithm, after executing each action A... t After that, the system will (S) t A t ,p(A t |S t ),V μ (S t ),R t ) is stored in memory for use in the next update, where V μ (S t ) is the output of the state-value network. Therefore, set A t The reward cannot be used for sub-actions. The update of the sub-action's reward is addressed in this embodiment by proportional allocation. and the output of the state value network The distribution is based on the probability of selecting each sub-action, as shown in equations (3) and (4) below:

[0149]

[0150]

[0151] Before training, the weights and experience replay regions of the sub-action selection network and the state-value network are initialized. In each round k∈{1,2,…,K}, the state S0 is first reset to prepare for the start of this round. In each time interval t∈0,1,…,T-1, based on the current state S... t The sub-action set A is obtained through sub-action selection. t The probability of selecting each sub-action The state value V output by the commentator network μ (S t ) Execute each sub-action New sample defense state data S is obtained t+1 And the intrinsic security reward parameter R of the defense strategy t For each sub-action Calculate the sub-actions based on probability ratios. Rewards and value Will The data is stored in the experience replay area B. Then, B samples are taken from B to update the gradients of the sub-action selection network and the state value network, with each sample taking U samples in the form of... t ,a t ,p t ,v t ,r t The tuple is defined by >. Finally, B is cleared. In each round of training, the parameters of the sub-action selection network and the state value network are updated using the loss function.

[0152] The combination of reinforcement learning and hierarchical network traffic mobile target defense has significant research value. In this embodiment, an intrinsic security standard and quantification model for the hierarchical network traffic mobile target defense mechanism are first constructed, balancing the salience and variability of proxy node defense feature allocation to select the optimal defense strategy. Then, a reinforcement learning algorithm is applied to hierarchical network traffic mobile target defense to maximize its performance. Furthermore, within an actor-critic framework, action space reduction and multi-sub-action selection mechanisms are implemented for the near-end policy optimization algorithm, thereby improving the learning effect.

[0153] Figure 5 This is a flowchart illustrating the defense strategy determination method provided in an embodiment of this application. The method is applied to a second device, such as... Figure 5 As shown, it includes:

[0154] S51: Obtain the defense status data of multiple servers in the current time period. For each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in the preset defense strategy set in the current time period.

[0155] ​In this embodiment, the second device and the first device can be the same device or different devices. Multiple servers are servers that may be vulnerable to network attacks; for example, they can be... Figure 1 The example shown is an edge proxy server, or other devices that may be vulnerable to cyberattacks.

[0156] For any given server, one or more defense policies may be deployed on that server. These one or more defense policies belong to the set of preset defense policies.

[0157] Let the current time period be t, then we can use a multidimensional vector S t ={s 1,1 ,…,s N,M This represents the defense status data of multiple servers during the current time period. i,j =1 indicates that feature j is deployed as a defense strategy on the i-th server, s i,j =0 indicates that feature j is not deployed as a defense strategy on the i-th server.

[0158] S52, input the defense status data of the current time period into the defense strategy model to obtain the action selection probability set corresponding to the current time period.

[0159] The defense strategy model in this embodiment is a model trained according to the scheme described in the above embodiments. After training, the defense state data of the current time period is input into the defense strategy model. The sub-action selection network in the defense strategy model processes the defense state data of the current time period to obtain the action selection probability set corresponding to the current time period.

[0160] In this embodiment, the process of processing the defense status data of the current time period through the sub-action selection network in the defense strategy model to obtain the action selection probability set corresponding to the current time period is similar to the process of processing the defense status data of multiple sample servers in the t-th time period through the sub-action selection network in the above embodiment to obtain the action selection probability set corresponding to the t-th time period, and will not be described again here.

[0161] S53, based on the action selection probability set corresponding to the current time period, determine the sub-actions that multiple servers will choose to execute for each defense strategy in the next time period. The sub-actions include deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0162] In this embodiment, the process of determining the sub-actions to be executed by multiple servers for each defense strategy in the next time period based on the action selection probability set corresponding to the current time period is similar to the process of determining the sub-actions to be executed by multiple sample servers for each defense strategy in the (t+1)th time period based on the action selection probability set corresponding to the tth time period in the above embodiment. For the specific implementation process, please refer to the relevant description in the above embodiment, which will not be repeated here.

[0163] S54, for any server among multiple servers, sends an instruction message to the server based on the sub-actions that the server selects to execute for each defense policy in the next time period. The instruction message is used to instruct the server to execute the sub-actions selected for each defense policy.

[0164] After determining the sub-actions that multiple servers will choose to execute for each defense strategy in the next time period, corresponding instruction information is sent to each server according to the sub-actions chosen by the servers for each defense strategy in the next time period. This instructs each server to execute the sub-actions chosen for each defense strategy, thereby realizing the deployment of the defense strategy and avoiding the adverse effects of network attacks on the servers.

[0165] The defense strategy model training method provided in this application embodiment is applied to a first device. First, it acquires sample defense status data of multiple sample servers in a first time period. For each sample server, the sample defense status data in the first time period indicates whether the sample server has deployed each defense strategy in the preset defense strategy set within the first time period. Then, according to the defense strategy model, the sample defense status data of the first time period is processed to obtain action selection probability sets and state values ​​corresponding to multiple time periods. The multiple time periods include the first time period. For any time period and any sample server within the multiple time periods, the action selection probability set corresponding to that time period indicates the probability that the sample server will choose to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy. Based on the action selection probability sets and state values ​​corresponding to multiple time periods, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model. The defense strategy model can determine the probability of each sample server choosing to deploy a defense strategy, thereby more rationally combining the characteristics of the server itself to deploy corresponding defense strategies to prevent network attacks, improving the efficiency and rationality of defense strategy deployment, and achieving good defense performance against intelligent attack erosion. The scheme in this application proposes a system model for a hierarchical network traffic mobile target defense framework. It constructs the defense strategy generation problem as a Markov decision process, transforming the selection of the optimal feature combination into finding the optimal strategy for this multi-objective programming. By comprehensively considering saliency and dissimilarity objectives, the strategy optimization problem is further transformed into a reinforcement learning problem, establishing a reinforcement learning quintuple and incorporating a newly proposed enhanced intrinsic security criterion into the reward function. Finally, the near-end strategy optimization algorithm is modified, reducing the original action space and adding a sub-action selection mechanism, resulting in a "security-performance" parallel maximization algorithm that outperforms existing network mobile target defense schemes in multiple defense and intrinsic security metrics, including global dissimilarity, average defense saliency, and defense success rate.

[0166] The defense strategy model training device provided in this application is described below. The defense strategy model training device described below can be referred to in correspondence with the defense strategy model training method described above.

[0167] Figure 6 This is a schematic diagram of the structure of the defense strategy model training device provided in the embodiments of this application, as shown below. Figure 6 As shown, it includes:

[0168] The acquisition module 61 is used to acquire sample defense status data of multiple sample servers in the first time period. For each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in the preset defense strategy set during the first time period.

[0169] Processing module 62 is used to process the sample defense status data of the first time period according to the defense strategy model to obtain action selection probability sets and state values ​​corresponding to multiple time periods respectively. The multiple time periods include the first time period. For any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability of the sample server selecting to execute a sub-action for each defense strategy. The sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy.

[0170] Training module 63 is used to adjust the parameters of the defense strategy model based on the action selection probability set corresponding to the multiple time periods and the state value corresponding to the multiple time periods, so as to obtain a trained defense strategy model.

[0171] In one possible implementation, the processing module 62 is specifically used for:

[0172] The sample defense state data of the t-th time period is input into the defense strategy model to obtain the action selection probability set corresponding to the t-th time period and the state value corresponding to the t-th time period; the initial value of t is 1, and t is a positive integer greater than or equal to 1;

[0173] Based on the action selection probability set corresponding to the t-th time period, determine the sample defense status data of the multiple sample servers in the (t+1)-th time period;

[0174] Update t to t+1 and repeat the above operation until t is greater than or equal to T, to obtain the action selection probability set and state value corresponding to T time periods respectively, where T is a positive integer.

[0175] In one possible implementation, the defense strategy model includes a sub-action selection network and a state value network; the processing module 62 is specifically used for:

[0176] The sample defense status data of the t-th time period is input into the sub-action selection network to obtain the sub-action selection probability of each sample server for each defense strategy in the t-th time period;

[0177] Based on the sub-action selection probability of each sample server for each defense strategy in the t-th time period, the action selection probability set corresponding to the t-th time period is obtained;

[0178] The sample defense state data of the t-th time period is input into the state value network to obtain the state value corresponding to the t-th time period.

[0179] In one possible implementation, the processing module 62 is specifically used for:

[0180] In the t-th time period, for any sample server, the total probability corresponding to the sample server is obtained based on the sub-action selection probability of the sample server for each of the defense strategies.

[0181] Based on the total probability, the sub-action selection probability of the sample server for each of the defense strategies is normalized to obtain the normalized probability of the sample server selecting sub-actions for each of the defense strategies.

[0182] Based on the normalized probabilities of sub-action selection for each of the aforementioned sample servers in response to each of the aforementioned defense strategies, the set of action selection probabilities corresponding to the t-th time period is obtained.

[0183] In one possible implementation, the processing module 62 is specifically used for:

[0184] Based on the action selection probability set corresponding to the t-th time period, determine the sub-actions that the multiple sample servers will choose to execute for each of the defense strategies in the (t+1)-th time period;

[0185] Based on the sub-actions selected by the multiple sample servers for each of the defense strategies during the (t+1)th time period, the multiple sample servers are instructed to update the deployed defense strategies.

[0186] Based on the updated defense strategies deployed by the multiple edge proxy servers, the sample defense status data of the multiple sample servers in the (t+1)th time period is determined.

[0187] In one possible implementation, the processing module 62 is specifically used for:

[0188] Determine the maximum number of defense strategies that can be deployed on each of the sample servers;

[0189] Based on the action selection probability set corresponding to the t-th time period and the maximum number of defense strategies that can be deployed on each of the sample servers, the sub-actions to be executed by the multiple sample servers for each of the defense strategies in the (t+1)-th time period are determined.

[0190] In one possible implementation, the training module 63 is specifically used for:

[0191] For any given time period, based on the sample defense status data for that time period, obtain the intrinsic security reward parameters of the defense strategy corresponding to that time period;

[0192] Based on the sample defense status data for each time period, the corresponding intrinsic security reward parameters of the defense strategy, the corresponding action selection probability set, the corresponding state value, and the sub-actions selected and executed by the multiple sample servers for each defense strategy, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

[0193] The defense strategy determination apparatus provided in this application is described below. The defense strategy determination apparatus described below can be referred to in correspondence with the defense strategy determination method described above.

[0194] Figure 7 This is a schematic diagram of the defense strategy determination device provided in the embodiments of this application, as shown below. Figure 7 As shown, it includes:

[0195] The acquisition module 71 is used to acquire the defense status data of multiple servers in the current time period. For each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in the preset defense strategy set in the current time period.

[0196] Processing module 72 is used to input the defense status data of the current time period into the defense strategy model to obtain the action selection probability set corresponding to the current time period;

[0197] The determining module 73 is used to determine, based on the action selection probability set corresponding to the current time period, the sub-actions that the multiple servers will choose to execute for each of the defense strategies in the next time period, the sub-actions including deploying the corresponding defense strategy or not deploying the corresponding defense strategy;

[0198] The instruction module 74 is used to send instruction information to any of the plurality of servers, based on the sub-actions selected by the server for each of the defense strategies in the next time period. The instruction information is used to instruct the server to perform the sub-actions selected for each of the defense strategies.

[0199] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions from the memory 830 to execute a defense strategy model training method or a defense strategy determination method. A defense strategy model training method is applied to a first device, comprising: acquiring sample defense status data of multiple sample servers in a first time period; for each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in a preset defense strategy set during the first time period; processing the sample defense status data in the first time period according to the defense strategy model to obtain action selection probability sets and state values ​​corresponding to multiple time periods respectively, wherein the multiple time periods include the first time period; for any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability that the sample server selects to execute a sub-action for each defense strategy, wherein the sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy; and adjusting the parameters of the defense strategy model according to the action selection probability sets and state values ​​corresponding to the multiple time periods respectively to obtain a trained defense strategy model. The defense strategy determination method is applied to a second device and includes: acquiring defense status data of multiple servers in the current time period; for each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in a preset defense strategy set during the current time period; inputting the defense status data in the current time period into a defense strategy model to obtain an action selection probability set corresponding to the current time period; determining, based on the action selection probability set corresponding to the current time period, the sub-actions to be executed by the multiple servers for each defense strategy in the next time period, the sub-actions including deploying the corresponding defense strategy or not deploying the corresponding defense strategy; for any server among the multiple servers, sending indication information to the server based on the sub-actions to be executed by the server for each defense strategy in the next time period, the indication information being used to instruct the server to execute the sub-actions selected for each defense strategy.

[0200] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0201] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the defense strategy model training method or defense strategy determination method provided by the above methods. A defense strategy model training method is applied to a first device, comprising: acquiring sample defense status data of multiple sample servers in a first time period; for each sample server, the sample defense status data in the first time period is used to indicate whether the sample server has deployed each defense strategy in a preset defense strategy set during the first time period; processing the sample defense status data in the first time period according to the defense strategy model to obtain action selection probability sets and state values ​​corresponding to multiple time periods respectively, wherein the multiple time periods include the first time period; for any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability that the sample server selects to execute a sub-action for each defense strategy, wherein the sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy; and adjusting the parameters of the defense strategy model according to the action selection probability sets and state values ​​corresponding to the multiple time periods respectively to obtain a trained defense strategy model. The defense strategy determination method is applied to a second device and includes: acquiring defense status data of multiple servers in the current time period; for each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in a preset defense strategy set during the current time period; inputting the defense status data in the current time period into a defense strategy model to obtain an action selection probability set corresponding to the current time period; determining, based on the action selection probability set corresponding to the current time period, the sub-actions to be executed by the multiple servers for each defense strategy in the next time period, the sub-actions including deploying the corresponding defense strategy or not deploying the corresponding defense strategy; for any server among the multiple servers, sending indication information to the server based on the sub-actions to be executed by the server for each defense strategy in the next time period, the indication information being used to instruct the server to execute the sub-actions selected for each defense strategy.

[0202] Furthermore, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the defense strategy model training method or defense strategy determination method provided by the above methods. The defense strategy model training method is applied to a first device and includes: acquiring sample defense status data of multiple sample servers in a first time period; for each sample server, the sample defense status data of the first time period is used to indicate whether the sample server has deployed each defense strategy in a preset defense strategy set during the first time period; processing the sample defense status data of the first time period according to the defense strategy model to obtain action selection probability sets and state values ​​corresponding to multiple time periods respectively, wherein the multiple time periods include the first time period; for any time period and any sample server in the multiple time periods, the action selection probability set corresponding to the time period is used to indicate the probability that the sample server selects to execute a sub-action for each defense strategy, wherein the sub-action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy; adjusting the parameters of the defense strategy model according to the action selection probability sets and state values ​​corresponding to the multiple time periods respectively, to obtain a trained defense strategy model. The defense strategy determination method is applied to a second device and includes: acquiring defense status data of multiple servers in the current time period; for each server, the defense status data in the current time period is used to indicate whether the server has deployed each defense strategy in a preset defense strategy set during the current time period; inputting the defense status data in the current time period into a defense strategy model to obtain an action selection probability set corresponding to the current time period; determining, based on the action selection probability set corresponding to the current time period, the sub-actions to be executed by the multiple servers for each defense strategy in the next time period, the sub-actions including deploying the corresponding defense strategy or not deploying the corresponding defense strategy; for any server among the multiple servers, sending indication information to the server based on the sub-actions to be executed by the server for each defense strategy in the next time period, the indication information being used to instruct the server to execute the sub-actions selected for each defense strategy.

[0203] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0204] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for training a defense strategy model, characterized in that, Applied to a first device, the method comprises: Obtaining sample defense state data of a plurality of sample servers in a first time period, for each sample server, the sample defense state data of the first time period is used to indicate whether each defense strategy in a preset defense strategy set is deployed by the sample server within the first time period; According to the defense strategy model, the sample defense state data of the first time period is processed to obtain a set of action selection probabilities corresponding to a plurality of time periods and a state value corresponding to a plurality of time periods, the plurality of time periods include the first time period, for any time period and any sample server in the plurality of time periods, the set of action selection probabilities corresponding to the time period is used to indicate the probability of the sample server selecting to perform a sub action for each defense strategy, the sub action includes deploying the corresponding defense strategy or not deploying the corresponding defense strategy; According to the set of action selection probabilities corresponding to the plurality of time periods and the state value corresponding to the plurality of time periods, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

2. The method of claim 1, wherein, According to the defense strategy model, the sample defense state data of the first time period is processed to obtain a set of action selection probabilities corresponding to a plurality of time periods and a state value corresponding to a plurality of time periods, comprising: Inputting the sample defense state data of the tth time period into the defense strategy model to obtain the set of action selection probabilities corresponding to the tth time period and the state value corresponding to the tth time period; t is initially 1, t is a positive integer greater than or equal to 1; According to the set of action selection probabilities corresponding to the tth time period, the sample defense state data of the t+1th time period of the plurality of sample servers is determined; Updating t to t+1, and repeating the above operation until t is greater than or equal to T, to obtain a set of action selection probabilities corresponding to T time periods and a state value, T is a positive integer.

3. The method of claim 2, wherein, The defense strategy model comprises a sub action selection network and a state value network; the inputting the sample defense state data of the tth time period into the defense strategy model to obtain the set of action selection probabilities corresponding to the tth time period and the state value corresponding to the tth time period, comprising: Inputting the sample defense state data of the tth time period into the sub action selection network to obtain the sub action selection probability of each sample server in the tth time period for each defense strategy; According to the sub action selection probability of each sample server in the tth time period for each defense strategy, the set of action selection probabilities corresponding to the tth time period is obtained; Inputting the sample defense state data of the tth time period into the state value network to obtain the state value corresponding to the tth time period.

4. The method of claim 3, wherein, According to the sub action selection probability of each sample server in the tth time period for each defense strategy, the set of action selection probabilities corresponding to the tth time period is obtained, comprising: In the tth time period, for any sample server, a total probability corresponding to the sample server is obtained according to the action selection probabilities of the sample server for each defense strategy; According to the total probability, the action selection probabilities of the sample server for each defense strategy are normalized to obtain normalized action selection probabilities of the sample server for each defense strategy; According to the normalized action selection probabilities of each sample server for each defense strategy, an action selection probability set corresponding to the tth time period is obtained.

5. The method of claim 2, wherein, The determination of the sample defense state data of the plurality of sample servers in the t+1th time period according to the action selection probability set corresponding to the tth time period comprises: According to the action selection probability set corresponding to the tth time period, the sub-actions selected and executed by the plurality of sample servers for each defense strategy in the t+1th time period are determined; According to the sub-actions selected and executed by the plurality of sample servers for each defense strategy in the t+1th time period, the plurality of sample servers are instructed to update the deployed defense strategies; Based on the updated defense strategies deployed by the plurality of sample servers, the sample defense state data of the plurality of sample servers in the t+1th time period is determined.

6. The method of claim 5, wherein, The determination of the sub-actions selected and executed by the plurality of sample servers for each defense strategy in the t+1th time period according to the action selection probability set corresponding to the tth time period comprises: Determining the maximum number of deployable defense strategies on each sample server; According to the action selection probability set corresponding to the tth time period and the maximum number of deployable defense strategies on each sample server, the sub-actions selected and executed by the plurality of sample servers for each defense strategy in the t+1th time period are determined.

7. The method according to any one of claims 2 to 6, characterized in that, The adjustment of the parameters of the defense strategy model according to the action selection probability sets corresponding to the plurality of time periods and the state values corresponding to the plurality of time periods to obtain a trained defense strategy model comprises: For any time period, the endogenous security reward parameters of the defense strategy corresponding to the time period are obtained according to the sample defense state data of the time period; According to the sample defense state data of each time period, the corresponding endogenous security reward parameters of the defense strategy, the corresponding action selection probability set, the corresponding state value, and the sub-actions selected and executed by the plurality of sample servers for each defense strategy, the parameters of the defense strategy model are adjusted to obtain a trained defense strategy model.

8. A defense policy determination method characterized by comprising: The method applied to the second device comprises: Obtaining defense state data of a plurality of servers in a current time period, for each server, the defense state data of the current time period is used to indicate whether each defense strategy in a preset defense strategy set is deployed on the server in the current time period; Inputting the defense state data of the current time period into a defense strategy model to obtain an action selection probability set corresponding to the current time period, wherein the defense strategy model is a model trained according to any one of the methods of claims 1-7; determining, according to the action selection probability set corresponding to the current time period, a sub-action selected and executed by the plurality of servers for each of the defense strategies in a next time period, the sub-action including deploying or not deploying the corresponding defense strategy; for any server in the plurality of servers, sending, to the server, indication information according to the sub-action selected and executed by the server for each of the defense strategies in the next time period, the indication information being used to instruct the server to execute the sub-action selected and executed for each of the defense strategies.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the defense strategy model training method according to any one of claims 1 to 7 when executing the program, or the processor implements the defense strategy determination method according to claim 8 when executing the program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the defense strategy model training method according to any one of claims 1 to 7, or the computer program is executed by the processor to implement the defense strategy determination method according to claim 8.

Citation Information

Patent Citations

  • Depth reinforcement learning strategy optimization defense method and device based on imitation learning

    CN112884131A

  • Improved DQN fault diagnosis method and system for gas turbine rotor system

    CN115270867A