Network attack and defense strategy optimization method based on random game and Rainbow algorithm

By constructing a network attack and defense stochastic game model and combining it with the Rainbow deep Q-network algorithm, the defense strategy is optimized, which solves the problems of high computational complexity and slow strategy learning in existing technologies. It realizes the rapid generation of adaptive and accurate defense strategies and improves the real-time response capability of the network defense system.

CN120896731APending Publication Date: 2025-11-04ZHONGYUAN ENGINEERING COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511009081.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing network defense strategies are ill-equipped to deal with complex and ever-changing network attacks, especially advanced persistent threats and zero-day exploits. Traditional methods suffer from high computational complexity, slow policy learning convergence, weak generalization ability, and are unable to generate efficient, accurate, and adaptive real-time defense strategies.

Method used

A network attack and defense stochastic game model is constructed. Combined with the Rainbow deep Q network algorithm, the defense strategy is optimized through distributed Q learning, multi-step learning, priority experience replay, competitive network structure and noisy network. The Rainbow deep Q network algorithm is used for Q function learning and strategy optimization.

Benefits of technology

It enables the rapid generation of adaptive and precise defense strategies in high-dimensional, continuous, and partially observable network attack and defense scenarios, thereby improving the real-time response capability and robustness of the defense system against complex attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120896731A_ABST
    Figure CN120896731A_ABST
Patent Text Reader

Abstract

The invention provides a network attack and defense strategy optimization method based on a random game and a Rainbow algorithm. The method comprises the following steps: constructing a network attack and defense random game model; performing Q function learning and strategy optimization by using a Rainbow deep Q network algorithm; and inputting the current state features into the trained Rainbow deep Q network, calculating Q values corresponding to all defense actions, and selecting the defense action which maximizes the Q values as an optimal defense strategy in the current state. According to the method, game analysis is carried out on the earnings of players by using a Rainbow algorithm, and the two game parties obtain benefit maximization through continuous learning and strategy adjustment. In addition, the Rainbow algorithm is added, so that the defense strategy can be adaptively adjusted, and the Nash equilibrium of the two game parties can be obtained without presetting the state transition probability of the network system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network attack and defense confrontation, and particularly provides a network attack and defense strategy optimization method based on random game and Rainbow algorithm. BACKGROUND

[0002] With the rapid development of network technology, network security problems are increasingly prominent, and the complexity and concealment of network attacks are increasing. Traditional network defense mechanisms, such as firewalls, intrusion detection systems (IDS) and defense strategies based on signature matching, have been difficult to effectively cope with the increasingly complex network attacks. These traditional methods usually rely on static rule libraries or matching of known attack patterns, and cannot adapt to the changing strategies of attackers, especially when facing advanced persistent threats (APTs) and zero-day exploit, their limitations are particularly obvious. In recent years, dynamic defense strategies have gradually attracted attention. These strategies attempt to respond to the dynamic behavior of attackers by adjusting defense measures in real time. However, existing dynamic strategy decision-making methods still have significant limitations when dealing with high-dimensional, continuous, and partially observable real network attack and defense scenarios. For example, although the classic random game model can depict the antagonism and uncertainty of the attack and defense sides, it faces the problem of exponential explosion of computational complexity when solving the Nash equilibrium or optimal strategy of complex large-scale games, making it difficult to respond in real time. In addition, traditional reinforcement learning algorithms perform poorly in dealing with high-dimensional feature representation of network state, efficient exploration and utilization balance, value function estimation bias, and low sample utilization efficiency, resulting in slow convergence of strategy learning, weak generalization ability, poor robustness in dealing with complex and variable attacks, and inability to generate efficient, accurate, and adaptive real-time defense strategies. In the prior art, although some research attempts to combine game theory and reinforcement learning to solve network attack and defense problems, these researches mostly have the following deficiencies: (1) assuming that the players have complete information. In real network attack and defense events, it is difficult for the defender to obtain all the information of the attacker due to the attacker's concealment. (2) Assuming that the players are completely rational. In real network attack and defense events, both sides are difficult to achieve complete rationality as they pursue their own maximum benefit. (3) Assuming a deterministic state transition model. However, in many cases, the state transition of the network system is affected by random behavior, and its transition probability is also random.

[0003] Therefore, there is an urgent need for a network attack and defense strategy optimization method based on random game and Rainbow algorithm to solve the above problems. SUMMARY

[0004] In order to overcome the above-mentioned defects, the present application is proposed to provide a solution or partial solution to the above-mentioned problems.

[0005] The application provides a network attack and defense strategy optimization method based on a random game and a Rainbow algorithm, comprising: constructing a network attack and defense random game model, wherein the network attack and defense random game model comprises an attacker and defender set, a network system state set, an attacker and defender behavior set, an attacker type set, a defender judgment on an attacker type, an attack and defense strategy set, an immediate return, and a benefit function set; performing Q function learning and strategy optimization by using a Rainbow deep Q network algorithm, wherein the algorithm improves strategy selection in the network attack and defense random game model by fusing distributed Q learning, multi-step learning, priority experience playback, a competitive network structure, and a noise network; inputting a current state feature into the trained Rainbow deep Q network, calculating Q values corresponding to all defense actions, and selecting a defense action maximizing the Q value as an optimal defense strategy under the current state.

[0006] In one technical solution of the network attack and defense strategy optimization method based on the random game and the Rainbow algorithm, the network attack and defense random game model is a 9-tuple , wherein: N={Attacker, Defender} is an attacker and defender set, that is, a game player; S={s1, s2, …, s n} is a network system state set; A={A1, A1, …, A n} is an attacker behavior set, wherein A i =(a1, a2, …, a j ) is an action set of the attacker in the network system state s i ; D={D1, D1, …, D n} is a defender behavior set, wherein D i =(d1, d2, …, d j ) is an action set of the attacker in the network system state s i ; } is an attacker type set; P A ={p A (s i , θ1), p A (s i , θ2), …, p A (s i , θ m} represents a defender judgment set on an attacker type when the network system is in s i ; π=(π a , π d ) represents a pair of attack and defense strategy sets; wherein π a (s i , θ j )=(σ a (si ,a1,θ j ),…,σ a (s i ,a m ,θ j )) represents θ i Type of attacker in network system state s i The strategy at that time, σ a (s i ,a i ,θ j Choose behavior a for the attacker i The probability of π; similarly, π d (s i This indicates that the defender is in the s position of the network system. i Defense strategy at that time, σ d (s i ,d j Choose behavior d for the defender i The probability of R. a,d (s i ,a,d,θ j ) indicates the state s of the network system. i Attacker type θ j At that time, the immediate reward for both the attacker and the defender to take action (a,d); U=(Q,V) is the set of payoff functions for the attacker and the defender, where Q=(Q a Q d Q represents the set of state-behavior reward functions. d =(s i ,a,d,θ j ) indicates that the network system state is s during a certain phase of the attack and defense event. i Attacker type θ j At that time, the payoff function of the defender after both sides take actions (a,d); V = (V a V d V represents the set of state-policy value reward functions. d (s i ,π a (s i ,θ j ),π d (s i ),θ j ) indicates that the network system state is s during a certain phase of the attack and defense event. i Attacker type θ j At that time, both sides adopted strategies (π) a (s i ,θ j ),π d (s i The payoff function for the defender.

[0007] In one technical solution of the network attack and defense strategy optimization method based on random game theory and Rainbow algorithm, the distributed Q-learning uses a value distribution instead of a single expected Q-value; multi-step learning uses n-step rewards as the objective to accelerate learning, and considers the cumulative discount benefits of the next n steps when calculating the objective; priority experience replay assigns priority based on the absolute value of the temporal difference error of the samples in the experience replay pool; competitive network structure splits the Q-network into value streams and advantage streams, and finally merges them to calculate the Q-value; noisy network adds parameter space noise to the linear layer parameters of the Q-network, replacing the action space noise added in the traditional action selection.

[0008] In one of the technical solutions of the network attack and defense strategy optimization method based on random game theory and Rainbow algorithm, the Q-function learning and strategy optimization are achieved by constructing a Rainbow DQN agent using the Rainbow deep Q network algorithm.

[0009] In one technical solution of the network attack and defense strategy optimization method based on random game theory and Rainbow algorithm, the agent is implemented by a deep neural network structure. The network structure takes the state feature vector and the defender's action as input and outputs the predicted state-defense action value function.

[0010] In one technical solution of the network attack and defense strategy optimization method based on random game theory and Rainbow algorithm, the following training process is adopted for Q-function learning and strategy optimization using Rainbow deep Q-network algorithm: A. Initialization: Initialize online Q-network parameters, target network parameters, and empty experience replay pool; B. Interaction and experience collection: In the simulated network environment, select defensive actions based on the current strategy. Attacker actions are generated by predefined scripts, another agent, or real attackers. Execute joint action (a,d), observe the next state s′, and the defender's immediate gain R. a,d (s i ,a,d,θ j The empirical tuple (s,(a,d),s′,R) a,d C. Sampling and Learning: Sample a batch of experiences from the experience replay pool according to priority, and use the target network ω. - Calculate the target distribution for each sample; calculate the loss between the online network's predicted distribution ω and the target distribution; update the sample priorities using priority empirical replay; minimize the loss using gradient descent and update the online network parameters ω; D. Target network update: periodically update the target network parameters ω. - Update the online network parameters ω; repeat step BD until the Q network converges or meets the preset conditions.

[0011] In a technical scheme of the network attack and defense strategy optimization method based on the random game and the Rainbow algorithm, before the Rainbow deep Q network algorithm is used for Q function learning and strategy optimization, the following is further included: preprocessing and feature engineering are performed on original state information of the network system.

[0012] The network attack and defense strategy optimization method based on the random game and the Rainbow algorithm has the following beneficial effects: The method uses the Rainbow algorithm to perform game analysis on the income of players, and the two parties in the game adjust their own strategies by constantly learning to achieve the goal of maximizing benefits. In addition, the introduction of the Rainbow algorithm gives the defense strategy the ability to adaptively adjust, without needing to preset the state transition probability of the network system, the Nash equilibrium of the two parties in the game can be obtained, so that the whole game analysis and strategy optimization process is more scientific, efficient and in line with the dynamic characteristics of the actual network attack and defense scene. BRIEF DESCRIPTION OF DRAWINGS

[0013] The disclosure of the present application will become more readily understood by referring to the accompanying drawings. It will be readily understood to those skilled in the art that the drawings are only intended to illustrate the present application and are not intended as limitations on the scope of the present application. In addition, similar numbers in the figures represent similar components, wherein:

[0014] Figure 1 is a main step flow diagram of the network attack and defense strategy optimization method based on the random game and the Rainbow algorithm according to an embodiment of the present application;

[0015] Figure 2 is a business network topology diagram according to an embodiment of the present application;

[0016] Figure 3 is a network state transition diagram according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] Some embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art will understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0018] As shown in Figure 1 A network attack and defense strategy optimization method based on the random game and the Rainbow algorithm in an embodiment of the present application mainly includes the following steps S1-S3.

[0019] Step S1, a network attack and defense stochastic game model is constructed, the network attack and defense stochastic game model includes a set of attackers and defenders, a set of states of a network system, a set of behaviors of the attackers and the defenders, a set of types of the attackers, a judgment of the defenders on the types of the attackers, a set of attack and defense strategies, an immediate return, and a set of revenue functions;

[0020] Step S2, Q function learning and strategy optimization are performed by using a Rainbow deep Q network algorithm, the algorithm improves the strategy selection in the network attack and defense stochastic game model by fusing distributed Q learning, multi-step learning, priority experience replay, a competitive network structure, and a noise network;

[0021] Step S3, for a given current state feature φ(s), φ(s) is input into the trained Rainbow Q network. Q values corresponding to all defense actions D i =(d1,d2,…,d j ) are calculated (or expected Q values are calculated from a value distribution). A defense action that maximizes the Q value is selected as the optimal defense strategy under the current state. The selected optimal defense action is converted into specific and executable defense instructions (such as calling an API to modify a firewall rule or sending a host isolation command), and acts on an actual network defense system.

[0022] In this embodiment, distributed Q learning: the traditional Q learning process is distributed to multiple agents or nodes for processing, which can accelerate the learning progress, improve the learning efficiency, and enable the model to process more data samples in a shorter time, thereby converging to a better strategy effect more quickly.

[0023] Multi-step learning: instead of considering only the reward and state change of the current step to update the Q value, the cumulative reward of subsequent steps is considered, which makes the strategy selection more forward-looking and more suitable for the network attack and defense scenario with certain time sequence and long-term goal orientation.

[0024] Priority experience replay: in the reinforcement learning process, the past experience replay mechanism randomly extracts samples from the stored experience for learning, while the priority experience replay gives different priorities according to the importance of the experience (for example, those with large Q value changes or special key state transitions), and preferentially enables the model to learn and utilize these more important experiences, thereby improving the learning effect and enabling the model to master key attack and defense strategy knowledge more quickly.

[0025] Competitive network structure: Introducing a competitive mechanism into the network structure can allow the model to better weigh and compare different strategy options during the learning process, promote competition between different strategies, and thus filter out better strategies, enhancing the model's adaptability and decision-making ability in complex and changing network attack and defense environments.

[0026] Noise network: Introducing noise into the neural network can increase the model's exploration ability, avoid the model from prematurely falling into a local optimal strategy, and enable the model to more fully explore various possible strategy spaces during the learning process, which is helpful for discovering more global and superior strategies and for dealing with the characteristics of constantly changing attacker and defender strategies in network attack and defense.

[0027] In one embodiment, the network attack and defense stochastic game model is a 9-tuple wherein:

[0028] N = {Attacker, Defender} is the set of attackers and defenders, i.e., the game players;

[0029] S = {s1, s2, …, s n} is the set of network system states;

[0030] A = {A1, A1, …, A n} is the set of attacker actions, where A i = (a1, a2, …, a j ) is the set of attacker actions in the network system state s i ;

[0031] D = {D1, D1, …, D n} is the set of defender actions, where D i = (d1, d2, …, d j ) is the set of defender actions in the network system state s i ;

[0032] is the set of attacker types;

[0033] P A = {p A (s i , θ1), p A (s i , θ2), …, p A (s i , θ m} represents the set of judgments made by the defender on the attacker types when the network system is in s i ;

[0034] π = (π a , πd ) represents a pair of attack-defense strategy sets; where, π a (s i ,θ j ) = (σ a (s i ,a1,θ j ),…,σ a (s i ,a m ,θ j )) represents the strategy of the attacker of type θ i at the network system state s i , σ a (s i ,a i ,θ j ) is the probability of the attacker choosing action a i ; similarly, π d (s i ) represents the defense strategy of the defender at the network system state s i , σ d (s i ,d j ) is the probability of the defender choosing action d i .

[0035] R a,d (s i ,a,d,θ j ) represents the immediate reward of the attack-defense pair (a, d) at the network system state s i when the attacker is of type θ j ;

[0036] U = (Q, V) is the set of payoff functions of the attacker and the defender, where Q = (Q a , Q d ) represents the set of state-action payoff functions, Q d = (s i ,a,d,θ j ) represents the payoff function of the defender after the attack-defense event at a certain stage, when the network system state is s i , the attacker is of type θ j , and the pair (a, d) is taken; V = (V a , V d ) represents the set of state-strategy value payoff functions, V d (s i ,π a (s i ,θ j ),π d (s i ),θ j) represents the defender's payoff function when the network system state is s i , the attacker type is θ j , and both sides take strategies (π a (s i , θ j ), π d (s i )) after a certain stage of attack-defense events.

[0037] In one embodiment, the defender's state-strategy payoff function V d is:

[0038]

[0039] In one embodiment, a powerful Rainbow DQN agent is trained (usually representing the defender's perspective, but theoretically also applicable to the attacker's perspective), learning the optimal defense strategy (Q function) under the stochastic game model.

[0040] Network structure: build a deep neural network (usually a multi-layer perceptron MLP, or combined with CNN / RNN to process specific features) as the Q network. The network takes the state feature vector φ(s) and the defender's action d j as input, and outputs the predicted state-defense action value function Q d = (s i , a, d, θ j ).

[0041] Various improvement techniques are used to improve the learning efficiency and stability of the Rainbow DQN fusion, as follows:

[0042] Distributed Q-learning: use value distribution instead of a single expected Q value. The Q network outputs a value distribution corresponding to each defense action D i . The learning goal is to approximate the distribution of returns. This better models the uncertainty and sparsity of rewards in the environment.

[0043] Multi-step learning: use n-step returns as the target to accelerate learning and alleviate the short-sightedness of single-step updates. When calculating the target, consider the cumulative discounted returns of the next n steps.

[0044] Priority experience replay: in the experience replay pool, assign priority according to the absolute value of the sample's temporal difference error (TD-error). Preferentially replay samples with large TD-error (i.e. "unexpected" or "important" experiences), improving data utilization efficiency.

[0045] Competitive network structure: Split the Q network into value stream (V stream) and advantage stream (A stream), and finally combine to calculate Q value, which helps to learn the relative advantage of actions more stably.

[0046] Q d = (s i , a, d, θ j ) = V(s i , θ j ) + (A(s i , a, d, θ j ) - mean_a' A(s i , a, d', θ j )).

[0047] Noise network: Add parameter space noise in the linear layer parameters of the Q network, instead of the traditional action space noise (such as ε-greedy) added during action selection. This encourages the policy to explore more effectively in the parameter space.

[0048] In an embodiment, the following training process is used for Q function learning and policy optimization using the Rainbow deep Q network algorithm:

[0049] ① Initialization: Initialize the online Q network parameters ω, the target network parameters ω - = ω, and an empty experience replay pool D.

[0050] ② Interaction and experience collection: In the simulated network environment, select the defense action d based on the current policy (such as the policy combined with ε-greedy or noise network). The attacker action a can be generated by a pre-defined script, another agent (such as using other RL algorithms), or a real attacker. Perform the joint action (a, d), observe the next state s', and the defender's immediate reward R a,d (s i , a, d, θ j ). Store the experience tuple (s, (a, d), s', R a,d ) in D.

[0051] ③ Sampling and learning: Sample a batch of experiences from D according to the priority. Use the target network ω - to calculate the distribution target of each sample (combined with distributed Q learning and multi-step return calculation). Calculate the loss (such as cross-entropy or KL divergence) between the online network ω prediction distribution and the target distribution. Update the priority of the sample based on the TD-error using priority experience replay. Update the online network parameters ω by minimizing the loss through gradient descent.

[0052] ④ Target network update: Update the target network parameters ω - to the online network parameters ω periodically (such as every C steps).

[0053] Repeat: repeat steps ②-④ until the Q-network converges or reaches a preset training round / performance index.

[0054] To verify the effectiveness of the method, a typical enterprise network environment is constructed for experiments, and the network topology structure is shown in Figure 2 The attack is initiated by an untrusted user from the external network, and the network administrator protects the security of the internal terminal through certain defense strategies. Due to the existence of the firewall, the external user can only access the data resources of the internal network through the Web server.

[0055] In the network attack and defense event, the attacker launches an attack on the target network. Once the attack is successful, the attacker will obtain the permission of the attacked component, and the network state will also change immediately. The attacker host is A, the Web server is W, the bastion host is H, the file server is F, and the database server is D. The permissions are divided into three types: none, user, and root. The network state set is S={s1, s2, s3, s4} in this experiment, and the attacker has different permissions in different states. The specific meaning is shown in Table 1, and the state transition diagram is shown in Figure 3 .

[0056] Table 1 Network state

[0057]

[0058] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the accompanying drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the original technical features without deviating from the principles of the present application, and the technical solutions after these changes or replacements will all fall within the protection scope of the present application.

Claims

1. A network attack and defense strategy optimization method based on random game theory and Rainbow algorithm, characterized in that, include: A network attack and defense stochastic game model is constructed, which includes a set of attackers and defenders, a set of network system states, a set of attacker and defender behaviors, a set of attacker types, a defender's judgment of attacker types, a set of attack and defense strategies, and a set of immediate rewards and payoff functions. The Rainbow deep Q-network algorithm is used for Q-function learning and policy optimization. The Rainbow deep Q-network algorithm improves the policy selection in the network attack and defense stochastic game model by integrating distributed Q-learning, multi-step learning, priority experience replay, competitive network structure, and noisy network. The current state features are input into the trained Rainbow deep Q network, the Q values ​​corresponding to all defensive actions are calculated, and the defensive action that maximizes the Q value is selected as the optimal defensive strategy for the current state.

2. The method according to claim 1, characterized in that, The network attack and defense random game model is a 9-tuple II-ADSGM={N,S,A,D,θ,P} A ,π,R,U}, where: N = {Attacker, Defender} is the set of attackers and defenders, i.e., the players in the game; S = {s1, s2, ..., s} n } represents the set of states of the network system; A = {A1, A1, ..., A1} n Let} be the set of attacker behaviors, where A i =(a1,a2,…,a j ) represents the attacker's state in the network system. i The set of behaviors at time; D = {D1, D2, ..., D} n Let} be the set of actions of the defender, where D i =(d1,d2,…,d j ) represents the attacker's state in the network system. i The set of behaviors at time; A collection of attacker types; P A ={p A (s i ,θ1),p A (s i ,θ2),…,p A (s i ,θ m )} indicates that the network system is in state s i At that time, the set of judgments the defender makes regarding the attacker's type; π=(π a ,π d ) represents a set of offensive and defensive strategies; where π a (s i ,θ j )=(σ a (s i ,a1,θ j ),…,σ a (s i ,a m ,θ j )) represents θ i Type of attacker in network system state s i The strategy at that time, σ a (s i ,a i ,θ j Choose behavior a for the attacker i The probability of π; similarly, the probability of π. d (s i This indicates that the defender is in the s position of the network system. i Defense strategy at that time, σ d (s i ,d j Choose behavior d for the defender i The probability of; R a,d (s i ,a,d,θ j ) indicates the state s of the network system. i Attacker type θ j At that time, the immediate feedback from both the attacker and defender on the actions (a,d); U = (Q, V) is the set of payoff functions for the attacker and the defender, where Q = (Q a Q d Q represents the set of state-behavior reward functions. d =(s i ,a,d,θ j ) indicates that the network system state is s during a certain phase of the attack and defense event. i Attacker type θ j At that time, the payoff function of the defender after both sides take actions (a,d); V = (V a V d V represents the set of state-policy value reward functions. d (s i ,π a (s i ,θ j ),π d (s i ),θ j ) indicates that the network system state is s during a certain phase of the attack and defense event. i Attacker type θ j At that time, both sides adopted strategies (π) a (s i ,θ j ),π d (s i The payoff function for the defender.

3. The method according to claim 1, characterized in that, The distributed Q-learning: uses a value distribution instead of a single expected Q value; Multi-step learning: Use n-step rewards as the goal to accelerate learning, and consider the cumulative discount benefits of the next n steps when calculating the goal; Priority Experience Replay: In the experience replay pool, samples are assigned priorities based on the absolute value of their temporal difference error. Competitive network structure: The Q network is split into value streams and advantage streams, and then the Q value is calculated by combining them. Noisy network: Parameter space noise is added to the linear layer parameters of the Q network, instead of the action space noise added during action selection in the traditional way.

4. The method according to claim 3, characterized in that, Defender's state-policy payoff function V d for:

5. The method according to claim 1, characterized in that, The Q-function learning and policy optimization using the Rainbow Deep Q-Network algorithm are achieved by constructing a Rainbow DQN agent.

6. The method according to claim 5, characterized in that, The agent is implemented using a deep neural network structure, which takes the state feature vector and the defender's actions as inputs and outputs a predicted state-defense action value function.

7. The method according to claim 5, characterized in that, The following training process is used for Q-function learning and policy optimization using the Rainbow deep Q-network algorithm: A. Initialization: Initialize online Q network parameters, target network parameters, and empty experience replay pool; B. Interaction and Experience Gathering: In a simulated network environment, the defender selects defensive actions based on the current strategy. Attacker actions are generated by predefined scripts, another agent, or a real attacker, executing a joint action (a,d), observing the next state s′, and receiving an immediate benefit R. a,d (s i ,a,d,θ j ), and the empirical tuple (s,(a,d),s′,R) a,d Stored in the experience replay pool; C. Sampling and Learning: A batch of experiences is sampled from the experience replay pool according to priority, and then used by the target network ω. - Calculate the target distribution for each sample; calculate the loss between the online network's predicted distribution ω and the target distribution; Update the priority of samples using priority experience replay; update the online network parameters ω by minimizing the loss through gradient descent; D. Target network update: Periodically update the target network parameters ω - Updated to online network parameter ω; Repeat step BD until the Q network converges or the preset conditions are met.

8. The method according to claim 1, characterized in that, Before using the Rainbow deep Q-network algorithm for Q-function learning and policy optimization, the following steps are required: preprocessing and feature engineering of the original state information of the network system.