Adaptive congestion control method and device based on multi-target reinforcement learning, and medium
Through an adaptive congestion control method based on multi-objective reinforcement learning, using the Actor-Critic network and environment-adaptive preference selection, the problem that the congestion control algorithm in the existing technology cannot adapt to the dynamically changing environment is solved, and efficient and timely transmission of information networks is achieved.
Patent Information
- Application Number
- CN202510963684.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-03
AI Technical Summary
Existing congestion control algorithms cannot adapt to the highly dynamically changing information communication environment, resulting in poor or even ineffective congestion control.
An adaptive congestion control method based on multi-objective reinforcement learning is adopted. By establishing a multi-objective Markov decision process (MOMDP) with delayed actions, the optimal control strategy of the intelligent agent is trained using the Actor-Critic network. Combined with the parallel multi-objective reinforcement learning training framework and the environment-adaptive preference selection method, it dynamically identifies changes in the information communication environment and selects appropriate preferences to achieve efficient and timely network information transmission.
It achieves adaptive and accurate perception of information networks in different network environments, ensures the efficiency and timeliness of end-to-end network information transmission, and avoids congestion problems in information transmission links.
Smart Images

Figure CN120750853A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of congestion control, and in particular relates to an adaptive congestion control method, device and medium based on multi-objective reinforcement learning. Background Art
[0002] Due to the unique network environment, end-to-end transmission delays in different regions are significant, and there are many demand terminals and few available nodes. When multiple demand terminals simultaneously send communication requests to the other end node, the transmission process is prone to congestion, resulting in information transmission failure.
[0003] Existing congestion control algorithms are often formulated for common characteristics in a wide range of scenarios. The control mode is fixed and single, and cannot be applied to highly dynamically changing information and communication environments, resulting in poor or ineffective congestion control. Summary of the Invention
[0004] The purpose of this invention is to provide an adaptive congestion control method, device, and medium based on multi-objective reinforcement learning, providing a flexible and dynamic flow control solution for the dual-end full-effect connection migration process, ensuring the efficiency and timeliness of end-to-end network information transmission. The technical solutions adopted are:
[0005] An adaptive congestion control method based on multi-objective reinforcement learning includes the following steps:
[0006] Establish a multi-objective Markov decision process MOMDP for delayed action: introduce the preference space Ω and preference function f in the Markov decision process Ω ;f Ω Used to convert the selected indicator preference w∈Ω into an indicator scalar; the indicator preference represents the weight vector of the indicator;
[0007] Model the congestion control problem as a MOMDP;
[0008] Based on the reinforcement learning algorithm, the Actor-Critic network is trained. The strategy learned by the trained Actor network is the optimal control strategy of the intelligent agent; among them, the Actor-Critic network takes state and indicator preferences as input.
[0009] Preferably, the preference function f Ω Any preference function in is:
[0010]
[0011] Among them, w represents the indicator preference;
[0012] w T represents the transpose of w;
[0013] R represents the multi-objective reward function;
[0014] S represents the state space of MOMDP; s represents the historical state;
[0015] A represents the action space of MOMDP; a represents the historical action;
[0016] f w represents the preference function.
[0017] Preferably, the indicator preference is specifically:
[0018] w=[w t ,w d ];
[0019] Among them, w t The weight vector representing the throughput, w t A weight vector representing the delay.
[0020] Preferably, the expected return relationship between the strategy π and the given preference w∈Ω is:
[0021]
[0022] in, represents the relationship between the expected return of strategy π and the given preference w∈Ω;
[0023] expresses the expectation of return;
[0024] γ represents the discount factor of MOMDP;
[0025] t represents the monitoring time;
[0026] T represents the length of the monitoring time interval;
[0027] s t ~P represents state s t is from sample P;
[0028] a t ~π represents the action a selected according to the strategy π t .
[0029] Preferably, the structure of the Actor-Critic network is:
[0030] The Actor network and the Critic network share three fully connected layers. Each connection layer uses ReLU to extract features from the original input. The extracted features are connected with the indicator scalar and input into the subsequent fully connected layers of the Actor network and the subsequent fully connected layers of the Critic network respectively.
[0031] Preferably, the updating method during training is:
[0032] Stochastic gradient dθ based on network parameters π Update the Actor network; θ π The network parameters of the Actor network;
[0033] Stochastic gradient dθ based on network parameters v Update the Critic network; θ v are the network parameters of the Critic network; v represents the utility.
[0034] Preferably, dθ is calculated π The specific steps include:
[0035]
[0036] Among them, V(s j ,w j θ v ) represents the baseline for computing advantage;
[0037] V ij represents the estimated optimal value;
[0038] N w Batch size indicating sampling preference;
[0039] N τ Indicates the batch size of the sampling transformation;
[0040] V(s j+1 ,w';θ) represents the baseline at the next moment;
[0041] θ represents the network parameters;
[0042] w′ represents the current indicator preference;
[0043] s j+1 represents the j+1th state;
[0044] represents the transpose of the preference at time i;
[0045] γ represents the discount factor of MOMDP;
[0046] r j represents the reward vector at time j;
[0047] w j represents the preference at time j;
[0048] τ represents sampling transformation;
[0049] i represents the i-th moment;
[0050] j represents the jth moment;
[0051] w represents the historical indicator preference;
[0052] W represents the preference space.
[0053] Preferably, dθ is calculated v The specific steps include:
[0054]
[0055] where Q(·;θ) represents the multi-objective Q-value function parameterized by θ;
[0056] E stands for expectation;
[0057] Q π The multi-objective Q-value function representing the π strategy;
[0058] s′ represents the current state;
[0059] a′ represents the current action taken;
[0060] w indicates preference;
[0061] γ represents the discount factor of MOMDP;
[0062] y represents the target of a given change (s,a,s',r) in step k;
[0063] w′ represents the current indicator preference, which is obtained by sampling according to a fixed distribution; θ k are the network parameters in step k;
[0064] r represents the reward vector;
[0065] L A () represents the first loss function;
[0066] L B () represents the second loss function;
[0067] a represents historical action; s represents historical state;
[0068] w represents the historical indicator preference;
[0069] λ represents a parameter.
[0070] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the adaptive congestion control method based on multi-objective reinforcement learning is implemented.
[0071] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the adaptive congestion control method based on multi-objective reinforcement learning.
[0072] A computer program, when executed by a processor, implements the adaptive congestion control method based on multi-objective reinforcement learning.
[0073] Compared with the prior art, the advantages of the present invention are:
[0074] A multi-objective reinforcement learning agent is used to generate the optimal strategy for all possible terminal communication preferences. By building a preference adaptation model, the state sequence is used as input to dynamically identify changes in the information communication environment, and appropriate preferences are automatically selected for each communication environment to achieve adaptive and accurate perception of the information network environment, ensuring the efficiency and timeliness of end-to-end network information transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 Diagram of the multi-objective reinforcement learning training framework in the adaptive congestion control method based on multi-objective reinforcement learning. DETAILED DESCRIPTION
[0076] The following diagrams describe the adaptive congestion control method based on multi-objective reinforcement learning in more detail, illustrating preferred embodiments of the present invention. It should be understood that those skilled in the art may modify the invention described herein while still achieving the beneficial effects of the invention. Therefore, the following description should be understood as a general guide for those skilled in the art and not as a limitation of the present invention.
[0077] like Figure 1 As shown in FIG, an adaptive congestion control method based on multi-objective reinforcement learning includes the following steps:
[0078] (1) Congestion Control Strategy Based on Multi-Objective Reinforcement Learning
[0079] 1) Congestion Control Modeling
[0080] In order to apply reinforcement learning algorithms to information network congestion control, congestion control is first modeled as a multi-objective Markov decision process (MOMDP) with delayed actions.
[0081] A MOMDP can be described as a tuple (S, A, P, R, Ω, γ, f Ω ),in:
[0082] It is a continuous state space;
[0083] A is a discrete action set;
[0084] P is a Markov transition model;
[0085] R:S×A→R 2 It is a set of two reward functions, corresponding to the two indicators of congestion control (throughput and delay);
[0086] It is a preference space;
[0087] γ∈[0,1) is a discount factor;
[0088] f Ω is a set of preference functions that transforms the reward vector into a scalar according to the chosen preferences w∈Ω.
[0089] The above preference w∈Ω represents a weight vector of two values.
[0090] For example, w=[w t ,w d ] represent the weights of throughput and delay respectively.
[0091] At the same time, f Ω The function in is a linear function, that is:
[0092]
[0093] In a MOMDP, a policy π is associated with the expected reward given preferences w∈Ω, expressed as:
[0094]
[0095] Among them, s t ~P represents state s t is a sample from P, a t ~π represents the action a selected according to the strategy π t .
[0096] The goal of solving a MOMDP is to find the set of Pareto optimal strategies for all possible preferences, which is described as:
[0097]
[0098] The delayed action MOMDP used in this technology is an extended form of MOMDP, that is, the action will not take effect immediately, and its state s t+1 Not only affected by action a t+1 and state s t The impact will also be affected by historical actions and states.
[0099] Considering the characteristics of delayed action MOMDP, the state space, reward function, action set and monitoring time interval are described as:
[0100] (a) State space (S) considering delayed actions.
[0101] To accurately identify network congestion without explicit network information, the following features are selected to model the network state. To further improve the robustness and generalization of the trained model, parameters such as latency and send rate use relative values instead of absolute values.
[0102] Delay Rate The delay ratio is the ratio of the average delay in the current monitoring interval to the minimum observed average delay, and is one of the most important features for detecting congestion. Where t represents the time.
[0103] Sending rate The sending rate is the ratio of the number of sent packets to the number of acknowledged packets, which can help the agent adjust the sending rate.
[0104] Packet loss rate Affected by the ACK frame during the node flight, in order to avoid the agent from over-inferring the congestion level, add parameter.
[0105] Delayed Gradient The delay gradient is the derivative of the delay with respect to time. If The agent can then infer that the sending rate should be greater than the capacity.
[0106] Action a t Considering that the state is affected by historical actions (i.e., delayed action characteristics), the action history should be put into the state to speed up model convergence.
[0107] In order to enable the agent to more accurately capture the changing process of MOMDP, each state is taken as the history of network statistics so that the agent can capture the impact of historical actions on future states.
[0108] Using h to represent the historical length of the state, the state s t ∈S is expressed as:
[0109]
[0110] (b) Multi-objective reward function (R).
[0111] The goal of congestion control is to achieve high throughput and low queuing delay.
[0112] Accordingly, the average throughput is used t, and negative average RTT as the reward function, that is:
[0113] R(s t ,a t )=[throughput t ,-rtt t ]
[0114] (c) RTT adaptive monitoring time interval.
[0115] The state and reward values are calculated at the end of each monitoring interval.
[0116] A monitoring interval that is too short is not enough to capture the actual changes in the environment in long RTT scenarios, while setting a monitoring interval that is too long will reduce the real-time response capability to network changes in short RTT scenarios.
[0117] In order to cope with heterogeneous network environments, each action requires about 0.5 RTT t-1 To make it work, we use RTT adaptive monitoring time interval and set the length of each monitoring time interval to:
[0118]
[0119] Among them MI t represents the monitoring time interval at time t, It is the minimum observable RTT. It ensures that the monitoring interval is long enough and that the congestion control strategy can react quickly to congestion.
[0120] (d) Action set (A).
[0121] Each action is associated with a change in the sending rate. To adapt to heterogeneous environments with different bandwidths, the sending rate is adjusted proportionally and an action set consisting of seven discrete values is used:
[0122] A={÷2,÷1.3,÷1.1,×1,×1.1,×1.3,×2}
[0123] Here, "÷x" means "sending rate divided by x", and "×x" means "sending rate multiplied by x".
[0124] Since the reward function is a combination of throughput and latency, the agent will observe equally inferior rewards when the sending rate is increased to 50x or 100x the bandwidth and the buffer queue is saturated.
[0125] During the training phase, limiting the sending rate to a reasonable range will produce the same changes while reducing the CPU overhead caused by packet generation and sending. t The sending rate in is denoted as srt+1 ,Right now:
[0126] sr r+1 =min(apply(sr t ,a t ),500Mbps)
[0127] Where apply(sr t ,a t ) is the application sending rate sr t and action a t function.
[0128] 2) Parallel multi-objective reinforcement learning training framework
[0129] The goal of multi-objective reinforcement learning agents is to achieve consistently high performance across diverse preferences and environments. Because the optimal policy in each of these scenarios is unique, multi-objective reinforcement learning agents must be trained across a range of environments with varying network characteristics and input preferences. This leads to catastrophic forgetting, where knowledge from previously learned environments is suddenly destroyed by training in new environments, and prolongs training time.
[0130] In order to solve the above problems, an efficient parallel multi-objective reinforcement learning training framework is proposed, such as Figure 1 shown.
[0131] It is an extended version of the asynchronous advantage Actor-Critic framework and has the following key features.
[0132] First, its training process is asynchronous, that is, multiple agents are used to train multiple Actors in parallel and update global parameters periodically.
[0133] To address the catastrophic forgetting problem, this technique trains agents in different environments simultaneously, allowing each agent to interact with a specific type of environment (e.g., long latency, low bandwidth, high throughput weights).
[0134] Secondly, this technology is based on the Actor-Critic paradigm. During the training phase, the Actor updates the policy parameters according to the direction suggested by the Critic, while the Critic updates the value function parameters.
[0135] Furthermore, using advantage functions instead of raw action values allows for better comparison of actions in a given state, making the training process more stable.
[0136] The multi-objective reinforcement learning training framework adopted by this technology greatly improves the model learning efficiency.
[0137] On the one hand, the interaction of environmental information is strengthened through parallel training.
[0138] On the other hand, sampling efficiency is improved by adopting off-policy reinforcement learning, which is able to reuse any past events through the experience replay mechanism.
[0139] In addition, the adopted offline strategy brings more precise detection direction for sampling by following a behavioral strategy different from the target strategy, which effectively improves the detection efficiency.
[0140] like Figure 1 As shown in Figure 2, this technique uses two deep neural networks that take state s and preference w as input and 2|A|Q values as output.
[0141] The first deep neural network is the Actor Network, which outputs a probability distribution over the action space.
[0142] The second deep neural network is the Critic network, which outputs the utility value of the action.
[0143] The parameters in the two networks are denoted as θ π and θ v .
[0144] The two networks share three fully connected layers, each of which uses a rectified linear activation unit (ReLU) to extract features from the raw input.
[0145] The extracted features are then concatenated with the preference values and input into different fully connected layers for output.
[0146] like Figure 1 As shown in the figure, the Actor network receives the state data of the environment at a certain moment and generates its corresponding control action; the Critic network estimates the value function and evaluates the current strategy learned by the Actor network.
[0147] The goal of multi-agent joint training is achieved through centralized training of Critic and distributed execution of Actor.
[0148] Centralized training means using global state and action inputs to the Critic network of each agent during training. The Critic network outputs a function value to evaluate the strategy of the Actor network.
[0149] Decentralized execution means that the agent maps its observed state to actions through the Actor network.
[0150] According to the evaluation results output by the Critic network, the Actor network parameters are updated, that is, the congestion control strategy is updated, and the Critic network parameters are also updated.
[0151] 3) Agent training strategy
[0152] To train a multi-objective reinforcement learning agent with multiple reward functions, this technique updates a neural network model according to a generalized version of the Bellman equation:
[0153]
[0154] where Q π is a multi-objective Q-value function. The key idea is to use a vectorized value function to perform envelope updates.
[0155] This allows the technique to quickly combine one preference with the best rewards and trajectories that could be explored under other preferences, allowing for learning a single parameter representation of the optimal policy over the entire preference space.
[0156] Based on the above equation, calculate the target for a given change (s,a,s',r) in step k:
[0157]
[0158] Where s and s' are the historical state and current state, a is the action taken, r is the reward vector, w' is sampled from a fixed distribution, θ k are the network parameters in step k,
[0159] Q(·;θ) is a multi-objective Q-value function whose form is parameterized by θ.
[0160] Afterwards, we can use the mean squared error (MSE) between the predictions returned by the neural network and the targets:
[0161]
[0162] a) Homotopy optimization w T
[0163] Considering that the optimal boundary contains a large number of discrete solutions, and L A (·) is non-smooth, so directly optimize L A (·) Very challenging in practice.
[0164] To solve this problem, this technology uses homotopy optimization to construct an auxiliary loss function L B , used to directly optimize the normalized Q value:
[0165] L B (θ)=E s,a,w [|w T yw T Q(s,a,w;θ)|]
[0166] Therefore, the overall loss function is expressed as:
[0167] L(θ)=(1-λ)L A (θ)+λL B (θ)
[0168] Where λ will slowly increase from 0 to 1 during the training process, and the loss function will be changed from L A (·)Transfer to L B (·).
[0169] Finally, the stochastic gradient of the network parameters is calculated:
[0170]
[0171] where N w and N τ are the sampling transformation and the preferred mini-batch size, respectively,
[0172] T ij Calculated by the following formula:
[0173]
[0174] Here, V(s j ,w j θ v ) is the baseline for computing advantage, V ij is the estimated optimal value, which is calculated as follows:
[0175]
[0176] b) Early termination strategy
[0177] To improve training efficiency, it is necessary to address the needle-in-the-haystack problem, whereby during the exploration phase of training, the agent performs random walks and may become stuck in a congested state. When the sending rate exceeds the available network bandwidth and the buffer is saturated, the agent can only observe equally inferior rewards (i.e., a combination of throughput and latency). At this point, the agent provides meaningless gradients and leads to extremely low model training efficiency. To address this issue and stabilize the training process, this technique increases the training step size based on the time period index and uses a linear function to calculate the termination index for each training step:
[0178] f et (i) = δ a +(i modδ b )
[0179] where δ a is the starting step size, δ b is the step growth rate.
[0180] c) Agent training algorithm
[0181] In the adaptive congestion control strategy based on multi-objective reinforcement learning proposed in this technology, the specific agent training process is as follows:
[0182] 1) Initialize the replay buffer to an empty set and use uniform distribution for preference sampling.
[0183] 2) At the beginning of each time period, the agent sets the network parameters (θ π ,θ v ) to synchronize with global parameters.
[0184] 3) Interact with the environment to collect the running trajectory of the replay buffer.
[0185] Considering the needly-in-the-haystack problem, the step size is calculated based on the early termination strategy.
[0186] 4) Calculate the Actor network (dθ) based on sampling transformation and homology optimization π ) and Critic network (dθ v ) parameters.
[0187] 5) Based on dθ π and dθ v Asynchronously updates global parameters.
[0188] (2) Environmental Adaptive Preference Selection Method
[0189] This technology proposes a new environment-adaptive preference selection method based on the expert strategy training preference adaptation model.
[0190] When preferences are not determined by upper-layer applications, the proposed method uses state sequences as input to dynamically identify the network environment and adaptively select appropriate preferences.
[0191] 1) Expert Strategy
[0192] In the training phase, since the link capacity is a known item, an expert strategy π is constructed * , to adjust the sending rate as quickly as possible to fit the link capacity.
[0193] More specifically, assuming c is the link capacity, the action generated by the expert strategy in each monitoring time interval is described as:
[0194]
[0195] in is a binary function. If apply(sr t,a) is greater than the link capacity c, the output is 1. The action that does not make the sending rate exceed the limit will be given priority.
[0196] 2) Strategy Similarity Model
[0197] This technique further proposes a policy similarity model based on cumulative reward distribution, which quantifies the similarity between expert policies and policies given a given preference w. This model then finds the optimal preference for each training environment. This method achieves higher accuracy than conventional methods that only use expected estimation policies.
[0198] The cumulative reward of strategy π is defined as:
[0199]
[0200] Due to the transition probability, yes For a given strategy π, we further define its Markov chain as M(π) = {S, A, P, R, π}, where π represents the strategy.
[0201] but The cumulative distribution of can be expressed as:
[0202]
[0203] Assuming that the distribution obeys a multivariate Gaussian distribution G(ε;θ) with θ as the parameter, and using the Monte Carlo sampling technique to collect the data set D, the distribution parameters can be obtained as:
[0204] Θ * =arg Θ maxE D (G(ε;Θ))
[0205] Afterwards, the Kullback-Leibler divergence between the expert and the strategies generated with different preferences is used as the similarity measure, which is formulated as follows:
[0206]
[0207] Among them G * =G(ε;Θ * ) and G w =G(ε;Θ w ) is the Gaussian distribution of the expert strategy and the preference w generating strategy. And D KL A larger value means a lower similarity between the two strategies.
[0208] 3) Preference Adaptive Model
[0209] Based on the policy similarity model, this technique first finds the most appropriate preference w in each training environment, under which the reward distribution is most similar to the expert policy (i.e., D KL (G * ||G w ) minimum). Afterwards, use the indicator preference Conduct a set of experiments and collect training data sets Among them S i is the state trajectory, Is the most suitable for S i preferences.
[0210] Based on this dataset, the policy adaptation problem can be formulated as a classification problem, where each candidate preference can be considered a class. To this end, this technique uses supervised learning to train a preference adaptation model called AdaM. AdaM is a three-layer neural network. Its first two layers are long short-term memory layers, which can effectively process sequential data, and the last layer is a fully connected layer.
[0211] Among them, cross entropy loss is used to train AdaM, and the cross entropy loss is expressed as follows:
[0212]
[0213] in is the true value of training sample i, and w i is the predicted value.
[0214] After training AdaM, a preference adaptation algorithm is used to infer the most appropriate preferences. It first initializes the state sequence buffer and observes the initial state.
[0215] Afterwards, by using the indicator preferences The generated policy interacts with the environment to collect state sequences.
[0216] Finally, the optimal preferences are inferred using the trained adaptive model AdaM.
[0217] The steps of the preference adaptation algorithm are as follows:
[0218] Input: training model AdaM and π; indicator preference Sample count SC.
[0219] Output: The optimal w under the current environment.
[0220] 1) Initialize sequence buffer S = {};
[0221] 2) According to S t Observe the current state s;
[0222] 3) Get guidance strategies
[0223] 4) When k∈[1,SC], sample an action Observe a new state s'; let S = S∪s' and s = s';
[0224] 5) Return w = AdaM(S).
[0225] In summary, the adaptive congestion control strategy for information communication networks proposed in this technology is
[0226] First, by constructing a congestion control model based on multi-objective deep reinforcement learning, a flexible and dynamic flow control processing solution is provided for the dual-end full-effect connection migration process.
[0227] Then, an environmental adaptive preference selection strategy was proposed, which combined with the information network environment self-perception algorithm to provide the intelligent agent with real-time information such as the location of the peer network nodes and terminals, changes in the natural environment, etc., and can adaptively adjust the control strategy according to the information transmission status of the information communication network, effectively avoiding the congestion problem of the information transmission link and ensuring the smoothness and efficiency of information transmission in the information communication network.
[0228] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. An adaptive congestion control method based on multi-objective reinforcement learning, characterized in that: The following steps are involved: Establish a multi-objective Markov decision process MOMDP for delayed action: introduce the preference space Ω and preference function f in the Markov decision process Ω ;f Ω Used to convert the selected indicator preference w∈Ω into an indicator scalar; the indicator preference represents the weight vector of the indicator; Model the congestion control problem as a MOMDP; Based on the reinforcement learning algorithm, the Actor-Critic network is trained. The strategy learned by the trained Actor network is the optimal control strategy of the intelligent agent; among them, the Actor-Critic network takes state and indicator preferences as input.
2. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 1, characterized in that: The preference function f Ω Any preference function in is: Among them, w represents the indicator preference; w T represents the transpose of w; R represents the multi-objective reward function; S represents the state space of MOMDP; s represents the historical state; A represents the action space of MOMDP; a represents the historical action; f w represents the preference function.
3. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 1, characterized in that: The indicator preferences are specifically: in=[in t ,In d ]; Among them, w t The weight vector representing the throughput, w t A weight vector representing the delay.
4. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 3, characterized in that: The expected return relationship between the strategy π and the given preference w∈Ω is: in, represents the relationship between the expected return of strategy π and the given preference w∈Ω; expresses the expectation of return; γ represents the discount factor of MOMDP; t represents the monitoring time; T represents the length of the monitoring time interval; s t ~P represents state s t is from sample P; a t ~π represents the action a selected according to the strategy π t .
5. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 3, characterized in that: The structure of the Actor-Critic network is: The Actor network and the Critic network share three fully connected layers. Each connection layer uses ReLU to extract features from the original input. The extracted features are connected with the indicator scalar and input into the subsequent fully connected layers of the Actor network and the subsequent fully connected layers of the Critic network respectively.
6. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 1, characterized in that: The update method during training is: Stochastic gradient dθ based on network parameters π Update the Actor network; θ π The network parameters of the Actor network; Stochastic gradient dθ based on network parameters v Update the Critic network; θ v are the network parameters of the Critic network; v represents the utility.
7. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 6, characterized in that: Calculate dθ π The specific steps include: Among them, V(s j ,w j θ v ) represents the baseline for computing advantage; V ij represents the estimated optimal value; N w Batch size indicating sampling preference; N τ Indicates the batch size of the sampling transformation; V(s j+1 ,w';θ) represents the baseline at the next moment; θ represents the network parameters; w′ represents the current indicator preference; s j+1 represents the j+1th state; represents the transpose of the preference at time i; γ represents the discount factor of MOMDP; r j represents the reward vector at time j; w j represents the preference at time j; τ represents sampling transformation; i represents the i-th moment; j represents the jth moment; w represents the historical indicator preference; W represents the preference space.
8. The adaptive congestion control method based on multi-objective reinforcement learning according to claim 7, characterized in that: Calculate dθ v The specific steps include: where Q(·;θ) represents the multi-objective Q-value function parameterized by θ; E stands for expectation; Q π The multi-objective Q-value function representing the π strategy; s′ represents the current state; a′ represents the current action taken; w indicates preference; γ represents the discount factor of MOMDP; y represents the target of a given change (s,a,s',r) in step k; w′ represents the current indicator preference, which is obtained by sampling according to a fixed distribution; θ k are the network parameters in step k; r represents the reward vector; L A () represents the first loss function; L B () represents the second loss function; a represents historical action; s represents historical state; w represents the historical indicator preference; λ represents a parameter.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the adaptive congestion control method based on multi-objective reinforcement learning according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the adaptive congestion control method based on multi-objective reinforcement learning according to any one of claims 1 to 8 is implemented.