A real-time deep reinforcement learning approach
By introducing a width learning system, the problems of long training time and slow state convergence of deep reinforcement learning methods are solved, real-time decision-making and fast state convergence are achieved, and it is suitable for practical engineering applications.
Patent Information
- Application Number
- CN202411024024.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-29
AI Technical Summary
The existing deep reinforcement learning methods are difficult to adapt to the needs of actual engineering applications due to the long training time and low efficiency, especially the inability to select decision actions in real time, resulting in slow state convergence speed and difficult to adapt to the needs of actual engineering applications.
A width learning system is introduced to quickly process and learn information by extending the width of the network rather than depth. The core structure consists of an input layer, an enhancement node and an output layer. The output weight is directly solved through linear equations, improving real-timeness and reducing the complexity of iterative computing.
Real-time performance of deep reinforcement learning methods is achieved, training time is reduced, and the state is fast converged, which is suitable for practical engineering application scenarios.
Smart Images

Figure CN119005288B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deep reinforcement learning method, and in particular to a real-time deep reinforcement learning method. Background Art
[0002] With the rapid development of artificial intelligence (AI), deep reinforcement learning (RL) is becoming an effective approach for solving complex decision-making problems. Traditional deep reinforcement learning algorithms primarily rely on deep neural networks for feature extraction and policy optimization. However, deep neural networks typically require extensive training data and computing resources, resulting in long training times and low efficiency. Training time and computational efficiency, as fundamental and critical technical indicators for complex decision-making problems, have become a research hotspot due to their importance in practical applications.
[0003] In existing deep reinforcement learning methods, the output of a deep neural network is a Q-table of decision actions. This means that decision actions are selected within a limited space and have the same step size. However, it should be noted that in existing deep reinforcement learning methods, the decision actions output by the deep neural network all have the same step size, which does not guarantee rapid state convergence. This results in the decision problem not being solved in real time, posing a challenge to practical engineering applications.
[0004] On the one hand, learning from large amounts of data requires long training times, which results in online data becoming offline data. On the other hand, decision actions with the same step length within a limited space result in states converging only at a fixed step size. Therefore, for deep reinforcement learning algorithms, reducing training time and ensuring rapid state convergence are of practical engineering significance. Summary of the Invention
[0005] The purpose of the present invention is to provide a real-time deep reinforcement learning method to solve complex decision-making problems and ensure that decision actions are selected in real time to achieve rapid state convergence. Given that width learning systems can be trained quickly and have good generalization capabilities, the present invention achieves rapid information processing and learning by expanding the width rather than the depth of the network. Its core structure consists of an input layer, enhancement nodes, and an output layer. The output weights are directly solved through linear equations, thereby improving the real-time performance of deep reinforcement learning methods and reducing the complexity of iterative calculations.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A real-time deep reinforcement learning method includes the following steps:
[0008] Step 1: Estimate the mean of the decision action
[0009] The agent starts from any given initial estimated state and uses the width learning system to learn the state increment from the latest data. The specific steps are as follows:
[0010] Step 1.1, initialize the width learning system;
[0011] Step 1.2: The agent updates the width learning system using the arrival time difference and state difference of the base station signals collected online, and uses the width learning system to learn the state increment from the latest data;
[0012] Step 2: Select a decision action
[0013] The decision action is selected based on the Gaussian distribution strategy with the output vector of the width learning system as the mean and the smaller value of the output value of the double Q network as the covariance. The specific steps are as follows:
[0014] Step 2.1: The output vector of the width learning system is the decision action is regarded as the mean of the Gaussian distribution strategy, and the smaller value of the output value of the double Q network is regarded as the covariance of the Gaussian distribution strategy;
[0015] Step 2.2: To evaluate the decision action performance, define a one-step reward function for:
[0016]
[0017] in, is the vector of arrival times, Indicates decision action Time after execution, Q t and Q u is a symmetric positive definite matrix;
[0018] Step 2.3: Define the total Q function of the double Q network for:
[0019]
[0020] where γ∈(0,1) is the discount factor, and represents the estimated state of agentj at the kth iteration step and the k+1th iteration step;
[0021] Step 2.4: Learn the output vector of the system based on the value and width of the Q function From Gaussian distribution strategy Randomly select decision actions Get the time difference Reward function and Q function;
[0022] Step 2.5: Tuple Stored in the memory pool for updating the dual Q network, where l c is the total number of iteration steps;
[0023] Step 3: Update status
[0024] Step 3.1: Model the state estimation process as a Markov decision process and establish the state update process as:
[0025]
[0026] Step 3.2, until Less than Δt or Estimated status Considered as the state vector of the agent, otherwise return to step 2, every interval l e Return to step 1.2 to update the width learning system, where Δt represents the desired accuracy, is the upper bound of the iteration step.
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] 1. In response to the time-consuming process of deep neural network training in deep reinforcement learning methods, a wide learning system is introduced. When a new set of data is collected, it can quickly learn from it, ensuring that changes in the environment are learned from online data in real time, thereby selecting decision actions.
[0029] 2. In deep reinforcement learning methods, selecting actions with the same step length within a limited action space may slow down the convergence of the state. In this paper, the output of the width learning system is used as the mean of the strategy for selecting decision actions (i.e., Gaussian distribution), and the smaller value of the output of the dual Q network is used as the covariance of the Gaussian distribution. The infinite decision action space ensures that the state can converge quickly.
[0030] 3. The proposed deep reinforcement learning method, which incorporates a breadth learning system, ensures rapid learning from newly collected data. This capability ensures the proposed method's real-time performance, reducing neural network training time while significantly learning from the latest data, perceiving real-time changes in the environment, and enabling real-time adjustments to decision-making actions. This makes the proposed deep reinforcement learning method, which incorporates a breadth learning system, more suitable for practical application scenarios.
[0031] 4. Based on the proposed deep reinforcement learning method that introduces a wide learning system, this method considers the case of selecting decision actions from a Gaussian distribution strategy. By using the output of the wide learning system as the mean of the Gaussian distribution strategy and the smaller value of the dual Q network output as the covariance of the Gaussian distribution strategy, the present invention has an infinite action space and can select increments of different step sizes to ensure state convergence. The Gaussian distribution strategy significantly reduces the time it takes for the agent to converge, improving the real-time performance of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Flowchart of a real-time deep reinforcement learning method;
[0033] Figure 2 The flowchart of the real-time deep reinforcement learning method is shown in Figure 2. DETAILED DESCRIPTION
[0034] The technical solution of the present invention is further described below with reference to the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention that does not depart from the spirit and scope of the technical solution of the present invention should be included in the scope of protection of the present invention.
[0035] This invention introduces a wide learning system for deep reinforcement learning methods, enabling agents to quickly learn from collected data. Furthermore, it considers an infinite decision action space to improve the speed of agent state convergence. This invention designs a real-time deep reinforcement learning method for deep reinforcement learning methods, enabling faster resolution of complex decision problems. The method includes three main stages: estimating the mean of decision actions, selecting decision actions, and updating states. Furthermore, the symbols used in this invention are: is the randomly given initial estimated state of agentj; and represents the estimated state of agentj at the kth iteration step and the k+1th iteration step; Represents the decision action of agentj at the kth iteration step. represents the arrival time vector of the signal from the base station of the I auxiliary agentj estimated state to the estimated state of agentj in the kth iteration step, The arrival time of the signal from the base station representing the estimated state of the I-th auxiliary agent j to the estimated state of agent j in the k-th iteration step. Figure 1 and Figure 2 As shown, the specific steps include:
[0036] Step 1: Estimate the mean of the decision action
[0037] The agent starts from any given initial estimated state and iteratively estimates the actual state value by selecting decision actions. The specific steps are as follows:
[0038] Step 1.1, initialize the width learning system;
[0039] Step 1.2: The agent uses the arrival time difference and state difference of the base station signals collected online to update the width learning system, and uses the width learning system to learn the state increment from the latest data. The specific steps are as follows:
[0040] In order to achieve real-time decision-making of the agent, an incremental algorithm is used to quickly calculate the output weight of the width learning system. The arrival time difference vector of the base station signal collected is The corresponding state difference is expressed as The weight matrix from the input layer to the enhancement layer of the width learning system is The bias of the neurons in the enhancement layer is Then the output of the enhancement layer of the width learning system for:
[0041]
[0042] The present invention defines:
[0043]
[0044] and
[0045]
[0046] Among them, H is given in the initialization stage ADD,ZH,0 Then, the present invention defines:
[0047]
[0048] According to the matrix defined by formulas (1) to (4) in the present invention, it can be obtained:
[0049]
[0050] Below, the present invention defines:
[0051]
[0052] According to the matrix defined by formulas (4) to (6) in the present invention, it is defined as follows:
[0053]
[0054] Among them, λ→0 is the regularization parameter, which is given in the initialization stage Assume that express situation, assuming express In this invention, the matrix is defined as The update expression is:
[0055]
[0056] Therefore, the weight matrix from the enhancement layer to the output layer of the width learning system of the present invention is Updated to:
[0057]
[0058] Among them, the initial value W of the weight matrix from the enhancement layer to the output layer of the width learning system is given in the initialization stage on,ho,0 In the present invention, it is defined as follows:
[0059]
[0060] in, is the difference between the predicted arrival time of the base station signal at the origin of the coordinate system and the estimated arrival time in the kth iteration step, is the predicted arrival time of the base station signal at the origin of the coordinate system, is the estimated arrival time of the base station signal at the origin of the coordinate system in the kth iteration step, It is the time difference vector formed by the difference between the predicted arrival time and the estimated arrival time in the kth iteration step. Below, we define:
[0061]
[0062] According to this definition, we can get: Then, the estimated state increment is:
[0063]
[0064] in, is the state increment of agent j output by the width learning system.
[0065] Step 2: Select a decision action
[0066] If the output of the width learning system is always selected as the state increment of the agent, the estimated state value can easily fall into a local optimal value. To this end, it is necessary to randomly select a decision action related to the state increment output by the width learning system. Therefore, a Gaussian distribution strategy is adopted to select the decision action. Specifically, the output of the width learning system is regarded as the mean of the Gaussian distribution strategy, and the smaller value of the output value of the double Q network is regarded as the covariance of the Gaussian distribution strategy. The agent's decision action (i.e., the state increment) is selected by minimizing the difference between the estimated time and the measured time to update the agent's estimated state to be close to the agent's true state value. In order to evaluate the action performance, define a one-step reward function for:
[0067]
[0068] in, is the vector of arrival times, Indicates action Time after execution, Q t and Q u is a symmetric positive definite matrix. In order to reduce the deviation, the real-time deep reinforcement learning framework proposed in this invention adopts two independent Q networks, namely: Q1 network and Q2 network. The smaller of the output values of the two Q networks is regarded as the Q value. In view of this, the total Q function is defined as for:
[0069]
[0070] Where γ∈(0,1) is the discount factor. Based on the value and width of this Q function, the output of the learning system is From Gaussian distribution strategy Randomly select decision actions Then, the time difference can be obtained Reward function and Q function. Therefore, the tuple Stored in the memory pool for updating the dual Q network, where l c is the total number of iterations. In order to make the training process of the double Q network more stable, each l e Update the network once.
[0071] Step 3: Update status
[0072] After the width learning system and the dual Q network are optimized, the output vector of the width learning system is and the smaller value associated with the output of the double Q network As Gaussian distribution strategies Based on this Gaussian distribution strategy, randomly select actions The present invention considers the case where no state information can be collected through movement in space. The state estimation process is modeled as a Markov decision process, and the state update process is established as:
[0073]
[0074] Among them, action Determined by the width learning system and the Q network, the action space is continuous. Less than Δt or Estimated status is regarded as the state vector of the agent. Otherwise, repeat the above process. Where Δt represents the desired accuracy, is the upper bound of the iteration step.
[0075] Example:
[0076] Consider a surface communication system consisting of j ships and three surface base stations. The method proposed in this invention is used to estimate the position of surface ship j. in Respectively represent the coordinates of the surface vessel on the X-axis, Y-axis, and Z-axis. The state in the present invention is the position of the surface vessel, and the decision action is the position increment of the surface vessel. Start updating the estimated position of ship j. First, before the ship goes out to sea, use a device with strong computing power to use offline data to train the width learning system and initialize the matrix H ADD,ZH,0 , and W on,ho,0 Based on this width, the system learns the time difference vector between the measured Get position increment Use this position increment as a Gaussian distribution strategy Then, the outputs of the two Q networks are compared and the smaller value is used as the covariance of the Gaussian distribution strategy. The position increment of the ship is randomly selected on the Gaussian distribution strategy to update the estimated position of the ship. Stored in the replay memory pool, it is used to update the width learning system and the two Q networks. The details of the position estimation process of ship j are as follows:
[0077] 1. Estimate the mean of decision actions
[0078] No. The time difference vector of the surface base station signal collected times arriving at ship j is: The corresponding position coordinate difference is expressed as The weight matrix from the input layer to the enhancement layer of the width learning system is The bias of the neurons in the enhancement layer is Then the output of the enhancement layer of the width learning system is:
[0079]
[0080] You can get:
[0081]
[0082] and
[0083]
[0084] Then, we get:
[0085]
[0086] According to the present invention, the matrix obtained above can be calculated:
[0087]
[0088] Below, there are:
[0089]
[0090] According to the matrix solved above in the present invention, calculate:
[0091]
[0092] Then, update the matrix for:
[0093]
[0094] Therefore, the weight matrix from the enhancement layer to the output layer of the width learning system is obtained as:
[0095]
[0096] The predicted arrival time of the surface base station signal at the origin of the coordinate system and the estimated arrival time in the kth iteration We can obtain:
[0097]
[0098] Then, we get:
[0099]
[0100] Based on this matrix, we can know:
[0101]
[0102] Therefore, the estimated position increment of ship j is:
[0103]
[0104] 2. Select a decision action
[0105] The position increment of ship j As the mean of the Gaussian distribution strategy, the smaller value of the output value of the double Q network is used as the covariance of the Gaussian distribution strategy. Randomly select position increment Then, the time difference can be obtained Minimize the difference between the estimated time and the measured time. The rewards are calculated as follows:
[0106]
[0107] in, is the vector sum of arrival times Indicates position increment Time after execution. t and Q u is a symmetric positive definite matrix. In addition, the overall Q function is:
[0108]
[0109] Where γ∈(0,1) is the discount factor. Stored in the memory pool for updating the width learning system and the double Q network, where l c is the total number of iterations. In order to make the training process of the double Q network more stable, each l e Update the network once.
[0110] 3. Update status
[0111] Based on Gaussian distribution strategy Increment of a randomly selected position on The position update process of ship j is:
[0112]
[0113] Among them, action Determined by the width learning system and the Q network, the action space is continuous. Less than Δt or Estimated position of ship j is regarded as the actual position of ship j. Otherwise, repeat the above process. Where Δt represents the desired accuracy, is the upper bound of the iteration step.
Claims
1. A real-time deep reinforcement learning method, characterized in that The method comprises the following steps: Step 1: Estimate the mean of decision actions The agent starts from any given initial estimated state and uses the width learning system to learn the state increment from the latest data. The specific steps are as follows: Step 1.1, initialize the width learning system; Step 1.2, the agent updates the width learning system using the arrival time difference and state difference of the base station signals collected online, and uses the width learning system to learn the state increment from the latest data; Step 2: Select a decision action The decision action is selected based on the Gaussian distribution strategy with the output vector of the width learning system as the mean and the smaller value of the output value of the double Q network as the covariance. The specific steps are as follows: Step 2.1: The output vector of the width learning system, i.e., the decision action is regarded as the mean of the Gaussian distribution strategy, and the smaller value in the output value of the double Q network is regarded as the covariance of the Gaussian distribution strategy; Step 2.2: To evaluate the decision action performance, define a one-step reward function for: in, is the vector of arrival times, Indicates decision action Time after execution, Q t and Q u is a symmetric positive definite matrix; Step 2.3: Define the total Q function of the double Q network for: Among them, γ∈(0,1) is the discount factor, and represents the estimated state of agentj at the kth iteration step and the k+1th iteration step; Step 2.4: Learn the output vector of the system based on the value and width of the Q function From Gaussian distribution strategy Randomly select decision actions in Get the time difference Reward function and Q function; Step 2.5: Tuple Stored in the memory pool for updating the dual Q network, where l c is the total number of iterations; Step 3: Update status Step 3.1: Model the state estimation process as a Markov decision process and establish the state update process as: Step 3.2, until Less than Δt or Estimated status is regarded as the state vector of the agent, otherwise return to step 2, and e Return to step 1.2 to update the width learning system, where Δt represents the desired accuracy, is the upper bound of the iteration step.
2. The real-time deep reinforcement learning method according to claim 1, characterized in that The specific steps of step 1.2 are as follows: Definition The arrival time difference vector of the base station signal collected is The corresponding state difference is expressed as The weight matrix from the input layer to the enhancement layer of the width learning system is The bias of the neurons in the enhancement layer is Then the output of the enhancement layer of the width learning system for: definition: and Among them, H is given in the initialization stage ADD,ZH,0 The value of is then defined as: According to the matrix defined by formulas (1) to (4), we get: Below, define: According to the matrix defined by formulas (4) to (6), we define: Among them, λ→0 is the regularization parameter, which is given in the initialization stage The value of express situation, assuming express The situation; define the matrix The update expression is: The weight matrix from the enhancement layer to the output layer of the width learning system Updated to: Among them, the initial value W of the weight matrix from the enhancement layer to the output layer of the width learning system is given in the initialization stage on,ho,0 The value of; definition: in, is the difference between the predicted arrival time of the base station signal at the origin of the coordinate system and the estimated arrival time in the kth iteration step, is the predicted arrival time of the base station signal at the origin of the coordinate system, is the estimated arrival time of the base station signal at the origin of the coordinate system in the kth iteration step, It is the time difference vector formed by the difference between the predicted arrival time and the estimated arrival time in the kth iteration step; below, it is defined as: According to the definition of formula (12), we get: Then, the estimated state increment is: in, is the state increment of agent j output by the width learning system.
Citation Information
Patent Citations
Knowledge strategy selection method and device based on reinforcement learning
CN112990485A
Parameterized deep reinforcement learning algorithm based on value function
CN113569466A