Goal-driven navigation experience replay augmented reinforcement learning method, computer device and medium
By optimizing deep reinforcement learning through the Actor-Critic architecture and greedy experience replay mechanism, and designing reward functions and hyperparameters, the problems of low navigation efficiency and poor adaptability in existing technologies are solved, enabling robots to navigate efficiently in dynamic environments.
Patent Information
- Application Number
- CN202510657141.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing goal-driven navigation technologies based on deep reinforcement learning suffer from low data utilization efficiency, conservative navigation behavior, and poor adaptability to dynamic environments, resulting in slow model convergence, long training cycles, and low navigation success rates.
A deep reinforcement learning algorithm is designed using an Actor-Critic architecture. It combines a greedy experience replay mechanism and a delayed policy update technique, designs a reward function with a reward and punishment mechanism, ranks the importance of experience data by calculating the TD error, dynamically adjusts the probability of experience extraction, and controls the training process through hyperparameter optimization.
It improves the navigation efficiency and success rate of robots in dynamic and complex environments, reduces the behavior of robots remaining stationary and getting too close to obstacles, and enhances obstacle avoidance capabilities and the stability and adaptability of navigation algorithms.
Smart Images

Figure CN120540080B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot autonomous navigation technology, and in particular to a reinforcement learning method, computer device and medium for target-driven navigation with experience playback enhancement. Background Technology
[0002] Target-guided autonomous navigation technology is one of the core research directions in the field of robot control and automation, and is widely used in scenarios such as industrial warehousing robots, unmanned vehicles, and intelligent service robots. In these applications, robots need to perceive environmental information (such as LiDAR data and visual information) in real time in dynamic and complex environments, and quickly and safely reach the target location through efficient path planning and obstacle avoidance strategies.
[0003] Currently, navigation methods based on Deep Reinforcement Learning (DRL) have become the mainstream technology. Typical implementations include:
[0004] 1. Deep Q-Network (DQN): Improves training stability by mitigating data correlation through an Experience Replay pool and a Target Network;
[0005] 2. Policy Gradient-Based Navigation Algorithm (DDPG): Combines policy optimization and value assessment to balance exploration and utilization efficiency.
[0006] However, existing target-driven navigation techniques based on deep reinforcement learning still have the following key shortcomings in practical applications:
[0007] 1. Inefficient data utilization: Traditional experience replay mechanisms use uniform sampling, ignoring the differences in the importance of different experiences for training. For example, high-value experiences (such as successful obstacle avoidance and rapid approach to the target) cannot be fully learned because they are sampled with equal probability, resulting in slow model convergence and long training cycles.
[0008] 2. Conservative navigation behavior: Due to the simple design of the reward function (such as relying solely on the end reward), the robot is prone to getting trapped in local optima (such as wandering in place or frequently triggering obstacle avoidance), which reduces the navigation success rate;
[0009] 3. Poor adaptability to dynamic environments: In scenarios with densely distributed obstacles or frequent changes in target points (such as logistics warehouses and hospital corridors), the obstacle avoidance response speed and path optimization capability of existing methods are significantly reduced. Summary of the Invention
[0010] The purpose of this invention is to provide a reinforcement learning method, computer device, and medium for target-driven navigation with enhanced experience playback, which can improve the navigation efficiency and success rate of robots in dynamic and complex environments, and is applicable to various scenarios such as industrial warehousing robots, unmanned vehicles, and intelligent service robots.
[0011] To achieve the above objectives, this invention provides a reinforcement learning method for goal-driven navigation with experience playback enhancement, comprising the following steps:
[0012] Build a robot simulation environment and design a robot model and sensor data acquisition module;
[0013] Design neural networks based on the Actor-Critic architecture, including Actor networks and Critic networks;
[0014] A deep reinforcement learning algorithm is designed using the Actor-Critic architecture, and a delayed update strategy is employed.
[0015] Design a reward function, including reward and penalty mechanisms for collisions, reaching the target point, linear velocity, and obstacle avoidance behavior;
[0016] A greedy experience replay mechanism is designed, which ranks the importance of experience data by calculating the TD error and dynamically adjusts the experience extraction probability by combining greedy sampling and random sampling strategies.
[0017] By optimizing the training process using hyperparameters, the robot's navigation performance can be maximized.
[0018] Preferably, the Actor network consists of a linear layer, an activation function, and a normalization layer. The input is laser data and waypoint polar coordinates, and the output is linear velocity and angular velocity. The Critic network adopts a dual-Q network structure. The input is state and action, and the output is Q-value evaluation.
[0019] Preferably, the delay policy update includes:
[0020] The main policy network parameters are synchronized to the target policy network using a soft update method.
[0021] The Q-value network parameters are updated to the target Q-value network using a soft update method.
[0022] Preferably, the reward function is as follows:
[0023]
[0024] The reward value is +120 when the robot reaches the target point, -120 when the robot collides with an obstacle, and in other cases, the reward value consists of linear velocity reward, obstacle avoidance penalty and small negative penalty.
[0025] action(0) represents the robot's linear velocity, goal represents the robot reaching the target point, collision represents the robot colliding with an obstacle, m(.) represents the obstacle avoidance penalty function, and min_laser represents the distance to the nearest obstacle detected by the lidar, normalized to 0-1. When min_laser<1: m(min_laser)=1-min_laser, when min_laser≥1: m(min_laser)=0.
[0026] Preferably, in step S5, the greedy experience replay mechanism specifically includes:
[0027] Calculate the TD error δ for each experience:
[0028] δ=|Q1(s t ,a t )-Q target1 |+|Q2(s t ,a t )-Q target1 |;
[0029] Q target1 =reward+((1-done)×discount×Q target );
[0030] Where Q1(s) t ,a t ) and Q2(s t ,a t ) represent the target policy network's response to state s. t and behavior a t The assessment, Q target This indicates that after the target policy network evaluates the next state and behavior, it takes the minimum score among them, Q. target1 The value represents the median, reward represents the reward value, done indicates whether a segment has ended after the robot takes action (if it has ended, done = True; if it has not ended, done = False), and discount represents the discount factor.
[0031] The empirical data are sorted based on the TD error, and sampling weights are assigned using a probability distribution function.
[0032] Set the greedy hyperparameters greedy_count and random_count to control the ratio of greedy sampling to random sampling;
[0033] Weight parameters are introduced into the gradient, and the sample distribution shift is balanced by adjusting the coefficients.
[0034] Preferably, the probability distribution function is:
[0035]
[0036] Where P(i) represents the i-th probability value, rank(i) represents the sorted order, and α represents a hyperparameter between 0 and 1;
[0037] After obtaining the sampling probability of each experience, normalization is performed to obtain the final sampling probability Pnormalized(i) for each data point as follows:
[0038]
[0039] Where P(j) represents the j-th probability value.
[0040] Preferably, the formula for calculating the weight parameters is as follows:
[0041] W i = (N×P(i)) -β ;
[0042] Among them, W i Let represent the i-th weight parameter, N represent the number of samples, and β represent the adjustment coefficient. The formula for calculating β is as follows:
[0043]
[0044] Where beta_start represents the initial value of β, frame represents the current number of training frames, and beta_frames is used to adjust the number of frames of β;
[0045] After obtaining the weight parameters, normalization is performed. The calculation process is as follows:
[0046]
[0047] Among them, W i ' represents the normalized weight parameters, W j This represents the j-th weight parameter.
[0048] Preferably, the hyperparameter optimization metrics include: discount factor and learning rate.
[0049] The present invention also provides a computer device including a memory and a processor, the memory being used to store instructions and the processor being used to execute the instructions to implement the experience playback-enhanced reinforcement learning method for target-driven navigation as described above.
[0050] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the reinforcement learning method for experience-replay enhancement of target-driven navigation as described above.
[0051] Therefore, the present invention employs the above-described reinforcement learning method, computer device, and medium for target-driven navigation with enhanced experience playback, and the beneficial technical effects are as follows:
[0052] (1) The designed reward function effectively reduces the robot's behavior of remaining stationary and getting too close to obstacles by setting reward and punishment mechanisms for collision, reaching the target point, linear velocity, and obstacle avoidance behavior. This makes the robot more actively seek the target point during navigation and significantly improves its obstacle avoidance ability.
[0053] (2) The proposed greedy experience replay method ranks the importance of experience data by calculating TD error, and dynamically adjusts the experience extraction probability by combining greedy sampling and random sampling strategies, thereby improving the efficiency of experience utilization, accelerating the convergence speed of navigation algorithm, and thus improving the overall performance of model.
[0054] (3) By optimizing the training process through hyperparameters, the stability and adaptability of the navigation algorithm are further improved, enabling it to exhibit better navigation efficiency and success rate in dynamic and complex environments. Attached Figure Description
[0055] Figure 1 This is the Actor-Critic architecture diagram;
[0056] Figure 2 This is an overall structural diagram of the present invention;
[0057] Figure 3 This is a cumulative reward value distribution chart; among which, Figure 3 (a) in the figure is a distribution chart of the cumulative reward value in each round during the G-TD3 test; Figure 3 (b) in the figure is a distribution chart of the cumulative reward value in each round during the GER-TD3 test; Figure 3 (c) in the figure is a distribution chart of the cumulative reward value in each round during the GDAE test;
[0058] Figure 4 It is a step size distribution chart; among which, Figure 4 (a) in the figure is the cumulative step size distribution of each round during the G-TD3 test; Figure 4 (b) in the figure is the cumulative step size distribution of each round during the GER-TD3 test; Figure 4 (c) in the figure is the cumulative step size distribution of each round during the GDAE test;
[0059] Figure 5 It is a time distribution chart; among which, Figure 5 (a) in the figure is a time distribution diagram of each round during the G-TD3 test; Figure 5 (b) in the figure is the time distribution diagram of each round during the GER-TD3 test; Figure 5 (c) in the figure is the time distribution diagram of each round during the GDAE test;
[0060] Figure 6 It is a success rate change curve; among which, Figure 6 (a) in the figure is a graph showing the change in success rate during the G-TD3 test; Figure 6 (b) in the figure is a graph showing the change in success rate during the GER-TD3 test; Figure 6 (c) in the figure is a graph showing the change in success rate during the GDAE test process;
[0061] Figure 7 It is a graph showing the change in collision rate; among which, Figure 7 (a) The collision rate change curve during the G-TD3 test process; Figure 7 (b) in the figure is the collision rate change curve during the GER-TD3 test; Figure 7 (c) in the figure is the collision rate change curve during the GDAE test. Detailed Implementation
[0062] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0063] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0064] Example 1
[0065] A reinforcement learning method for enhancing experience playback in goal-driven navigation includes the following steps:
[0066] Step 1: Construct a robot simulation environment and design the robot model and sensor data acquisition module;
[0067] The sensor data acquisition module uses lidar to detect obstacle information within its range;
[0068] Step 2: Design a neural network based on the Actor-Critic architecture, including the Actor network and the Critic network.
[0069] Neural network architecture such as Figure 1 As shown, given a state s t The state, composed of laser-read data and waypoint polar coordinates relative to the robot's position, is input into the Actor network. The Actor network consists of a series of linear layers, activation functions, and normalization layers. After processing by the Actor network, the input state ultimately outputs an action 'a'. t The action a tIt consists of linear velocity and angular velocity. The agent executes the corresponding actions, and the quality of the actions is evaluated by the Critic network.
[0070] Overall architecture as follows Figure 2 As shown, the robot interacts with the environment and generates data (s t ,a t ,r t ,d t ,s t+1 ), where s t For the environmental state, a t For the actions taken by the robot, r t The reward obtained for the action taken, d t s serves as a marker indicating whether a segment has ended. t+1 The next environmental state after taking an action is defined. A neural network is trained using sampled data obtained through greedy experience replay. The policy network π(s; θ) represents the policy for selecting an action based on the current state s and its parameters θ. The value network Q(s, a; w) evaluates the selected action based on the current state s, action a, and its parameters w, thus assisting in the training of the policy network. The target policy network π(s; θ) is also defined. * ) and the target value network Q(s,a;w * θ is obtained through soft updates of the policy network and value network. * w represents the network parameters of the target policy. * This represents the network parameters of the target value.
[0071] The Critic network employs a double Q-network structure, approximating the value function. It consists of multiple linear layers and activation functions. Given an input state and the behavior generated by the Actor network, the input is fed into the Critic network, ultimately outputting a Q-value. The Q-value represents the Critic network's evaluation of the policy selected by the Actor network. Through feedback from the Critic network, the Actor network continuously optimizes its own policy, gradually improving its chosen strategy. Simultaneously, the Critic network continuously refines its evaluation process to enhance the accuracy of its policy evaluation.
[0072] Step 3: Design a deep reinforcement learning algorithm that combines the master policy network, the target policy network, the master Q-value network, and the target Q-value network, and adopts a delayed policy update technique;
[0073] The policy network consists of 3 linear layers, 2 layer normalization layers, 2 ReLU activation functions, and 1 Tanh activation function. The target policy network has the same structure as the main policy network. The main Q-value network consists of 4 linear layers and 2 ReLU activation functions. The target Q-value network has the same structure as the main Q-value network.
[0074] Deep reinforcement learning algorithms improve learning stability and performance by introducing delayed policy update techniques. Delayed policy update involves periodically synchronizing the parameters of the main policy network to the target policy network, while using a soft update method to update the parameters of the target Q-value network.
[0075] Step 4: Design the reward function, including reward and punishment mechanisms for collisions, reaching the target point, linear velocity, and obstacle avoidance behavior;
[0076] The reward function is as follows:
[0077]
[0078] The reward value is +120 when the robot reaches the target point, -120 when the robot collides with an obstacle, and in other cases, the reward value consists of linear velocity reward, obstacle avoidance penalty and small negative penalty.
[0079] action(0) represents the robot's linear velocity, goal represents the robot reaching the target point, collision represents the robot colliding with an obstacle, m(.) represents the obstacle avoidance penalty function, and min_laser represents the distance to the nearest obstacle detected by the lidar, normalized to 0-1. When min_laser<1: m(min_laser)=1-min_laser, when min_laser≥1: m(min_laser)=0.
[0080] Step 5: Design a greedy experience replay mechanism. Rank the importance of experience data by calculating the TD error, and dynamically adjust the experience extraction probability by combining greedy sampling and random sampling strategies.
[0081] The greedy experience replay mechanism specifically includes:
[0082] Calculate the TD error δ for each experience:
[0083] δ=|Q1(s t ,a t )-Q target1 |+|Q2(s t ,a t )-Q target1 |;
[0084] Q target1 =reward+((1-done)×discount×Q target );
[0085] Where Q1(s) t ,a t ) and Q2(s t ,a t) represent the target policy network's response to state s. t and behavior a t The assessment, Q target This indicates that after the target policy network evaluates the next state and behavior, it takes the minimum score among them, Q. target1 The value represents the median, reward represents the reward value, done indicates whether a segment has ended after the robot takes action (if it has ended, done = True; if it has not ended, done = False), and discount represents the discount factor.
[0086] The empirical data are sorted based on the TD error, and sampling weights are assigned using a probability distribution function.
[0087] The probability distribution function is:
[0088]
[0089] Where P(i) represents the i-th probability value, rank(i) represents the sorted order, and α represents a hyperparameter between 0 and 1;
[0090] After obtaining the sampling probability of each experience, normalization is performed to obtain the final sampling probability Pnormalized(i) for each data point as follows:
[0091]
[0092] Where P(j) represents the j-th probability value.
[0093] Set the greedy hyperparameters greedy_count and random_count to control the ratio of greedy sampling to random sampling;
[0094] Weight parameters are introduced into the gradient, and the sample distribution shift is balanced by adjusting the coefficients.
[0095] The formula for calculating the weight parameters is as follows:
[0096] W i = (N×P(i)) -β ;
[0097] Among them, W i Let represent the i-th weight parameter, N represent the number of samples, and β represent the adjustment coefficient. The formula for calculating β is as follows:
[0098]
[0099] Where beta_start represents the initial value of β, frame represents the current number of training frames, and beta_frames is used to adjust the number of frames of β;
[0100] After obtaining the weight parameters, normalization is performed. The calculation process is as follows:
[0101]
[0102] Among them, W i ' represents the normalized weight parameters, W j This represents the j-th weight parameter.
[0103] Step 6: Optimize the training process using hyperparameters (discount factor, learning rate) to achieve the best robot navigation performance.
[0104] The invention will be further illustrated by simulation experiments below.
[0105] The experiments were conducted in a simulation environment, specifically Gazebo. The robot model was defined using URDF or SDF files. Linkages described rigid components (including visual, collision, and inertial attributes), joints constrained the motion relationships between components, sensor plugins simulated LiDAR / camera and other sensing devices, and control plugins implemented the motion logic. Within the simulation environment, the robot interacted with the environment, continuously collecting data during this interaction. This collected data was stored in a buffer pool for subsequent training. This process was dynamic; the robot continuously collected data, and the buffer pool continuously added new data and deleted some older data. A greedy experience replay method was adopted, continuously collecting data from the buffer pool to train the robot.
[0106] To demonstrate the effectiveness of the proposed method, comparative experiments were conducted, primarily comparing three models. The first model, denoted as G-TD3, uses the neural network and reward function designed in this invention. The second model, GER-TD3, builds upon G-TD3 by incorporating a greedy experience replay mechanism. The third model is the currently advanced GDAE model.
[0107] Five evaluation metrics were selected: accuracy, average step size, average reward, collision rate, and average duration.
[0108] To present the experimental results more intuitively, Figures 3 to 5 The average of the data is calculated, and the results are shown in Table 1. Figure 3 As shown in Table 1, the proposed model GER-TD3 has the highest average reward, indicating that the proposed method is feasible and effective. Figure 4 As shown in Table 1, the GER-TD3 model proposed in this invention has the smallest average step size, indicating that the proposed method is feasible and effective. Figure 5 As shown in Table 1, the average duration of the proposed model GER-TD3 is the smallest, indicating that the proposed method is feasible and effective.
[0109] To present the experimental results more intuitively, Figure 6 , Figure 7 The results are shown in Table 2. Figure 6 As shown in Table 2, GER-TD3 has the highest accuracy. Figure 7 As shown in Table 2, GER-TD3 has the lowest collision rate, indicating that the proposed method is feasible and effective.
[0110] Table 1 Performance Evaluation
[0111]
[0112] Table 2 Performance Evaluation
[0113]
[0114] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0115] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0116] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0117] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.
[0118] Therefore, the present invention employs the above-mentioned reinforcement learning method, computer equipment and medium that enhances the experience playback of target-driven navigation, which can improve the navigation efficiency and success rate of robots in dynamic and complex environments, and is applicable to various scenarios such as industrial warehousing robots, unmanned vehicles, and intelligent service robots.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A reinforcement learning method for goal-driven navigation with experience playback enhancement, characterized in that, Includes the following steps: Build a robot simulation environment and design a robot model and sensor data acquisition module; Design neural networks based on the Actor-Critic architecture, including Actor networks and Critic networks; A deep reinforcement learning algorithm is designed using the Actor-Critic architecture, and a delayed update strategy is employed. Design a reward function, including reward and penalty mechanisms for collisions, reaching the target point, linear velocity, and obstacle avoidance behavior; A greedy experience replay mechanism is designed, which ranks the importance of experience data by calculating the TD error and dynamically adjusts the experience extraction probability by combining greedy sampling and random sampling strategies. The training process is optimized by hyperparameter control to achieve the best robot navigation performance. The delay policy update includes: The main policy network parameters are synchronized to the target policy network using a soft update method. The Q-value network parameters are updated to the target Q-value network using a soft update method. The reward function is as follows: The reward value is +120 when the robot reaches the target point, -120 when the robot collides with an obstacle, and in other cases, the reward value consists of linear velocity reward, obstacle avoidance penalty and small negative penalty. action(0) represents the robot's linear velocity, goal represents the robot reaching the target point, collision represents the robot colliding with an obstacle, m(.) represents the obstacle avoidance penalty function, and min_laser represents the distance to the nearest obstacle detected by the lidar, normalized to 0-1. When min_laser<1: m(min_laser)=1-min_laser, when min_laser≥1: m(min_laser=0; In step S5, the greedy experience replay mechanism specifically includes: Calculate the TD error δ for each experience: δ=|Q1(s t ,a t )-Q target1 |+|Q2(s t ,a t )-Q target1 |; Q target1 =reward+((1-done)×discount×Q target ); Where Q1(s) t ,a t ) and Q2(s t ,a t ) represent the target policy network's response to state s. t and behavior a t The assessment, Q target This indicates that after the target policy network evaluates the next state and behavior, it takes the minimum score among them, Q. target1 The value represents the median, reward represents the reward value, done indicates whether a segment has ended after the robot takes action (if it has ended, done = True; if it has not ended, done = False), and discount represents the discount factor. The empirical data are sorted based on the TD error, and sampling weights are assigned using a probability distribution function. Set the greedy hyperparameters greedy_count and random_count to control the ratio of greedy sampling to random sampling; Weight parameters are introduced into the gradient, and the sample distribution shift is balanced by adjusting the coefficients.
2. The reinforcement learning method for enhancing experience playback in target-driven navigation according to claim 1, characterized in that, The Actor network consists of linear layers, activation functions, and normalization layers. Its inputs are laser data and waypoint polar coordinates, and its outputs are linear velocity and angular velocity. The Critic network adopts a dual-Q network structure. Its inputs are state and action, and its output is Q-value evaluation.
3. The reinforcement learning method for enhancing experience playback in target-driven navigation according to claim 1, characterized in that, The probability distribution function is: Where P(i) represents the i-th probability value, rank(i) represents the sorted order, and α represents a hyperparameter between 0 and 1; After obtaining the sampling probability of each experience, normalization is performed to obtain the final sampling probability Pnormalized(i) for each data point as follows: Where P(j) represents the j-th probability value.
4. The reinforcement learning method for enhancing experience playback in target-driven navigation according to claim 1, characterized in that, The formula for calculating the weight parameters is as follows: IN i =(N×P(i)) -β ; Among them, W i Let represent the i-th weight parameter, N represent the number of samples, and β represent the adjustment coefficient. The formula for calculating β is as follows: Where beta_start represents the initial value of β, frame represents the current number of training frames, and beta_frames is used to adjust the number of frames of β; After obtaining the weight parameters, normalization is performed. The calculation process is as follows: Among them, W i ' represents the normalized weight parameters, W j This represents the j-th weight parameter.
5. The reinforcement learning method for enhancing experience playback in target-driven navigation according to claim 1, characterized in that, Hyperparameter optimization metrics include: discount factor and learning rate.
6. A computer device, characterized in that, It includes a memory and a processor, the memory being used to store instructions and the processor being used to execute the instructions to implement the reinforcement learning method for experience-replay-enhanced target-driven navigation as described in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning method for experience-replay-enhanced target-driven navigation as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Reinforced learning unmanned aerial vehicle flight path planning method based on delayed experience-first playback mechanism
CN116974299A