An end-to-end autonomous motion decision-making method based on deep reinforcement learning

By combining Mamba-YOLOv10 and HER-DQN algorithms, an end-to-end autonomous motion decision-making method is designed, which solves the specular interference and fragility problems of glass product identification and transport in biochemistry laboratories, and realizes stable and efficient transport of robots in complex environments.

CN119782825BActive Publication Date: 2025-07-25CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510267110.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-25
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In the biochemistry laboratory, the identification and transportation of glass products have problems such as specular interference, fragility of glass and unsmooth robot movement, resulting in low transport efficiency and safety risks.

Method used

The end-to-end autonomous motion decision-making method based on deep reinforcement learning is adopted, combined with the Mamba-YOLOv10 algorithm and the empirical playback mechanism HER, and the HER-DQN algorithm is designed, glass products are identified through dynamic sparse convolutional neural networks, and reward functions are designed to improve the robot's obstacle avoidance and decision stability.

Benefits of technology

It realizes accurate identification of glass products under complex lighting and environmental interference, improves the stability and safety of the robot's autonomous movement, and ensures the smoothness and efficiency of the transport process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782825B_ABST
    Figure CN119782825B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end autonomous motion decision-making method based on deep reinforcement learning. This method involves fields such as multi-robot motion, machine learning, and integrated decision-making. First, an improved Mamba-YOLOv10 is used to achieve accurate recognition of glass products. Then, aiming at the data dependence problem of end-to-end autonomous motion, an experience replay mechanism HER and a designed reward function are proposed to be applied to the deep Q-network to implement an integrated decision-making framework of HER-DQN. Compared with other methods, the present invention can not only accurately recognize glass products under complex lighting conditions and environmental interferences, but also improve the stability and robustness of the transfer robot in a biochemistry laboratory during autonomous motion decision-making. The designed end-to-end autonomous motion decision-making method has real-time performance and high efficiency, and can enable the transfer robot in a biochemistry laboratory to safely and autonomously complete the execution task, and has beneficial effects when applied to fields such as agriculture, industry, and service industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of multi-robot motion, machine learning, integrated decision-making, etc., and specifically relates to an end-to-end autonomous motion decision-making method based on deep reinforcement learning. Background Art

[0002] With the rapid development of biochemical research and the wide application of high-throughput analysis technology, the cross-regional transportation of experimental vessels has become a key link in modern scientific research processes. In order to improve transportation efficiency, transportation robots in biochemical laboratories have also developed rapidly. In the sample transportation of biochemical laboratories, the accurate identification of glassware is the core link to ensure safe transportation. However, the current biochemical laboratories generally adopt a transparent layout and use experimental tables made of transparent acrylic materials. Although the space utilization rate and operation convenience are maximized, the refraction characteristics of the acrylic material, as well as the light transmission and high reflectivity of the glassware itself, will cause specular reflection on the surface of the experimental vessels. At the same time, the visual interference caused by transmitted light will significantly affect the accurate identification of the characteristic points of the vessels by the transportation robot. Secondly, since glassware is prone to breakage due to collision, vibration or improper operation during transportation, it will cause liquid sample leakage, pollute the environment and threaten the health of experimental personnel. At the same time, the uneven motion modes such as sudden speed change, emergency obstacle avoidance and vibration of the transportation robot during transportation may also increase the risk of glassware breakage and leakage problems. Therefore, there are extremely high requirements for the safety and smoothness of the transportation process.

[0003] Existing autonomous motion systems based on robotic arms or mobile platforms generally adopt a multi-module architecture of perception-planning-decision-making-execution. Although they can achieve autonomous and safe navigation of robots, the heterogeneous communication between modules causes time delay, which affects the real-time control of robots. The non-deterministic delay accumulation between subsystems will cause trajectory deviation, resulting in a decrease in the positioning accuracy of the robot's end effector. In contrast, the end-to-end integrated decision control method, due to its closed-loop characteristic of directly driving the output of motion instructions by sensor data, can effectively improve the working efficiency of the transportation robot when performing tasks and can perform real-time feedback processing on complex and changing environments, becoming a hot technology for mobile robots to complete autonomous motion.

[0004] In view of the above problems and actual requirements, the present invention designs an end-to-end autonomous motion decision-making method based on deep reinforcement learning. First, in view of the unique environmental characteristics of the biochemical laboratory mentioned above, the Mamba model is combined with the YOLOv10 algorithm. Because of its dynamic sparse convolutional neural network structure, it can suppress specular reflection artifacts through an adaptive weight allocation mechanism in a strong reflection interference scenario. At the same time, it can enhance the edge sharpness perception of the glass contour under low-light conditions by using multi-scale residual connections, and can effectively identify glass doors, glassware, etc. in the laboratory, preventing collisions with transparent obstacles during the robot transportation process. Secondly, in view of the problem of data dependence in end-to-end autonomous motion, experience replay (Hindsight Experince Replay, HER) is introduced to improve the stability and sample utilization rate during the training process. Finally, in view of the problems of poor real-time performance, low efficiency, and uneven motion caused by the traditional autonomous motion modular method, a reward function is designed that is safety-oriented and ensures the smoothness of motion and the high efficiency of task completion. The experience replay mechanism and the designed reward function are applied to the deep Q-network to realize the design of an integrated decision-making framework based on experience replay-deep Q-network (HindsightExperince Replay-Deep Q-Network, HER-DQN), enabling the transportation robot to use the HER-DQN model to achieve safe and efficient autonomous navigation in the biochemical laboratory, avoiding the problems of fragile sample breakage and liquid sample leakage, and realizing the safe transportation of samples. Summary of the Invention

[0005] The present invention designs an end-to-end autonomous motion decision-making method based on deep reinforcement learning.

[0006] The specific implementation steps are as follows:

[0007] Step 1: Construct a virtual environment and system architecture. In this method, the proposed algorithm needs to be repeatedly trained and tested in a dynamically changing environment. First, an initial environment is set, and elements such as a mobile robot system, pedestrians, and other glass experimental tools are added through software tools to construct and improve the virtual environment. The mobile robot system consists of a robot body, a sensor module, a processing module, and an execution module. The sensor module is responsible for identifying glass products and surrounding complex environment information using the object detection results of Mamba-YOLOv10; the processing module is used to perform tasks such as execution positioning, identification, and navigation on the information collected by the sensor; the execution module controls the movement of the robot according to the calculation results of the processing module.

[0008] Step 2: Dynamic recognition of glass products. The Mamba model adopts a dynamic sparse convolutional neural network structure. Combined with the YOLOv10 algorithm, it can perform feature analysis on glass products under complex lighting conditions and environmental interferences, providing sample information for the decision-making process. Train the Mamba-YOLOv10 model so that the robot can identify and locate glass products in a long-distance environment and extract information such as the position, category, and confidence of the glass products, providing data support for subsequent decision-making processes and obstacle avoidance.

[0009] The measure used to evaluate the matching degree between the predicted bounding box and the ground truth bounding box is called the consistent matching metric, and the formula is as follows:

[0010] ,

[0011] where, represents the spatial prior, is the classification confidence, represent the predicted bounding box and the ground truth bounding box respectively, is the intersection over union (IoU) of the predicted bounding box and the ground truth bounding box;

[0012] Object score The calculation formula is:

[0013] ,

[0014] is the object score of the nth positive sample, represents the maximum value of IoU among all positive samples, represents the alignment degree of the positive sample used to evaluate the predicted bounding box and the ground truth bounding box, Similarly, it is the maximum value;

[0015] Classification loss calculation formula:

[0016] l n =− ω n [ t n ⋅ log σ ( x n ) + ( 1 − t n ) ⋅ log ( 1 − σ ( x n ) ) ] ,

[0017] is the coefficient of the loss function, represents the predicted value of the nth sample, represents the activation function;

[0018] Coordinate loss calculation formula:

[0019] ,

[0020] is the intersection over union of the nth positive sample, object score is the weight coefficient, is the time when the nth sample bounding box is formed;

[0021] The total loss is the weighted sum of classification loss, coordinate loss, and confidence loss, and the formula is as follows:

[0022] ,

[0023] in, is the weighting coefficient, is the confidence loss.

[0024] Step 3: Design state space and action space. Use the Experience Replay-Deep Q Network (HER-DQN) algorithm to achieve navigation, set the state in the virtual environment built in step 1, set the pedestrian state, and set the robot state. The robot's state information is a key component of operation. This information includes the robot's precise position, the speed and direction of movement, and other data that may affect the robot's behavior. Similarly, setting the specific position of pedestrians, the speed of movement, and the direction of travel are important for understanding and predicting the behavior of pedestrians and robots.

[0025] Assume the position of the robot in the environment is , the location of the target point is , the position of the obstacle closest to the robot is ,but is the distance between the robot and the target location, is the distance between the robot and the obstacle. At the same time, the angle between the robot position and the target location is calculated as , the angle between the robot and the obstacle is .choose As the state space of the robot in the environment, it can express the state of the robot at a certain moment in the environment.

[0026] The action space is used to describe the robot's actions such as moving forward, backward, and changing direction in the current environment. After setting the robot's starting point and target point, the robot moves as a particle in the environment. In the actual movement process, the robot's movement process is a continuous state, so the robot's action space is a discrete action. The specific actions include moving forward, backward, turning left, turning right, turning left 45°, and turning right 45°. The purpose of adding corner actions is to improve the robot's ability to explore corner situations in the environment and avoid increasing the number of training times due to too little action space.

[0027] Step 4: Design a reward function. Designing a reward function mainly includes two aspects: reaching the target and avoiding obstacles. Designing a suitable reward function requires considering multiple factors such as path length, obstacle avoidance effect, and speed to the target point. In this method, a design scheme of three comprehensive reward functions is proposed, which considers target point reward, obstacle avoidance reward, and speed reward.

[0028] Step 4.1: Set the reward for reaching the target location. When the robot successfully reaches the target location, a relatively large positive reward should be given to encourage the robot to complete the decision-making task quickly and efficiently. Therefore, when the robot is near the target location, the reward function should output a relatively large positive value to encourage the robot to find the target point, and a negative value should be given as a penalty to the mobile robot;

[0029] R t a r g e t = { 1 , d m [ − 0 . 5 , 0 . 5 ] 20 ⋅ d m − 1 , − 1 < d m <− 0 . 5 , 0 . 5 < d m < 1 0 , Other ,

[0030] where is the reward for reaching the target location, is the maximum distance between the robot and the target location.

[0031] Step 4.2: Set the obstacle avoidance reward function. To ensure that the robot can effectively avoid obstacles in a complex and changing environment and reduce the risk of contact with surrounding objects and people, an obstacle avoidance reward function is designed. This function can not only help the robot accurately locate and avoid various obstacles, but also enhance the robot's autonomous learning ability. When the robot identifies an obstacle and takes measures to maintain a safe distance, the system will automatically give the robot a certain positive reward. When the robot fails to detect an obstacle in time or approaches or collides with an obstacle during the execution of the task, a negative reward will be given. The positive reward is an incentive for the robot's good performance, while the negative reward is a punishment for improper behavior. Through this strategy of combining rewards and punishments, the robot can continuously improve its obstacle avoidance ability and enhance its adaptability and reliability in different environments.

[0032] ,

[0033] where is the obstacle avoidance reward function, is the distance between the robot and the obstacle.

[0034] Step 4.3: Set the speed reward. The speed of the robot in a complex road environment should not be too fast to avoid behaviors such as moving too fast, colliding, and slipping. At the same time, the robot should be encouraged to reach the destination more efficiently, and a speed reward mechanism is added to the decision-making process. When the robot can move forward quickly in a way that exceeds the predetermined speed, the system should give positive feedback and rewards in a timely manner. By this method, it can be ensured that the robot can complete the task according to the set goal in the shortest time and the best path.

[0035] Set the speed reward function as:

[0036] ,

[0037] Among all the reward functions of the above design, the safety of the robot during the transfer of products is the most important. Secondly, it is required to reach the transfer location. When the above three reward functions conflict, safety and no collision are ensured first. Secondly, the speed reward will be controlled to avoid collision. Finally, it reaches the target point as required. The overall system reward function :

[0038] ,

[0039] where is the absolute distance of the robot to the target location, is time, , , are the reward functions for reaching the target point, obstacle avoidance, and speed respectively, , , are the coefficients of each reward function, satisfying .

[0040] Step 5: Calculate the Temporal Difference Target (TD-Target). In HER-DQN, the TD-Target is used to calculate the loss function, and the network parameters of the Q value are updated by minimizing the mean square error between the predicted Q value and the TD-Target to ensure that the optimal policy can be effectively learned.

[0041] The TD-Target is calculated by the following formula:

[0042] ,

[0043] Set the loss function of the HER-DQN algorithm:

[0044] L ( θ i ) = F [ ( R + λ max A ' Q ( S ' , A ' , θ ¯ i ) − Q ( S , A , θ i ) ) 2 ] ,

[0045] Substitute the TD-Target and simplify to get:

[0046] L ( θ i ) = F [ ( y i − Q ( S , A , θ i ) ) 2 ] ,

[0047] where is the discount factor, is the state, is the reward, is the action, is the coefficient of the function, is the state at the next moment, is the action at the next moment, is the parameter of the Q network.

[0048] Step 6: Train the experience replay system. The robot continues to interact with the environment, stores experiences to update the neural network. The robot can also continuously interact with and learn from the environment during the training of the HER-DQN algorithm, including the target position information detected by Mamba-YOLOv10, and can adjust the optimal path according to the ground complexity. Eventually, it learns to select the best action in a given state, forming a complete environmental state.

[0049] Step 7: Model evaluation and testing, control the execution output. Through simulation environment testing and experiments, comprehensively evaluate the performance of the HER-DQN model in the end-to-end autonomous motion decision-making of the robot. The reward function plays a key role in the testing and control execution phases. By reasonably designing the reward function and updating the Q value, it can guide the mobile robot to efficiently complete tasks. The specific steps are as follows:

[0050] Step1: Simulation environment testing. Observe whether the algorithm model can enable the robot to accurately identify glass products, conduct decision-making analysis during the transfer process, avoid obstacles, etc. Record the reward value of each test, calculate the average reward to evaluate the performance, and evaluate the stability and consistency of the model through multiple tests;

[0051] Step2: Visualization testing. Use visualization tools (such as MATLAB) to draw the reward curve and observe the learning process and convergence of the model;

[0052] Step3: Control the execution output. Control the execution to convert the output of the HER-DQN model into actual robot control instructions, and train the HER-DQN model to select actions by designing the reward function. Input the current state into the model and select the action with the highest Q value;

[0053] Step4: Convert the action into a control instruction. Convert the action output by the HER-DQN model into a control instruction that the robot can execute.

[0054] The present invention designs a method for autonomous motion decision-making of a robot, which is used to solve the problem of safe and efficient autonomous navigation of a transfer robot in a biochemistry laboratory. Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0055] 1. The improved Mamba-YOLOv10 target detection model of the present invention can accurately identify the characteristics of glass products in a biochemistry laboratory under complex lighting and environmental interference through algorithm optimization;

[0056] 2. The improved HER-DQN algorithm of the present invention improves the stability and robustness of the transfer robot in a biochemistry laboratory during autonomous motion decision-making, and also improves the smoothness of the transfer robot during operation;

[0057] 3. The end-to-end autonomous motion decision-making method designed by the present invention has real-time performance and high efficiency, enabling the biochemical laboratory transfer robot to safely and autonomously complete its execution tasks. This method has beneficial effects in fields such as agriculture, industry, and the service industry.

[0058] The present invention will be described in further detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is the overall flowchart in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] APPENDIX Figure 1 is the overall flowchart of the implementation of the present invention. This embodiment provides an end-to-end autonomous motion decision-making method based on deep reinforcement learning. Taking the biochemical laboratory sample transfer robot (BLSTR) as an example, the specific process includes the following: constructing a virtual environment and system architecture, dynamically identifying and precisely positioning glassware, designing a state space and an action space, designing a reward function, calculating a loss function according to the temporal difference target, training an experience replay system, evaluating and testing the model, and controlling the execution output.

[0062] The specific implementation is described as follows:

[0063] Implementation step 1: The environment and the agent are the main bodies of the deep reinforcement learning system. In this embodiment, the environment is set as a biochemical laboratory, and the agent is the biochemical laboratory sample transfer robot (BLSTR). The determination of the environment and the definition of the system architecture are as follows.

[0064] Step 1.1: Construct a closed biochemical laboratory environment. According to the actual environment of the biochemical laboratory, use the Gazebo simulator in the ROS system to construct the real environment of the biochemical laboratory. It can simulate the behavior of the robot in the biochemical laboratory to identify glassware and transfer samples.

[0065] Step 1.2: Add a biochemical laboratory transfer robot system, which consists of a robot body, a sensor module, a processing module, and an execution module. The sensor module is responsible for identifying glassware and surrounding interference environment information using the object detection results of Mamba-YOLOv10; the processing module is used to perform tasks such as execution positioning, identification, and navigation on the information collected by the sensor; the execution module controls the movement of the robot according to the calculation results of the processing module.

[0066] Implementation step 2: Dynamic identification of glassware. Analyze the characteristics of glassware under different lighting conditions using the Mamba-YOLOv10 algorithm on the collected data to provide sample information for the decision-making process. Train the Mamba-YOLOv10 model so that BLSTR can identify and locate glassware in a long-distance environment and also extract information such as the position, category, and confidence of the glassware, providing data support for subsequent decision-making processes and obstacle avoidance.

[0067] What is used to measure the matching degree between the predicted box and the ground truth box is called the consistent matching metric, and the formula is as follows:

[0068] ,

[0069] where, represents the spatial prior, is the classification confidence, represent the predicted box and the ground truth box respectively, is the intersection over union of the predicted box and the ground truth box;

[0070] Object score The calculation formula of is:

[0071] ,

[0072] is the object score of the nth positive sample, represents the maximum value of IoU among all positive samples, represents the alignment degree of the positive sample used to measure the predicted box and the ground truth box, Similarly, it is the maximum value;

[0073] Classification loss calculation formula:

[0074] l n =− ω n [ t n ⋅ log σ ( x n ) + ( 1 − t n ) ⋅ log ( 1 − σ ( x n ) ) ] ,

[0075] is the coefficient of the loss function, represents the predicted value of the nth sample, represents the activation function;

[0076] Coordinate loss calculation formula:

[0077] ,

[0078] is the intersection over union of the nth positive sample, and the target score is the weight coefficient, is the time when the nth sample bounding box is formed;

[0079] The total loss is the weighted sum of the classification loss, the coordinate loss, and the confidence loss, and the formula is as follows:

[0080] ,

[0081] where, is the weighting coefficient, is the confidence loss.

[0082] Implementation step 3: Design the state space and the action space. Use the Hindsight Experience Replay - Deep Q - Network (HER - DQN) algorithm to achieve navigation. Set the states of the researcher and the robot (BLSTR) in the virtual environment built in step 1. Setting the state information of the BLSTR is a key component of the operation. These information include the exact position of the BLSTR, and also involve the moving speed and direction, as well as the data that may affect the behavior of the BLSTR. Similarly, setting the specific position, moving speed, and traveling direction and other state information of the researcher is of great significance for understanding and predicting the behaviors of the researcher and the robot.

[0083] Let the position of the BLSTR in the laboratory be , the position of the target point be , and the position of the obstacle closest to the BLSTR be , then is the distance between the BLSTR and the target location, is the distance between the BLSTR and the obstacle. At the same time, calculate the angle between the position of the BLSTR and the target location as , and the angle between the BLSTR and the obstacle as . Select as the state space of the BLSTR in the laboratory, which can represent the state of the BLSTR at a certain moment in the laboratory.

[0084] The action space is used to describe the actions that BLSTR takes in the current laboratory environment, such as moving forward, moving backward, changing direction, etc. After setting the starting point and target point of BLSTR, BLSTR moves as a particle in the environment. In the actual movement process, the movement process of BLSTR is a continuous state, so the action space of BLSTR is a discrete action. The specific actions include moving forward, moving backward, turning left, turning right, turning left 45°, and turning right 45°. The purpose of adding corner actions is to improve BLSTR's ability to explore corner situations in the environment and avoid increasing the number of training times due to too little action space.

[0085] Implementation step 4: Design reward function. Designing reward function mainly includes two aspects: reaching the target and avoiding obstacles. Designing a suitable reward function needs to consider multiple factors such as path length, obstacle avoidance effect, speed to reach the target point, etc. In this method, a design scheme of three comprehensive reward functions is proposed, which considers target point reward, obstacle avoidance reward, and speed reward.

[0086] Step 4.1: Set the reward for reaching the target location. When BLSTR successfully reaches the target location, a large positive reward should be given to encourage BLSTR to complete the decision-making task quickly and efficiently. Therefore, when the robot is near the target location, the reward function should output a large positive value to encourage BLSTR to find the target point, and when a negative value is obtained, it is a penalty for the mobile robot;

[0087] R t a r g e t = { 1 , d m [ − 0 . 5 , 0 . 5 ] 20 ⋅ d m − 1 , − 1 < d m <− 0 . 5 , 0 . 5 < d m < 1 0 , Other ,

[0088] in It is the reward for reaching the destination. is the maximum distance between the robot and the target location.

[0089] Step 4.2: Set the obstacle avoidance reward function. In order to ensure that the robot can effectively avoid obstacles in a complex and changing environment and reduce the risk of contact with surrounding objects and people, an obstacle avoidance reward function is designed. This function can not only help the robot accurately locate and avoid various obstacles, but also enhance the autonomous learning ability of BLSTR. When BLSTR identifies an obstacle and takes measures to maintain a safe distance, the system will automatically give BLSTR a certain positive reward. When BLSTR fails to find the obstacle in time or approaches or collides with the obstacle during the task, it will give a negative reward. Positive rewards are incentives for BLSTR to perform well, while negative rewards are penalties for inappropriate behavior. Through this combination of reward and punishment strategy, the robot can continuously improve its obstacle avoidance ability and improve its adaptability and reliability in different environments.

[0090] ,

[0091] in is the obstacle avoidance reward function, is the distance from the robot to the obstacle.

[0092] Step 4.3: Set the speed reward. BLSTR should not move too fast in a complex road environment to avoid behaviors such as moving too fast, colliding, and slipping. At the same time, the robot should be encouraged to reach the destination more efficiently, and a speed reward mechanism should be added to the decision-making process. When BLSTR can move forward quickly beyond the predetermined speed, the system should give positive feedback and rewards in a timely manner. By this method, it can be ensured that the robot can complete the task according to the set goal in the shortest time and the best path.

[0093] Set the speed reward function as:

[0094] ,

[0095] Among all the reward functions designed above, the safety of the robot when transporting products is the most important. Secondly, it is required to reach the transfer location. When the above three reward functions conflict, first ensure safety and no collision. Secondly, the speed reward will be controlled to a certain extent to avoid collision. Finally, reach the target point as required. The overall system reward function :

[0096] ,

[0097] where, is the absolute distance from the robot to the target location, is the time, , , are the reward functions for reaching the target point, obstacle avoidance, and speed respectively, , , are the coefficients of each reward function, satisfying .

[0098] Step 5: Calculate the Temporal Difference Target (TD-Target). In HER-DQN, the TD-Target is used to calculate the loss function, and the network parameters of the Q value are updated by minimizing the mean square error between the predicted Q value and the TD-Target to ensure that the optimal policy can be effectively learned.

[0099] The TD-Target is calculated by the following formula:

[0100] ,

[0101] Set the loss function of the HER-DQN algorithm:

[0102] L ( θ i ) = F [ ( R + λ max A ' Q ( S ' , A ' , θ ¯ i ) − Q ( S , A , θ i ) ) 2 ] ,

[0103] Substituting into TD - Target and simplifying gives:

[0104] L ( θ i ) = F [ ( y i − Q ( S , A , θ i ) ) 2 ] ,

[0105] where is the discount factor, is the state, is the reward, is the action, is the coefficient of the function, is the state at the next moment, is the action at the next moment, are the parameters of the Q - network.

[0106] Implementation step 6: Train the experience replay system. BLSTR continuously interacts with the environment and stores experiences to update the neural network. The robot can also continuously interact with and learn from the environment during the training of the HER - DQN algorithm, including the target position information detected by Mamba - YOLOv10, and can adjust the optimal path according to the ground complexity. Eventually, it learns to select the best action in a given state, forming a complete environmental state.

[0107] Implementation step 7: Model evaluation and testing, control the execution output. Through simulation environment testing and experiments, comprehensively evaluate the performance of the HER - DQN model in the end - to - end autonomous motion decision of BLSTR. The reward function plays a key role in the testing and control execution phases. By reasonably designing the reward function and updating the Q - value, the mobile robot can be guided to efficiently complete tasks. The specific steps are as follows:

[0108] Step1: Simulation environment testing. Observe whether the algorithm model can enable the robot to accurately identify glass products, make decision - making analysis during the transfer process, avoid obstacles, etc. Record the reward value of each test, calculate the average reward to evaluate the performance, and evaluate the stability and consistency of the model through multiple tests;

[0109] Step2: Visualization testing. Use visualization tools (such as MATLAB) to plot the reward curve and observe the learning process and convergence of the model;

[0110] Step3: Control the execution output. Control the execution to convert the output of the HER - DQN model into actual robot control instructions. Train the HER - DQN model to select actions by designing the reward function. Input the current state into the model and select the action with the highest Q - value;

[0111] Step 4: Convert the action into a control instruction, and convert the action output by the HER-DQN model into a control instruction that the robot can execute.

Claims

1. An end-to-end autonomous motion decision-making method based on deep reinforcement learning, characterized in that It includes the following steps: Step 1: Construct a virtual environment and system architecture; Use the Gazebo simulator to set an initial environment, and use software tools to add a mobile robot system, pedestrians, and other glass experiment tools to construct and improve the virtual environment. The mobile robot system consists of a robot body, a sensor module, a processing module, and an execution module. The sensor module is responsible for using the object detection results of Mamba-YOLOv10 to identify glass products and complex environmental information around; The processing module is used to process the information collected by the sensors to perform positioning, recognition, and navigation tasks; The execution module controls the movement of the robot according to the calculation results of the processing module. Step 2: Dynamic recognition of glass products; The Mamba model adopts a dynamic sparse convolutional neural network structure, combined with the YOLOv10 algorithm to perform feature analysis on glass products under complex lighting conditions and environmental interference, and train the Mamba-YOLOv10 model, enabling the robot to identify and locate glass products in a long-distance environment and also extract information on the position, category, and confidence of glass products, providing sample information for the decision-making and obstacle avoidance processes. Step 3: Design the state space and action space; Use the Experience Replay - Deep Q-Network HER-DQN algorithm to achieve navigation, set the state in the virtual environment built in Step 1, set the pedestrian state and the robot state, and the action space is used to describe the actions of the robot to move forward, backward, and change direction in the current environment. Step 4: Design three reward functions: reward for reaching the target location, reward for obstacle avoidance, and reward for speed. Set the reward function for reaching the target location as: where R target is the reward for reaching the target location, d m is the maximum distance between the robot and the target location; Set the reward function for obstacle avoidance as: where R obstacle is the obstacle avoidance reward function, and d n is the distance from the robot to the obstacle; Set the reward function for speed as: Among all the above-designed reward functions, the safety of the robot during the transfer of products is the most important, followed by the requirement to reach the transfer location. When the above three reward functions conflict, first ensure safety and no collision. Secondly, to avoid collision, the speed reward will be controlled. Finally, reach the target location as required. The final overall system reward function R: where d f is the absolute distance from the robot to the target location, T is the time, and R target , R obstacle , R v are the reward functions for reaching the target location, obstacle avoidance, and speed respectively. β, μ are the coefficients of each reward function, satisfying Step 5: Calculate the temporal difference target; In HER-DQN, calculate the temporal difference target TD-Target for calculating the loss function, and update the network parameters of the Q value by minimizing the mean square error between the predicted Q value and TD-Target, effectively learning the optimal policy. TD-Target is calculated by the following formula: Set the loss function of the HER-DQN algorithm: Substitute TD-Target and simplify to get: L(θ i ) = F[(y i - Q(S, A, θ i )) 2 , where λ is the discount factor, S is the state, R is the reward, A is the action, F is the coefficient of the function, and S' = S t+1 is the state at the next moment, and A' = A t+1 is the action at the next moment, and θ i are the parameters of the Q-network; Step 6: Train the experience replay system; The robot continues to interact with the environment, stores experiences to update the neural network. The robot continuously interacts and learns in the training of the HER-DQN algorithm, including the target position information detected by Mamba-YOLOv10, adjusts the optimal path, and finally learns to select the best action in a given state. Step 7: Model evaluation and testing; evaluate the performance of the HER-DQN model in the end-to-end autonomous motion decision-making of the robot, use the reward curve generated by the visualization tool MATLAB, select the action with the highest Q value according to the model input of the current state; convert the action output by the HER-DQN model into a control instruction that the robot can execute.

2. The end-to-end autonomous motion decision-making method based on deep reinforcement learning according to claim 1, characterized in that Dynamic recognition of glass products described in Step 2; the Mamba model adopts a dynamic sparse convolutional neural network structure and combines with the YOLOv10 algorithm to perform feature analysis on glass products under complex lighting conditions and environmental interference, train the Mamba-YOLOv10 model, enabling the robot to identify and locate glass products in a long-distance environment and also extract information on the position, category, and confidence of the glass products, providing sample information for the decision-making and obstacle avoidance processes, and specifically implemented according to the following steps: Measuring the matching degree between the predicted box and the ground truth box is called the consistent matching metric, and the formula is as follows: where s represents the spatial prior, p α is the classification confidence, b represents the predicted bounding box and the ground truth bounding box respectively, is the intersection over union of the predicted bounding box and the ground truth bounding box; Target score t n The calculation formula is as follows: t n is the target score of the nth positive sample, μ * represents the maximum value of IoU among all positive samples, m n represents the degree of alignment between the predicted bounding box and the ground truth box for positive samples, m * Similarly, it is the maximum value; Classification loss calculation formula: l n = -ω n [t n ·logσ(x n ) + (1 - t n )·log(1 - σ(x n ))], ω n is the coefficient of the loss function, x n represents the predicted value of the nth sample, σ(x n ) represents the activation function; Coordinate loss calculation formula: IOU n is the intersection over union of the n-th positive sample, and the target score t n is the weight coefficient, M n is the time when the n-th sample bounding box is formed; The total loss is the weighted sum of the classification loss, coordinate loss, and confidence loss, and the formula is as follows: L = λ(l n + L box + σ(t0)), Among them, λ is the weighting coefficient, and σ(t0) is the confidence loss.

Citation Information

Patent Citations

  • Motion planning method based on machine learning in complex environment

    CN116551703A

  • Unmanned aerial vehicle navigation method based on deep reinforcement learning and PID controller

    CN117387635A