Transform-DQN-based navigation method and device for mobile robot
By adopting multi-sensor fusion perception and improved DQN algorithm in mobile robot navigation, using Transformer model and sigmoid function, the problem of poor path planning strategy performance in complex dynamic environments is solved, and more efficient autonomous navigation is achieved.
Patent Information
- Application Number
- CN202510178571.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-27
AI Technical Summary
The existing path planning algorithm has slow convergence time in complex dynamic environments, and the path planning strategy performance is poor, resulting in the inability of mobile robots to effectively plan better paths.
Using a navigation method based on Transformer-DQN, the environment is perceived through multi-sensor fusion, environment information and obstacle information are obtained, and the improved DQN algorithm is used to introduce sigmoid functions and Transformer models to obtain the optimal strategy and control the movement of the mobile robot.
It improves the path planning efficiency and strategy performance of mobile robots in complex dynamic environments, and achieves faster and more stable autonomous navigation.
Smart Images

Figure CN120215486A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of robot path planning. Specifically, it relates to a navigation method and device for a mobile robot based on Transformer-DQN. Background Art
[0002] Path planning technology refers to a mobile robot finding an optimal path from the starting state to the target state that can avoid obstacles in its working environment according to one or some optimization criteria.
[0003] Currently, most path planning algorithms can quickly plan path trajectories in a simple and known environment. However, when working in a complex and unknown environment, they have poor exploration ability, slow algorithm convergence time, and low environmental adaptability, resulting in the mobile robot being unable to effectively plan an optimal path. Summary of the Invention
[0004] In order to solve the problems such as slow algorithm convergence time and poor path planning strategy performance of existing path planning algorithms in complex dynamic environments, the present application provides a navigation method and device for a mobile robot based on Transformer-DQN.
[0005] The embodiments of the present application are implemented as follows:
[0006] In a first aspect, the present application provides a navigation method for a mobile robot based on Transformer-DQN, including:
[0007] Perceiving the environment through multi-sensor fusion to obtain environmental information and the state information of the mobile robot;
[0008] According to the environmental information and state information, obtaining the expected action of the current mobile robot according to the optimal strategy obtained by the improved DQN algorithm, where a sigmoid function and a Transformer model are introduced in the improved DQN algorithm;
[0009] Controlling the movement of the mobile robot according to the current expected action.
[0010] In a possible implementation, the step of perceiving the environment through multi-sensor fusion to obtain environmental information and the state information of the mobile robot further includes:
[0011] Obtaining environmental information and the state information of the current mobile robot through a camera sensor and a single-line lidar sensor in a dynamic environment.
[0012] In a possible implementation, the environmental information includes obstacle information and target point position information, and the state information of the mobile robot includes the linear velocity and angular velocity of the mobile robot at time t, as well as the distance and angle of the mobile robot relative to the target point.
[0013] In a possible implementation, obtaining the expected action of the current mobile robot according to the optimal policy obtained by the DQN algorithm based on the environmental information and state information further includes:
[0014] Input the current state information into the DQN model to obtain the Q values corresponding to all actions in the action set;
[0015] Select the action corresponding to the maximum Q value as the current expected action of the mobile robot.
[0016] 6. In a possible implementation, the training process of the DQN model includes:
[0017] a. Initialize the parameters of the estimated Q network and the target network, initialize the hyperparameters of the learning rate, discount factor, and greedy factor, and the experience replay pool;
[0018] b. According to the state of the mobile robot at time t in the historical data, select and execute an action with a certain probability according to the greedy factor to obtain the moved state, that is, the state at time t+1. At the same time, calculate the immediate reward value obtained by executing this action. The immediate reward value is evaluated by a reward function, and the reward function is:
[0019] R = w1R goal + w2R obstacle + w3R smooth ,
[0020] where R goal is the target reward, which encourages the mobile robot to approach the target; R obstacle is the collision penalty, which punishes the behavior of approaching or colliding with obstacles; R smooth is the path smoothness reward, which encourages the mobile robot to move smoothly and reduces sharp turns or jitters; w1, w2, w3 are the weight coefficients of each reward to balance safety, efficiency, and smoothness;
[0021] c. Store the motion data in the experience replay pool, and then extract a batch of samples from the experience replay pool;
[0022] d. According to the samples extracted from the experience pool, calculate the target Q value using the target network, calculate the estimated Q value using the estimated network, and calculate the loss function;
[0023] e. Repeat the process of b-d until the DQN model converges.
[0024] In a possible implementation, controlling the movement of the mobile robot according to the current desired action further includes:
[0025] Using the trained DQN model, obtaining the current desired action of the mobile robot and converting it into a control instruction;
[0026] Controlling the movement of the mobile robot through linear velocity and angular velocity.
[0027] In a possible implementation, in step b, introducing the sigmoid function to improve the epsilon decay method can flexibly control the decay speed of epsilon by adjusting parameters.
[0028] In a possible implementation, in step c, introducing the Transformer model into the experience replay mechanism can better model the relationship between experiences, improving learning efficiency and model stability.
[0029] In a possible implementation, changing the limiting conditions for the first one hundred rounds of training. In the initial stage of training, only when the agent's exploration reaches the maximum number of steps does the current round end and the next round start. In the later stage of training, the agent learns experience samples stably with a smaller greedy factor.
[0030] In a second aspect, the present application provides a navigation device for a mobile robot based on Transformer-DQN, including:
[0031] An information acquisition module for perceiving the environment through multi-sensor fusion to obtain environmental information and the state information of the mobile robot;
[0032] An obstacle avoidance decision module for obtaining the current desired action of the mobile robot according to the optimal strategy obtained by the improved DQN algorithm based on the environmental information and state information, where the sigmoid function and the Transformer model are introduced into the improved DQN algorithm;
[0033] A planning implementation module for controlling the movement of the mobile robot according to the current desired action.
[0034] The technical solution provided by the present application can at least achieve the following beneficial effects:
[0035] A navigation method and device for a mobile robot based on Transformer-DQN provided by the present application perceive the environment through multi-sensor fusion to obtain environmental information and obstacle information; then, use the improved DQN algorithm to obtain the optimal strategy and find the optimal path; finally, convert the obtained strategy into a control instruction to achieve the purpose of autonomous navigation of the mobile robot in a dynamic environment.
[0036] Aiming at the problems of low perception accuracy and unstable recognition caused by the limitations of single-sensor environmental perception, multi-sensor fusion is proposed to perceive the environment. The environmental information obtained by the single-line lidar sensor and the camera is fused to enhance the perception accuracy and stability of the mobile robot. Aiming at the problems of low sample utilization rate and low learning efficiency caused by uniform sampling learning in the DQN algorithm, the Transformer model is introduced to consider the temporal and spatial correlations between experience sequences, so that high-quality experience samples can be better utilized and the learning efficiency of DQN can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic flowchart of a navigation method and device for a mobile robot based on Transformer-DQN shown in an exemplary embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the overall design of a mobile robot navigation system shown in an exemplary embodiment of the present application;
[0040] Figure 3 It is a schematic block diagram of a navigation system shown in an exemplary embodiment of the present application;
[0041] Figure 4 It is a schematic flowchart of a navigation system shown in an exemplary embodiment of the present application;
[0042] Figure 5 It is a flowchart of network model training shown in an exemplary embodiment of the present application;
[0043] Figure 6 It is a schematic flowchart of the Transformer experience replay mechanism shown in an exemplary embodiment of the present application;
[0044] Figure 7 It is a schematic structural diagram of a navigation device for a mobile robot based on Transformer-DQN shown in an exemplary embodiment of the present application.
[0045] REFERENCE NUMERALS:
[0046] 1. Information acquisition module; 2. Obstacle avoidance decision module; 3. Planning implementation module. DETAILED DESCRIPTION
[0047] In order to make the purpose, implementation manner and advantages of this application clearer and more understandable, the following will clearly and completely describe the exemplary implementation manner of this application in combination with the accompanying drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only a part of the embodiments of this application, rather than all the embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0048] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the following described implementation manner, rather than intending to limit the implementation manner of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meanings.
[0049] The terms "first", "second", "third", etc. in the description, claims and the above-mentioned accompanying drawings of this application are used to distinguish similar or homogeneous objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0050] The terms "include" and "have" and any of their variations are intended to cover but not exclusively include. For example, a product or device that includes a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components that are not clearly listed or are inherent to these products or devices.
[0051] Before explaining the navigation method of the mobile robot based on Transformer-DQN provided in the embodiments of this application, the application scenarios and implementation environments of the embodiments of this application will be introduced first.
[0052] Path planning technology refers to that a mobile robot finds an optimal path from the starting state to the target state and can avoid obstacles in its working environment according to one or some optimization criteria.
[0053] Currently, most path planning algorithms can quickly plan the path trajectory in a simple and known environment. However, when working in a complex and unknown environment, they have poor exploration ability, slow algorithm convergence time, and low environmental adaptability, resulting in the mobile robot being unable to effectively plan an optimal path.
[0054] Based on this, this application provides a navigation method and device for a mobile robot based on Transformer-DQN. It uses multi-sensor fusion to perceive the environment to obtain environmental information and obstacle information; then, it uses an improved DQN algorithm to obtain the optimal strategy and find the optimal path; finally, it converts the obtained strategy into control instructions to achieve the purpose of autonomous navigation of the mobile robot in a dynamic environment.
[0055] Aiming at the problems of low perception accuracy and unstable recognition caused by the limitations of single-sensor environmental perception, multi-sensor fusion is proposed to perceive the environment, and the environmental information obtained by the single-line lidar sensor and the camera is fused to enhance the perception accuracy and stability of the mobile robot for the environment. Aiming at the problems of low sample utilization rate and low learning efficiency caused by uniform sampling learning in the DQN algorithm in the experience replay pool, the Transformer model is introduced to consider the temporal and spatial correlations between experience sequences, so that high-quality experience samples can be better utilized and the learning efficiency of DQN can be improved.
[0056] Next, through embodiments and in combination with the accompanying drawings, the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be specifically described in detail. Embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments.
[0057] Figure 1 It is a schematic flowchart of a navigation method and device for a mobile robot based on Transformer-DQN shown in an exemplary embodiment of the present application.
[0058] In an exemplary embodiment, as Figure 1 shown, a navigation method for a mobile robot based on Transformer-DQN is provided, and the method may include the following steps:
[0059] Step 100: Perceive the environment through multi-sensor fusion to obtain environmental information and the state information of the mobile robot.
[0060] Step 200: According to the environmental information and state information, obtain the expected action of the current mobile robot according to the optimal policy obtained by the improved DQN algorithm, and the sigmoid function and the Transformer model are introduced in the improved DQN algorithm.
[0061] Step 300: Control the movement of the mobile robot according to the current expected action.
[0062] Figure 2 It is a schematic diagram of the overall design of a mobile robot navigation system shown in an exemplary embodiment of the present application.
[0063] In a possible implementation manner, as Figure 2 shown, it includes three modules: information acquisition, obstacle avoidance decision-making, and planning implementation.
[0064] First is the environmental information processing part, which uses the camera and single-line lidar mounted on the mobile robot to obtain the state information.
[0065] In the information acquisition module, considering the issues of cost reduction and the limitations of the sensors themselves, a camera sensor and a single-line lidar sensor are used in combination to perceive the environment.
[0066] The camera and the single-line lidar are relatively low-cost compared to other sensors. The camera is greatly affected by light but can perceive multi-dimensional information, while the single-line lidar is less affected by light but can only perceive the information on the height plane where it is located. Therefore, combining the camera and the single-line lidar sensors can be complementary and perceive the environment more comprehensively.
[0067] The information obtained from perceiving the environment includes the position (x t , y t ) and speed (v t , ω t ) information of the mobile robot itself, as well as the position (x ot , y ot ) and speed (v ot , ω ot ) of the dynamic obstacle. The experimental environment is a simulated warehouse environment, with 8 1×5 blocks used as fixed obstacles, and 5 obstacles randomly placed.
[0068] Then comes the obstacle avoidance decision part. The mobile robot starts from the current state, randomly selects an action and interacts with the environment, and the environment feedbacks the new state and reward value according to the action executed by the robot.
[0069] The robot stores the current state, the selected action, the obtained reward, and the new state in the experience replay pool, then uses the Transformer model to obtain the importance score of the experience samples, and then samples a batch of experience samples (including state, action, reward, and next moment state) from the experience replay pool by weighted sampling. The neural network is used for training. The goal of the neural network is to predict the Q value of each possible action according to the current state, and update the network parameters by minimizing the loss function, so as to learn which actions should be selected in different states to obtain the maximum long-term return.
[0070] This process is carried out through multiple iterations. Each time, an action is selected according to the current policy and the Q network is updated, gradually optimizing the policy to find the optimal path. As the training progresses, the robot can gradually select the action with the highest Q value from the neural network and get closer and closer to the optimal decision-making strategy.
[0071] In the obstacle avoidance decision module, during the process of the mobile robot interacting with the environment, a reward function is used to evaluate the contribution of the action to completing the navigation task.
[0072] A reward function considering multiple factors is designed as R = w1R goal+w2R obstacle +w3R smooth , where R goal is the target reward, which encourages the mobile robot to approach the target; R obstacle is the collision penalty, which punishes the behavior of approaching or colliding with obstacles; R smooth is the path smoothness reward, which encourages the mobile robot to move smoothly and reduces sharp turns or jitters; w1, w2, and w3 are the weight coefficients of each reward to balance safety, efficiency, and smoothness. The specific design is as follows:
[0073] R goal = α(D t-1 - D t )
[0074] where D t is the Euclidean distance between the mobile robot and the target at time t, and α is the scaling coefficient to adjust the reward amplitude.
[0075]
[0076] where d min is the distance between the robot and the nearest obstacle, d safe is the safety distance, and β is the penalty coefficient.
[0077] R smooth = -γ|a t - a t-1 |
[0078] where a t is the action at time t, and γ is the weight coefficient;
[0079] Select and execute an action according to the current state and the value of the greedy factor to obtain the corresponding reward and the new state, and then store the current state, action, reward, and next state in the experience replay pool. Sample the samples in the experience replay pool and input them into the neural network for training and learning.
[0080] Finally, in the planning implementation stage, according to the learned optimal policy, the robot converts the policy into specific action instructions, and controls the movement of the robot through these instructions to complete the path planning task.
[0081] In the planning implementation module, according to the learned optimal policy, the robot converts the policy into specific action instructions, and controls the movement of the robot through these instructions to complete the path planning task.
[0082] Figure 3 is a schematic block diagram of the navigation system shown in an exemplary embodiment of the present application, Figure 4 is a schematic flow diagram of the navigation system shown in an exemplary embodiment of the present application.
[0083] In a possible implementation, as Figure 3 shown, the specific implementation process of the navigation method is as follows:
[0084] Step 1: Obtain environmental information and mobile robot state information;
[0085] In a dynamic environment, through a camera sensor and a single-line lidar sensor, obtain environmental information and the state information of the current mobile robot. The environmental information refers to obstacle information and target point position information, and the state information of the mobile robot includes the linear velocity v t and angular velocity ω t of the mobile robot at time t, as well as the distance d t and angle θ t of the mobile robot relative to the target point;
[0086] Step 2: According to the environmental information and state information, obtain the expected action of the current mobile robot according to the optimal policy obtained by the DQN algorithm;
[0087] Input the current state information into the DQN model, obtain the Q values corresponding to all actions in the action set, and select the action corresponding to the maximum Q value as the current expected action of the mobile robot.
[0088] As Figure 4 shown, the steps of training the DQN model include:
[0089] Initialize the estimated Q network and target network parameters, initialize hyperparameters such as the learning rate, discount factor, and greedy factor, and the experience replay pool.
[0090] According to the state S t at time t in the historical data of the mobile robot, select and execute action A t with a certain probability according to the greedy factor, and obtain the moved state S t+ 1 (the state at time t+1). At the same time, calculate the immediate reward value R t obtained by executing this action.
[0091] Store the motion data [S t , A t , R t , S t+1 into the experience replay pool, and then extract a batch of samples from the experience replay pool.
[0092] According to the samples extracted from the experience pool, calculate the target Q value using the target network, calculate the estimated Q value using the estimated network, and calculate the loss function.
[0093] Repeat the process of b-d until the DQN model converges.
[0094] Step 3: Control the movement of the mobile robot according to the current desired action.
[0095] Use the trained DQN model to obtain the current desired action of the mobile robot, convert it into a control instruction, and control the movement of the mobile robot through the linear velocity and angular velocity.
[0096] Figure 5 is the flowchart of network model training shown in an exemplary embodiment of the present application, Figure 6 is the schematic diagram of the Transformer experience replay mechanism process shown in an exemplary embodiment of the present application.
[0097] In a possible implementation, as Figure 5 shown, during the training process, at the beginning of each episode, the environment is first initialized. The agent senses the current state, inputs it into the neural network, selects an action according to the greedy factor, obtains a new state and a reward value after executing the action, and stores the experience in the experience replay pool.
[0098] When a collision occurs or the target point is reached, the next episode is entered. If the agent encounters an obstacle or exceeds the boundary in each episode, the episode will automatically end and the next episode will be entered. The agent continuously loops this process to optimize the strategy, and the decay of the greedy factor is closely related to the number of episodes. When the episode ends too quickly, the agent fails to fully explore the environment in these episodes and then enters the next episode. As the number of episodes increases, the greedy factor gradually decreases, and the agent explores the environment with a smaller probability and is more inclined to learn the existing experience. However, in the previous episodes, the experience obtained by the agent may be insufficient or of poor quality, which results in the agent being unable to learn an effective and high-quality strategy from these experiences.
[0099] Therefore, in some embodiments of the present application, the limiting conditions for the first one hundred episodes of training are changed. In the initial stage of training, only when the agent explores to the maximum number of steps does the current episode end and the next episode enter. The advantage of doing this is that it allows the agent to fully explore the environment with a larger greedy factor in the initial stage of training and obtain rich and high-quality experience. In the later stage of training, it stably learns the experience samples with a smaller greedy factor, so as to better optimize the strategy and improve the overall performance of the agent.
[0100] It should also be noted that in step b of training the DQN model, an action is selected with a certain probability according to the value of the greedy factor. The conventional greedy factor decays linearly, and the linear decay uses a fixed decay rate regardless of the complexity of the problem.
[0101] For complex problems, it may not be possible to fully explore the key states and actions of the environment before epsilon decreases. For simple problems, unnecessary random exploration may continue even after exploration is complete.
[0102] Some embodiments of the present application improve the epsilon decay method by introducing the sigmoid function, which can flexibly control the decay rate of epsilon by adjusting parameters. The shape of the sigmoid function is similar to an "S" curve, with slow decay in the initial stage, rapid decay in the middle stage, and tending to be flat in the later stage.
[0103] This non-linear change enables the algorithm to have sufficient time for exploration in the initial stage and smoothly transition to the exploitation stage in the later stage, avoiding overly drastic switching between exploration and exploitation.
[0104] For complex problems, the algorithm can fully explore the state space by slowing down the decay rate in the initial stage. For simpler problems, the parameters can be adjusted to enter the stable exploitation stage earlier.
[0105]
[0106] where ε, ε0, ε min are the immediate value, initial value, and minimum value of the greedy factor respectively; k is a constant that controls the steepness of the curve; c is a constant that controls the offset of the curve.
[0107] In step c of training the DQN model, the traditional experience replay mechanism has the following problems: ignoring the correlation between experiences, that is, experiences are independent during sampling and the temporal and spatial correlations in the experience sequence are not considered; being unable to fully utilize sequence information, that is, failing to capture the long-term dependencies between experiences, which affects the learning effect.
[0108] Some embodiments of the present application improve the experience replay mechanism by introducing the Transformer model, which can better model the relationships between experiences and improve learning efficiency and model stability.
[0109] As Figure 6 shown, the specific steps are:
[0110] Arrange the experiences in the experience replay pool in chronological order to form an experience sequence where each experience contains: state S t 、action A t 、reward R t 、next moment state S t+1 ;
[0111] Perform feature representation and encoding on each experience: e t =[f(St ), f(A t ), R t , f(S t+1 ), P t , where f(S t ) is the feature representation of state S t ; f(A t ) is the embedding representation of action A t ; R t is the reward, used directly or normalized; P t is the positional encoding, used to preserve sequence position information.
[0112] Combine all the experience feature representations into an input sequence: E = {e1, e2,..., e N}, and input the input sequence E into the Transformer encoder for the following operations: Self-attention mechanism: Calculate the correlation between each experience in the sequence to obtain the attention weight matrix A. Encoding output: Through multiple layers of the Transformer encoder, obtain the hidden representation h t of each experience.
[0113] Use the output of the Transformer to calculate the importance score I t for each experience: I t = MLP(h t ), where MLP represents a multi-layer perceptron.
[0114] According to the importance score I t , perform weighted sampling on the experiences, and the sampling probability is According to the sampling probability P t , sample a mini-batch of data from the experience sequence for training.
[0115] Among them, during the process of feature representation and encoding for each experience, P t is used for positional encoding. Represent a single experience sample with a vector, and the samples in a round can be represented by a matrix by adding positional information to a single experience sample. This is the input matrix of the Transformer encoder. This is different from RNN and CNN which input vectors one by one and retain them. In the encoder, each experience sample in a round is calculated in parallel, so positional information needs to be introduced in advance.
[0116] It can be seen that some embodiments of the present application provide a navigation method based on the Transformer model and the DQN algorithm, aiming to improve the experience replay mechanism by using the Transformer model and complete path planning in a dynamic scenario with the DQN algorithm, so that the mobile robot can make safer decisions more quickly during movement and provide guarantee for the safe navigation of the mobile robot. This method can perceive the dynamic environment, make obstacle avoidance decisions at the same time, and perform real-time navigation; use the Transformer to capture long-term dependencies in the experience sequence, improve the utilization rate of experience samples, and enhance the stability of the model.
[0117] It should be understood that although the steps in the flowcharts involved in the above embodiments are displayed in sequence as indicated, these steps are not necessarily executed in the order indicated. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0118] Corresponding to the embodiments of the navigation method of the mobile robot based on Transformer-DQN described above and adopting the same technical concept, the present application also provides an embodiment of a navigation device of a mobile robot based on Transformer-DQN.
[0119] Figure 7 It is a schematic structural diagram of a navigation device of a mobile robot based on Transformer-DQN shown in an exemplary embodiment of the present application.
[0120] In an exemplary embodiment, as Figure 7 shown, the navigation device of the mobile robot based on Transformer-DQN includes:
[0121] An information acquisition module 1, configured to sense the environment through multi-sensor fusion and acquire environment information and mobile robot state information.
[0122] An obstacle avoidance decision module 2, configured to obtain the expected action of the current mobile robot according to the optimal strategy obtained by the DQN algorithm based on the environment information and the state information.
[0123] A planning implementation module 3, configured to control the movement of the mobile robot according to the current expected action.
[0124] For the specific limitations of a navigation device for a mobile robot based on Transformer-DQN, reference may be made to the limitations of the navigation method for a mobile robot based on Transformer-DQN in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned navigation device for a mobile robot based on Transformer-DQN can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0125] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0126] The embodiments described above merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A navigation method for a mobile robot based on Transformer-DQN, characterized in that: include: Perceive the environment through multi-sensor fusion to obtain environmental information and mobile robot status information; According to the environmental information and the state information, the desired action of the current mobile robot is obtained according to the optimal strategy obtained by the improved DQN algorithm, wherein the improved DQN algorithm introduces the sigmoid function and the Transformer model; Control the movement of the mobile robot according to the current desired action.
2. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 1, characterized in that: The method of sensing the environment through multi-sensor fusion and obtaining environmental information and mobile robot status information further includes: In a dynamic environment, camera sensors and single-line lidar sensors are used to obtain environmental information and the current state information of the mobile robot.
3. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 2, characterized in that: The environmental information includes obstacle information and target point position information, and the state information of the mobile robot includes the linear velocity and angular velocity of the mobile robot at time t, and the distance and angle of the mobile robot relative to the target point.
4. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 1, characterized in that: The method of obtaining the desired action of the current mobile robot according to the optimal strategy obtained by the DQN algorithm based on the environmental information and the state information further includes: Input the current state information into the DQN model to obtain the Q values corresponding to all actions in the action set; The action corresponding to the largest Q value is selected as the current expected action of the mobile robot.
5. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 4, characterized in that: The training process of the DQN model includes: a. Initialize the estimated Q network and target network parameters, initialize the learning rate, discount factor and greed factor hyperparameters, and the experience replay pool; b. According to the state of the mobile robot at time t in the historical data, the action is selected and executed with a certain probability according to the greed factor to obtain the state after the move, that is, the state at time t+1, and the instant reward value obtained by executing the action is calculated. The instant reward value is evaluated by the reward function, which is: R=w1R goal +w2R obstacle +w3R smooth , Among them, R goal is the target reward, which encourages the mobile robot to approach the target; R obstacle A collision penalty is a penalty for approaching or colliding with an obstacle; R smooth is the path smoothness reward, which encourages the mobile robot to move smoothly and reduce sharp turns or jitters; w1, w2, w3 are the weight coefficients of each reward to balance safety, efficiency and smoothness; c. Store the motion data in the experience replay pool, and then extract batch samples from the experience replay pool; d. Based on the samples drawn from the experience pool, the target Q value is calculated using the target network, the estimated Q value is calculated using the estimation network, and the loss function is calculated; e. Repeat the bd process until the DQN model converges.
6. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 1, characterized in that: The controlling the movement of the mobile robot according to the current desired action further comprises: Use the trained DQN model to obtain the current expected action of the mobile robot and convert it into control instructions; Control the motion of the mobile robot by linear velocity and angular velocity.
7. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 5, characterized in that: In step b, the sigmoid function is introduced to improve the epsilon attenuation method, and the attenuation speed of epsilon can be flexibly controlled by adjusting parameters.
8. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 5, characterized in that: In step c, introducing the Transformer model into the experience replay mechanism can better model the relationship between experiences and improve learning efficiency and model stability.
9. The navigation method of a mobile robot based on Transformer-DQN as claimed in claim 5, characterized in that: The restriction conditions of the first 100 rounds of training are changed. In the early stage of training, the current round ends and the next round begins only when the agent's exploration reaches the maximum number of steps. In the later stage of training, the experience samples are stably learned with a smaller greed factor.
10. A navigation device for a mobile robot based on Transformer-DQN, applied to the navigation method for a mobile robot based on Transformer-DQN according to any one of claims 1 to 9, characterized in that: include: The information acquisition module is used to perceive the environment through multi-sensor fusion and obtain environmental information and mobile robot status information; An obstacle avoidance decision module, used to obtain the desired action of the current mobile robot according to the environmental information and state information and the optimal strategy obtained by the improved DQN algorithm, wherein the improved DQN algorithm introduces a sigmoid function and a Transformer model; The planning implementation module is used to control the movement of the mobile robot according to the current desired action.
Citation Information
Cited By
Interaction action decision optimization method of humanoid robot and related equipment
CN120461451A
An interaction action decision optimization method and related device of a humanoid robot
CN120461451B