Robot path planning method and system
By combining the ST-Transformer model and the TD3 algorithm, the spatiotemporal dependencies of the robot's state data are captured, solving the problem of inaccurate path planning in dynamic environments caused by traditional methods, and achieving efficient path planning for robots in complex environments.
Patent Information
- Application Number
- CN202510598552.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional path planning methods exhibit poor adaptability and real-time performance in dynamic environments, and cannot meet the needs of efficient and accurate path planning for robots in complex environments.
The ST-Transformer model is combined with the double-delayed deep deterministic policy gradient algorithm (TD3). The dual-path spatiotemporal attention mechanism is used to capture the spatiotemporal dependencies of the environment and generate the optimal action instructions for the robot at the next moment.
It improves the robot's path planning adaptability in dynamic environments, generates action strategies that are more in line with the current environmental status, and improves task execution efficiency.
Smart Images

Figure CN120593746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robot path planning method and system, and belongs to the field of robot control. Background Art
[0002] With the continuous advancement of intelligent technology, the application of mobile robots has become increasingly widespread in various industries, including industrial automation, intelligent logistics, autonomous driving, and medical services. Especially in complex dynamic environments, achieving efficient and accurate path planning has become a core technical challenge facing robotic systems. Traditional path planning methods, such as those based on graph search, artificial potential fields, and dynamic planning, although they perform well in static or known environments, often exhibit poor adaptability and real-time performance in rapidly changing, unknown dynamic environments, and are unable to meet the needs of real-time, flexible path decision-making. Traditional algorithms often fail to fully account for long-term environmental changes and have a lag in their response to dynamic environments. This can cause robots to deviate from the optimal path when performing tasks, reducing task execution efficiency. Summary of the Invention
[0003] The purpose of the present invention is to provide a robot path planning method and system to solve the problem of inaccurate robot path planning.
[0004] To achieve the above object, the solution of the present invention includes: A robot path planning method of the present invention comprises the following steps: 1) obtaining state data of the robot at the current moment and several consecutive moments before the current moment; 2) inputting the obtained state data into an ST-Transformer model, jointly encoding temporal dynamics and spatial correlations through the ST-Transformer model's dual-path spatiotemporal attention mechanism to generate a spatiotemporal fused feature representation, inputting the feature representation into an Actor network in a trained dual-delay deep deterministic policy gradient algorithm, and obtaining the optimal action of the robot for path planning at the next moment.
[0005] Furthermore, the untrained Actor network is trained according to the pre-collected state data to obtain the corresponding action. The Critic network in the double-delay deep deterministic policy gradient algorithm obtains the corresponding value based on the action and the pre-collected state data. The untrained Actor network updates its own parameters according to the value, thereby forming a trained Actor network in the double-delay deep deterministic policy gradient algorithm.
[0006] Furthermore, the status data includes: the distance between the robot and the obstacle, the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position, where the connecting line is the line between the robot's current position and the target position.
[0007] Furthermore, the real-time position and posture of the robot are determined using the AMCL algorithm to obtain the yaw angle, the angle between the robot's motion direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position.
[0008] A robot path planning system includes a processor that executes a computer program to implement the following steps: 1) obtaining state data of the robot at a current moment and several consecutive moments before the current moment; 2) inputting the obtained state data into an ST-Transformer model, jointly encoding temporal dynamics and spatial correlation through a dual-path spatiotemporal attention mechanism of the ST-Transformer model to generate a spatiotemporal fused feature representation, and inputting the feature representation into an Actor network in a trained dual-delay deep deterministic policy gradient algorithm to obtain the optimal action of the robot for path planning at the next moment.
[0009] Furthermore, the untrained Actor network is trained according to the pre-collected state data to obtain the corresponding action. The Critic network in the double-delay deep deterministic policy gradient algorithm obtains the corresponding value based on the action and the pre-collected state data. The untrained Actor network updates its own parameters according to the value, thereby forming a trained Actor network in the double-delay deep deterministic policy gradient algorithm.
[0010] Furthermore, the status data includes: the distance between the robot and the obstacle, the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position, where the connecting line is the line between the robot's current position and the target position.
[0011] Furthermore, the real-time position and posture of the robot are determined using the AMCL algorithm to obtain the yaw angle, the angle between the robot's motion direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position.
[0012] The beneficial effects of the present invention are as follows: the present invention is a pioneering invention. The present invention inputs the collected robot state data at continuous moments into the ST-Transformer model that can effectively jointly model spatiotemporal dynamic characteristics. The model synchronously captures the long-distance spatiotemporal dependencies of the environmental state through a dual-path attention mechanism, and then inputs the spatiotemporal features encoded by the model into the trained Actor network, thereby obtaining the robot's action instructions for path planning at the next moment. Since the input is the robot's state data at continuous moments and the ST-Transformer model performs spatiotemporal joint modeling on these robot continuous state data, the model can not only analyze the temporal evolution law, but also explore spatial correlation, thereby more comprehensively capturing the global characteristics of the dynamic changes in the environment, helping the TD3 algorithm to generate action strategies that are more in line with the current environmental state in a non-steady-state environment, and improving the intelligent agent's adaptability to complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a schematic diagram of a path planning network structure of the present invention; Figure 2 It is a training process diagram of the path planning network of the present invention; Figure 3 It is a robot in a simulation environment of the present invention; Figure 4 It is a schematic diagram of a simulation environment of the present invention; Figure 5 This is a schematic diagram of path planning in a simulation environment of the present invention; Figure 6 It is a robot in a real environment of the present invention; Figure 7 It is a real environment schematic diagram of the present invention; Figure 8 This is a schematic diagram of path planning in a real environment of the present invention. DETAILED DESCRIPTION
[0014] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.
[0015] The concept of this invention is to input the robot's state time series data into the ST-Transformer model. This model, with its dual-path spatiotemporal attention mechanism, has significant advantages in both time series modeling and spatial correlation analysis, effectively capturing long-range spatiotemporal dependencies. The spatiotemporal feature representation processed by the ST-Transformer model is then input into a trained Actor network to obtain the robot's next action command.
[0016] Method implementation method: This embodiment provides a robot path planning method, which first obtains environmental data through sensors such as laser rangefinders, heading angle sensors, and odometers. These environmental data are processed accordingly to generate sequence state data for reinforcement learning. Figure 1 As shown, these sequence state data, that is, the sequence state data of the current moment and several consecutive moments in history (for example, the state data S of the current moment and three consecutive moments immediately before the current moment) are combined. t-3 、S t-2 、S t-1 、S t ) is input into the ST-Transformer model, and then the sequence state data processed by the ST-Transformer model is input into the Actor network in the trained TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm to output the optimal action instruction at the next moment.
[0017] The ST-Transformer model is introduced to enhance the overall model's ability to jointly model spatiotemporal data. The spatiotemporal relative position encoding layer combines relative position biases based on time intervals and spatial distances with state information, effectively capturing the temporal dynamics and spatial correlations between states. The ST-Transformer encoder layer employs a dual-path spatiotemporal attention mechanism to separately compute cross-timestep dependencies within the same spatial block and cross-timestep interactions within the same spatial block, significantly improving the efficiency of spatiotemporal feature extraction. Finally, a spatiotemporal global pooling layer aggregates features across the spatiotemporal dimensions to generate a fixed-size spatiotemporal fused feature representation, providing fixed-size input features for subsequent fully connected layers. The actor network consists of three fully connected layers and combines sigmoid and tanh nonlinear activation functions to learn complex nonlinear relationships, adapting to the requirements of robot motion control. Based on the corresponding state data, the actor network outputs the robot's linear and angular velocities, using a Gaussian distribution to model the motion. To ensure the robot's physical properties, the linear velocity is limited by a sigmoid function, while the angular velocity is adjusted by a tanh function. In this way, the Actor network can output reasonable action instructions (linear velocity and angular velocity), ensuring that the robot can efficiently plan its path in a dynamic environment.
[0018] As a specific embodiment, the sequence state data S iIncluding 10-dimensional laser ranging results, yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, the relative position distance between the robot and the target position, that is, each sequence state data S i is a 14-dimensional vector. The laser ranging result is the distance between the robot and the obstacle collected by the laser radar during one rotation, thus forming a 10-dimensional laser ranging result. The connecting line refers to the line connecting the robot's current position and the target position.
[0019] As a specific example, after obtaining environmental data, the robot generates point cloud data and processes it using the AMCL (Adaptive Monte Carlo Localization) algorithm to generate time-series state data. Point cloud generation converts data from sensors like lidar into a point cloud format for describing the environmental structure. The AMCL algorithm is used to localize the robot within the environment, determining its posture and position to determine the yaw angle, the angle between the robot's motion direction and the connecting line, the angular difference between the yaw angle and the angle, and the relative distance between the robot and the target position.
[0020] like Figure 2 As shown in the figure, training data collection and construction: the robot collects environmental data through sensors such as lidar, generates point cloud data and processes the environmental data with the AMCL algorithm to obtain sequence state data, and initializes the sequence length to 4. The sequence state data is then input into the ST-Transformer model, and then the sequence state data processed by the ST-Transformer model is input into the Actor network for training.
[0021] TD3 algorithm training: The actor network generates actions based on the corresponding sequence state data. The critic network evaluates the path planning quality based on the sequence state data and the actions generated by the actor network, outputting a value Q. This value Q indicates the quality of the current path planning. The actor network then adjusts its network parameters based on the value Q, resulting in a higher-quality path plan, ultimately resulting in a fully trained actor network. This implementation collects reward values (e.g., approaching the target point, avoiding obstacles, etc.) through 400,000 interactions to continuously optimize the path planning strategy.
[0022] This embodiment has also been tested and verified in simulation environment and real environment respectively. The simulation platform adopts Gazebo simulation platform, such as Figure 3 As shown in the figure, a car with differential speed change is built on the simulation platform to simulate a crawler robot with differential speed in real scenes. Figure 4As shown in the figure, the simulation environment includes the settings of various typical scenarios, such as obstacle avoidance, narrow channel passage, complex dynamic environment, etc. The red box in the figure represents the corresponding obstacles, dynamic obstacles and corresponding narrow channels. The simulation results of one of them are as follows Figure 5 As shown in the figure, the red lines represent the corresponding path planning routes of the robot for different target positions.
[0023] The robots in real environment are tracked, such as Figure 6 As shown, the crawler robot chassis also uses drive differential control. The host computer runs the Linux operating system and is configured with ROS and Pytorch environments for control and data processing. The microcontroller, as the underlying hardware, is responsible for driving the robot chassis and communicating with the host computer through the serial port. As the core computing device, the host computer sends control signals to the chassis controller through the serial port, and at the same time receives sensor data from the lidar and controller. The lidar collects environmental information in real time, processes point cloud data through ROS, and generates the current state of the robot in combination with the odometer information. The trained ST-Transformer-TD3 model is ported to the host computer, and the algorithm network is loaded using the Pytorch environment. The path planning module is integrated in the ROS framework to realize the real-time input of lidar data and odometer data, the state sequence generation of the algorithm, and the action output. Real scenes such as Figure 7 As shown in the figure, the collected environmental information is used for path planning and action decision making. The output linear velocity and angular velocity instructions are received by the chassis controller and drive the crawler robot to move. The path planning performed by the crawler robot for this scene is shown in the figure. Figure 8 shown.
[0024] System implementation method: This embodiment provides a robot path planning system, in which the computer program executed by the processor is designed using the method described in the method implementation method. Since the introduction of this method is clear enough, it will not be repeated here.
[0025] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific embodiments of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A robot path planning method, characterized in that: The method includes the following steps: 1) obtaining the robot's state data at the current moment and several consecutive moments before; 2) inputting the obtained state data into the ST-Transformer model, jointly encoding temporal dynamics and spatial correlation through the ST-Transformer model's dual-path spatiotemporal attention mechanism, generating a spatiotemporal fused feature representation, and inputting the feature representation into the Actor network in the trained dual-delay deep deterministic policy gradient algorithm to obtain the robot's optimal action for path planning at the next moment.
2. The robot path planning method according to claim 1, characterized in that: The untrained Actor network is trained according to the pre-collected state data to obtain the corresponding action. The Critic network in the double-delay deep deterministic policy gradient algorithm obtains the corresponding value according to the action and the pre-collected state data. The untrained Actor network updates its own parameters according to the value, thereby forming the trained Actor network in the double-delay deep deterministic policy gradient algorithm.
3. The robot path planning method according to claim 1, characterized in that: The state data includes: the distance between the robot and the obstacle, the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position. The connecting line is the line between the robot's current position and the target position.
4. The robot path planning method according to claim 3, characterized in that: The AMCL algorithm is used to determine the real-time position and posture of the robot to obtain the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position.
5. A robot path planning system, comprising a processor, characterized in that: The processor executes a computer program to implement the following steps: 1) obtaining state data of the robot at the current moment and several previous consecutive moments; 2) inputting the obtained state data into the ST-Transformer model, jointly encoding temporal dynamics and spatial correlation through the dual-path spatiotemporal attention mechanism of the ST-Transformer model, generating a spatiotemporal fused feature representation, and inputting the feature representation into the Actor network in the trained dual-delay deep deterministic policy gradient algorithm to obtain the optimal action of the robot for path planning at the next moment.
6. The robot path planning system according to claim 5, characterized in that: The untrained Actor network is trained according to the pre-collected state data to obtain the corresponding action. The Critic network in the double-delay deep deterministic policy gradient algorithm obtains the corresponding value according to the action and the pre-collected state data. The untrained Actor network updates its own parameters according to the value, thereby forming the trained Actor network in the double-delay deep deterministic policy gradient algorithm.
7. The robot path planning system according to claim 5, characterized in that: The state data includes: the distance between the robot and the obstacle, the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position. The connecting line is the line between the robot's current position and the target position.
8. The robot path planning system according to claim 7, characterized in that: The AMCL algorithm is used to determine the real-time position and posture of the robot to obtain the yaw angle, the angle between the robot's movement direction and the connecting line, the angle difference between the yaw angle and the angle, and the relative position distance between the robot and the target position.