Adaptive Robot Trajectory Planning Method and System Based on Deep Reinforcement Learning

Deep reinforcement learning is used to enhance robot trajectory planning in dynamic environments by processing real-time data and optimizing trajectories, addressing inaccuracies and inefficiencies in existing methods.

CN120095834BActive Publication Date: 2025-07-15CHENGDU AEROSPACE KAITE ELECTROMECHANICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510582778.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-15
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing robot trajectory planning methods are difficult to fully obtain environmental interaction states and motion postures in dynamic environments, and lack adaptability, resulting in inaccurate trajectory planning and low execution efficiency, making it difficult to meet the adaptability requirements of dynamic environments.

Method used

Using a method based on deep reinforcement learning, dynamic trajectory feature extraction and multi-dimensional trajectory optimization are performed by obtaining real-time trajectory data sets, adaptive trajectory optimization strategies are generated, and incremental trajectory correction is performed, and synchronized to the robot motion control system to achieve accurate trajectory execution.

Benefits of technology

It significantly improves the motion adaptability and execution efficiency of the robot in a dynamic environment, and can adjust the trajectory in real time according to dynamic changes, ensuring efficient motion planning and execution of the robot in a complex environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120095834B_ABST
    Figure CN120095834B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to robot path planning. Specifically, it relates to an adaptive robot trajectory planning method and system based on deep reinforcement learning. The method can completely present the movement of the robot by obtaining a real-time trajectory data set covering environmental interaction states and motion posture sequences; key features can be accurately refined through dynamic trajectory feature extraction and processing; an adaptive trajectory optimization strategy is generated based on a pre-trained deep reinforcement learning policy network, enabling the strategy to be flexibly adjusted according to dynamic features; incremental trajectory correction processing is performed on the robot motion posture sequence, which can gradually optimize the trajectory and generate an optimized trajectory sequence that conforms to the adaptability of the dynamic environment; the optimized trajectory sequence is synchronized to the motion control system, enabling the robot to quickly and accurately execute the optimized trajectory, significantly enhancing the motion planning ability and execution efficiency of the robot in complex dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to robot path planning. Specifically, it relates to an adaptive robot trajectory planning method and system based on deep reinforcement learning. Background Art

[0002] In the field of robot trajectory planning, as the application scenarios of robots become increasingly complex and diverse, the demand for their ability to autonomously plan and execute trajectories in dynamic environments is continuously increasing. In the early stage, robot trajectory planning mainly targeted static environments and completed tasks by presetting fixed paths. With the development of technology, attempts have been made to deal with dynamic environments, but mostly simple reactive strategies are adopted.

[0003] However, in the face of complex dynamic environments, traditional methods are difficult to comprehensively obtain the environmental interaction states and accurate motion postures of robots, resulting in a lack of comprehensiveness in trajectory planning. Moreover, existing trajectory optimization strategies lack sufficient adaptability and are difficult to adjust in real time according to dynamic changes, and cannot generate precise and efficient trajectories. At the same time, traditional trajectory correction methods are relatively rough and cannot achieve incremental fine-tuning, making it easy for robots to deviate when executing trajectories and difficult to meet the requirements of dynamic environment adaptability. In view of this, how to improve the trajectory planning and execution capabilities of robots in dynamic environments is a technical problem that needs to be solved at present. Summary of the Invention

[0004] In order to at least overcome the above deficiencies in the prior art, one of the purposes of this application is to provide an adaptive robot trajectory planning method and system based on deep reinforcement learning.

[0005] An embodiment of this application provides an adaptive robot trajectory planning method based on deep reinforcement learning, including: obtaining a real-time trajectory data set of a target robot in a dynamic environment, where the real-time trajectory data set includes sequences of environmental interaction states and sequences of robot motion postures at multiple consecutive time stamps; performing dynamic trajectory feature extraction processing on the real-time trajectory data set to obtain a dynamic trajectory feature set; based on a pre-trained deep reinforcement learning policy network, performing multi-dimensional trajectory optimization strategy generation processing on the dynamic trajectory feature set to obtain an adaptive trajectory optimization strategy; performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that meets the dynamic environment adaptability, and synchronizing the optimized trajectory sequence to the robot motion control system to trigger a trajectory execution operation.

[0006] An embodiment of the present application further provides an adaptive robot trajectory planning system, including a processor, a memory, and a bus connected to the processor; wherein, the processor and the memory complete communication with each other through the bus; the processor is used to call program instructions in the memory to execute the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning.

[0007] An embodiment of the present application further provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning.

[0008] The adaptive robot trajectory planning method and system based on deep reinforcement learning provided by the embodiment of the present application can effectively improve the motion adaptability of the robot in a dynamic environment. Specifically, by obtaining a real-time trajectory data set covering environmental interaction states and motion pose sequences, the motion situation of the robot can be fully presented; through dynamic trajectory feature extraction and processing, key features can be accurately refined; based on a pre-trained deep reinforcement learning policy network, an adaptive trajectory optimization policy is generated, enabling the policy to be flexibly adjusted according to dynamic features; by performing incremental trajectory correction processing on the robot motion pose sequence, the trajectory can be gradually optimized to generate an optimized trajectory sequence that meets the dynamic environment adaptability; synchronizing the optimized trajectory sequence to the motion control system enables the robot to quickly and accurately execute the optimized trajectory, significantly enhancing the motion planning ability and execution efficiency of the robot in a complex dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of an adaptive robot trajectory planning method based on deep reinforcement learning provided by an embodiment of the present application.

[0011] Figure 2 It is a block diagram of an adaptive robot trajectory planning system provided by an embodiment of the present application.

[0012] ICON:

[0013] 100 - Adaptive robot trajectory planning system;

[0014] 101 - Processor; 102 - Memory; 103 - Bus. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Exemplary embodiments disclosed in the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0016] To better understand the above technical solution, the technical solution of the present application will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0017] Figure 1 FIG. is a flowchart of an adaptive robot trajectory planning method based on deep reinforcement learning according to an embodiment of the present application, which is applied to an adaptive robot trajectory planning system and includes steps 110-step 140.

[0018] Step 110: Obtain a real-time trajectory data set of a target robot in a dynamic environment, where the real-time trajectory data set includes an environmental interaction state sequence and a robot motion posture sequence under multiple consecutive time stamps.

[0019] In an embodiment of the present application, for a dynamic environment, taking a logistics warehouse scenario as an example. The target robot is responsible for cargo handling work in this environment. The process of obtaining the real-time trajectory data set is as follows: In a continuous time process, corresponding data is recorded at each time stamp.

[0020] For the environmental interaction state sequence, it details the interaction between the robot and the surrounding environment. For example, at one time stamp, information about obstacles is recorded, including the position information of the obstacles relative to the robot, such as the relative orientation in the front, back, left, and right, as well as the shape description of the obstacles, such as the length, width, height, and other dimensional information, and the motion state of the obstacles, such as the moving speed and moving direction. For the robot motion posture sequence, the angle information of each joint of the robot is recorded to describe the limb posture of the robot, and at the same time, information such as the overall moving speed and moving direction of the robot is recorded.

[0021] By continuously recording data at each time stamp, an environmental interaction state sequence and a robot motion posture sequence under multiple consecutive time stamps are finally formed, jointly constituting the real-time trajectory data set.

[0022] Step 120: Perform dynamic trajectory feature extraction processing on the real-time trajectory data set to obtain a dynamic trajectory feature set.

[0023] In this embodiment, the dynamic trajectory feature set includes the obstacle distribution feature, the motion constraint feature, and the trajectory continuity feature. The dynamic trajectory feature extraction process is performed on the real-time trajectory data set to obtain the dynamic trajectory feature set, including:

[0024] Step 121: Extract the obstacle distribution feature from the environmental interaction state sequence. The obstacle distribution feature includes the boundary of the obstacle aggregation area, the direction of the obstacle movement trend, and the safe avoidance distance of the obstacle.

[0025] When extracting the obstacle distribution feature from the environmental interaction state sequence, first determine the boundary of the obstacle aggregation area, which requires analyzing the position information of each obstacle recorded in the sequence. The DBSCAN density clustering algorithm is used to regard adjacent and closely spaced obstacles as an aggregation area. By calculating the distances between the positions of these obstacles and setting a distance threshold, if the distance between two obstacles is less than the threshold, it is considered that they belong to the same aggregation area. Then, through the comprehensive analysis of the positions of all obstacles in the aggregation area, determine the boundary range of the aggregation area, for example, use the minimum bounding rectangle method to determine the boundary. For the direction of the obstacle movement trend, observe the position changes of the obstacle at multiple time stamps, calculate its displacement vector, and the direction of the displacement vector is the direction of the obstacle movement trend. The determination of the safe avoidance distance of the obstacle needs to consider factors such as the size of the robot, its movement speed, and the possible collision consequences. According to the size information of the robot and its movement speed, calculate the distance at which the robot can safely brake at different speeds, and then comprehensively set an adaptive safety margin to determine the safe avoidance distance of the obstacle.

[0026] Step 122: Extract the motion constraint feature from the robot motion posture sequence. The motion constraint feature includes the range of joint motion angle limits, the threshold of the end path curvature limit, and the speed mutation suppression interval.

[0027] When extracting motion constraint features from the robot motion posture sequence, the determination of the joint motion angle limit range is based on the mechanical structure design of the robot. Different joints have different activity range limitations. By referring to the robot's mechanical design documents, the minimum and maximum angles that each joint can move are obtained, so as to determine the joint motion angle limit range. For the end path curvature limit threshold, the motion characteristics and working requirements of the robot end effector are considered. When the end effector executes a task, the curvature of its motion path cannot be too large, otherwise it may cause the task execution to fail or damage the robot. Through the motion analysis of the robot end effector in various typical tasks and combining with existing empirical data, a suitable end path curvature limit threshold is determined. The determination of the speed mutation suppression interval is to ensure the smoothness of the robot's motion. For example, the Kalman filtering algorithm is used to analyze the motion stability of the robot under different speed change conditions, considering the dynamic characteristics of the robot, such as inertia, friction and other factors, to determine a reasonable range of speed change, that is, the speed mutation suppression interval. Within this interval, the robot can smoothly change its speed.

[0028] Step 123: Perform trajectory continuity detection on the temporal alignment relationship between the environmental interaction state sequence and the robot motion posture sequence to generate trajectory continuity features.

[0029] In this embodiment, the performing trajectory continuity detection on the temporal alignment relationship between the environmental interaction state sequence and the robot motion posture sequence to generate trajectory continuity features includes:

[0030] Step 1231: Align the environmental interaction state sequence and the robot motion posture sequence according to timestamps to generate a set of synchronized trajectory data segments.

[0031] Optionally, when aligning the environmental interaction state sequence and the robot motion posture sequence according to timestamps, taking each timestamp as a reference, combine the data in the environmental interaction state sequence and the data in the robot motion posture sequence at the same timestamp. For example, at time t, combine the obstacle information, environmental light information, etc. in the environmental interaction state sequence at this moment with the joint angles, speeds, etc. in the robot motion posture sequence at this moment to form a synchronized data unit. In this way, all timestamps are processed, and the synchronized data units are arranged in sequence to generate a set of synchronized trajectory data segments.

[0032] Step 1232: Perform trajectory segmentation processing on each synchronized trajectory data segment to obtain multiple trajectory segments, and perform trajectory conflict detection on each trajectory segment to generate trajectory conflict features; the trajectory conflict features include the conflict time window, conflict position area and conflict type label of the dynamic obstacle invading the trajectory segment.

[0033] When performing trajectory segmentation processing on each synchronized trajectory data segment, segmentation is carried out according to the changes in the robot's motion posture and the changes in the environmental interaction state. For example, when the robot's motion direction changes significantly, or a new obstacle appears in the environment, affecting the robot's movement, the trajectory is segmented at these key points to obtain multiple trajectory segments. For each trajectory segment, trajectory conflict detection is performed. First, the motion trajectory of the dynamic obstacle is determined. By analyzing the obstacle position information in the environmental interaction state sequence at multiple time stamps, a linear prediction algorithm is used to predict the motion trajectory of the obstacle in the future for a period of time. Then, the trajectory segment of the robot is compared with the predicted trajectory of the dynamic obstacle to determine whether there is an intersection. If there is an intersection, the conflict time window is determined, that is, the time range from when the robot enters the area where a conflict may occur to when it leaves the area; the conflict position area is determined, that is, the spatial range where the trajectory intersection is located; according to the specific situation of the conflict, such as whether it is a frontal collision, a side collision, etc., a label is added to the conflict type, thereby generating trajectory conflict characteristics.

[0034] Step 1233: Evaluate the feasibility of smooth transition of the connection points between adjacent trajectory segments according to the trajectory conflict characteristics, and generate a trajectory interruption risk level and a trajectory direction mutation probability.

[0035] When evaluating the feasibility of smooth transition of the connection points between adjacent trajectory segments according to the trajectory conflict characteristics, for the evaluation of the trajectory interruption risk level, consider the severity of the conflict and the location where the conflict occurs. If the conflict occurs at a key position of the trajectory segment, such as a position close to the target point, or the conflict causes the robot to need to change its motion direction significantly, then the trajectory interruption risk level is high; on the contrary, if the conflict occurs at a relatively unimportant position and the robot can avoid the conflict through minor adjustments, then the trajectory interruption risk level is low. For the evaluation of the trajectory direction mutation probability, analyze the direction change situation of adjacent trajectory segments at the connection point. If the robot's motion direction changes sharply at the connection point, such as the angle change exceeds the target threshold, then the trajectory direction mutation probability is high; if the direction change is relatively gentle, then the trajectory direction mutation probability is low. By comprehensively considering these factors, a trajectory interruption risk level and a trajectory direction mutation probability are generated.

[0036] Step 1234: Generate the trajectory continuity characteristics based on the trajectory interruption risk level and the trajectory direction mutation probability, and the trajectory continuity characteristics are used to mark the set of trajectory segments that need to be continuously repaired in the optimized trajectory sequence.

[0037] When generating the trajectory continuity feature based on the trajectory interruption risk level and the trajectory direction mutation probability, corresponding evaluation criteria are set. For example, when the trajectory interruption risk level exceeds one of the set high-risk thresholds and the trajectory direction mutation probability also exceeds one of the set high-probability thresholds, the corresponding trajectory segment is marked as a segment that needs to be repaired for continuity. By performing the above evaluations and markings on all trajectory segments, the trajectory continuity feature is finally generated, which can characterize which trajectory segments in the optimized trajectory sequence need to be repaired for continuity, so as to perform subsequent optimization processing on the trajectory to ensure the continuity and stability of the robot's motion trajectory.

[0038] Step 130: Based on the pre-trained deep reinforcement learning policy network, perform multi-dimensional trajectory optimization policy generation processing on the dynamic trajectory feature set to obtain an adaptive trajectory optimization policy.

[0039] In one implementation, the performing multi-dimensional trajectory optimization policy generation processing on the dynamic trajectory feature set based on the pre-trained deep reinforcement learning policy network to obtain an adaptive trajectory optimization policy includes:

[0040] Step 131: Input the dynamic trajectory feature set into the deep reinforcement learning policy network, and generate a joint optimization decision vector through a multi-level feature fusion module.

[0041] The deep reinforcement learning policy network in the embodiment of this application adopts a deep deterministic policy gradient network (DDPG). When inputting the dynamic trajectory feature set into the deep reinforcement learning policy network, first preprocess each feature in the dynamic trajectory feature set. For the obstacle distribution feature, quantify and normalize information such as the boundary of the obstacle aggregation area, the obstacle movement trend direction, and the obstacle safe avoidance distance it contains to meet the requirements of network input. For the motion constraint feature, also quantify and normalize information such as the joint motion angle limit range, the end path curvature limit threshold, and the speed mutation suppression interval. For the trajectory continuity feature, adaptively transform and normalize information such as the trajectory interruption risk level and the trajectory direction mutation probability. Then, input the preprocessed dynamic trajectory feature set into the multi-level feature fusion module of the deep reinforcement learning policy network. In the multi-level feature fusion module, features at different levels are gradually fused. For example, first combine and weighted sum low-level features pairwise, and then further fuse the obtained results with higher-level features. Through multiple above fusion processes, a joint optimization decision vector is finally generated.

[0042] Step 132: Generate a trajectory optimization priority sequence according to the trajectory optimization weight allocation relationship in the joint optimization decision vector.

[0043] Among them, when generating the trajectory optimization priority sequence according to the trajectory optimization weight allocation relationship in the joint optimization decision vector, first analyze the trajectory optimization factors represented by each element in the joint optimization decision vector. For example, one of the elements in the vector can represent the weight of obstacle avoidance, and another element can represent the weight of trajectory smoothness, etc. Then, sort according to the magnitude relationship of these weights. The trajectory optimization task corresponding to the optimization factor with a larger weight has a higher priority, and the trajectory optimization task corresponding to the optimization factor with a smaller weight has a lower priority. Through the above sorting, a trajectory optimization priority sequence is generated to clarify the order of each task when performing trajectory optimization.

[0044] Step 133: Based on the trajectory optimization priority sequence, perform trajectory replanning on the trajectory conflict features in the dynamic trajectory feature set to generate a set of candidate trajectory correction solutions.

[0045] Among them, when performing trajectory replanning on the trajectory conflict features in the dynamic trajectory feature set based on the trajectory optimization priority sequence, process the trajectory conflict features in sequence according to the order of the trajectory optimization priority sequence. For trajectory conflicts with higher priorities, first analyze the specific situation of the conflict, such as the conflict location, conflict type, etc. According to the analysis results, combined with the motion ability of the robot and environmental information, try different trajectory adjustment methods. For example, if the conflict is caused by an obstacle in front, consider whether the robot can bypass the obstacle from the side, calculate new trajectory points, and generate a new trajectory segment to avoid the conflict. For each trajectory conflict, multiple possible trajectory correction solutions are generated in the above way, and all these solutions are collected to form a set of candidate trajectory correction solutions.

[0046] Step 134: Perform dynamic environment adaptability verification on the set of candidate trajectory correction solutions, screen out the optimized trajectory correction solutions that meet the preset constraint conditions, and encode the optimized trajectory correction solutions as the adaptive trajectory optimization strategy; the adaptive trajectory optimization strategy is used to synchronously indicate the trajectory smoothness optimization direction and the dynamic obstacle avoidance optimization direction.

[0047] In a preferred embodiment, the performing dynamic environment adaptability verification on the set of candidate trajectory correction solutions and screening out the optimized trajectory correction solutions that meet the preset constraint conditions includes:

[0048] Step 1341: Simulate the execution process of the candidate trajectory correction solutions in the dynamic environment and collect virtual environment interaction response data.

[0049] When simulating the execution process of a candidate trajectory correction scheme in a dynamic environment, a virtual environment model similar to the actual dynamic environment is constructed. In this virtual environment model, the same obstacle distribution, environmental interference factors, etc. as in the actual environment are set. The candidate trajectory correction scheme is input into the virtual environment model, and the simulation execution is carried out according to the robot motion trajectory and actions specified in the scheme. During the simulation execution process, virtual environment interaction response data is collected, and this data includes information such as the distance change between the robot and obstacles in the virtual environment, the force on each joint of the robot, and the speed and direction change of the robot.

[0050] Step 1342: Perform trajectory safety analysis on the virtual environment interaction response data to generate trajectory safety evaluation indicators; the trajectory safety evaluation indicators include the achievement rate of the minimum obstacle avoidance distance, the number of times the joint motion angle exceeds the limit, and the compliance rate of the end path curvature.

[0051] It can be understood that when performing trajectory safety analysis on the virtual environment interaction response data, for the achievement rate of the minimum obstacle avoidance distance, first determine the minimum distance between the robot and each obstacle during the simulation execution. Then, compare these minimum distances with the preset safe obstacle avoidance distance, and calculate the proportion of the actual minimum avoidance distance reaching the preset safe avoidance distance, that is, the achievement rate of the minimum obstacle avoidance distance. For the number of times the joint motion angle exceeds the limit, monitor the angle change of each joint of the robot during the simulation execution, and count the number of times the joint motion angle exceeds its limit range. For the compliance rate of the end path curvature, calculate the curvature of the motion path of the robot end effector during the simulation execution, compare it with the preset end path curvature limit threshold, and count the proportion of the path curvature meeting the threshold requirements, that is, the compliance rate of the end path curvature. Through these analyses, trajectory safety evaluation indicators are generated.

[0052] Step 1343: If any one of the trajectory safety evaluation indicators does not reach the corresponding preset safety threshold, perform trajectory parameter backtracking adjustment on the unqualified candidate trajectory correction scheme to generate an adjusted candidate trajectory correction scheme.

[0053] If any one of the trajectory safety evaluation indicators does not reach the corresponding preset safety threshold, for example, the achievement rate of the minimum obstacle avoidance distance is lower than the preset safety threshold, it means that the distance between the robot and the obstacle is too close during the simulation execution, and there is a collision risk. At this time, perform trajectory parameter backtracking adjustment on the unqualified candidate trajectory correction scheme. Analyze the reasons for the too-close distance. It may be that the motion prediction of the obstacle during trajectory planning is inaccurate, or the robot's motion ability is not fully considered when adjusting the trajectory. According to the analysis results, adjust the parameters of the trajectory, such as changing the turning angle of the robot, adjusting the speed, etc., to generate an adjusted candidate trajectory correction scheme.

[0054] Step 1344: Repeatedly execute the simulation and adjustment operations until the trajectory safety evaluation indicators of all candidate trajectory correction schemes meet the corresponding preset safety thresholds, and generate an optimized trajectory correction scheme that meets the preset constraint conditions.

[0055] When repeatedly executing the simulation and adjustment operations, input the adjusted candidate trajectory correction scheme into the virtual environment model again for simulation execution, collect new virtual environment interaction response data, and re - conduct trajectory safety analysis. If there are still indicators that do not reach the preset safety thresholds, continue to perform backtracking adjustment of the trajectory parameters of the scheme, and conduct simulation execution and analysis again. Loop like this until the trajectory safety evaluation indicators of all candidate trajectory correction schemes, namely the achievement rate of the minimum obstacle avoidance distance, the number of times the joint motion angle exceeds the limit, and the compliance rate of the end - path curvature, all meet the corresponding preset safety thresholds. At this time, generate an optimized trajectory correction scheme that meets the preset constraint conditions.

[0056] Step 140: Perform incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy, generate an optimized trajectory sequence that meets the dynamic environment adaptability, and synchronize the optimized trajectory sequence to the robot motion control system to trigger the trajectory execution operation.

[0057] In an optional embodiment, the performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that meets the dynamic environment adaptability includes:

[0058] Step 141: Perform local trajectory interpolation processing on the trajectory conflict feature points in the robot motion posture sequence according to the trajectory optimization direction in the adaptive trajectory optimization strategy to generate a smooth transition trajectory segment.

[0059] When performing local trajectory interpolation processing on the trajectory conflict feature points in the robot motion posture sequence according to the trajectory optimization direction in the adaptive trajectory optimization strategy, first determine the positions of the trajectory conflict feature points. Then, select several reference points within a preset range before and after the conflict feature points according to the direction change trend indicated by the trajectory optimization direction. Using the cubic spline interpolation algorithm, generate a trajectory segment that can smoothly transition at the conflict feature points based on the position and posture information of these reference points. During the interpolation process, ensure that the generated trajectory segment is within the motion ability range of the robot and meets the requirements of the motion constraint characteristics.

[0060] Step 142: Perform dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment to generate an obstacle avoidance trajectory offset.

[0061] When performing dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment, the obstacle information in the dynamic environment is combined. First, it is determined whether there are dynamic obstacles in the forward direction of the smooth transition trajectory segment. If there are, the motion trajectory and speed of the obstacles are analyzed. According to the motion of the obstacles and the robot's safety avoidance distance requirements, the direction and distance that the robot needs to offset are calculated, that is, the obstacle avoidance trajectory offset is generated. During the calculation process, the robot's steering ability and speed adjustment ability should be considered to ensure that the offset is within a reasonable range.

[0062] Step 143: Perform spatial superposition of the smooth transition trajectory segment and the obstacle avoidance trajectory offset to generate an intermediate optimized trajectory segment.

[0063] When performing spatial superposition of the smooth transition trajectory segment and the obstacle avoidance trajectory offset, the obstacle avoidance trajectory offset is adjusted in terms of its direction and magnitude to perform spatial position adjustment on the basis of the smooth transition trajectory segment. For example, if the obstacle avoidance trajectory offset is a distance offset to the right, then each point on the smooth transition trajectory segment is spatially moved to the right according to this offset, thereby generating an intermediate optimized trajectory segment. During the superposition process, the continuity and feasibility of the trajectory should be ensured to avoid unreasonable trajectory mutations.

[0064] Step 144: Perform inverse kinematic verification on the intermediate optimized trajectory segment. When the inverse kinematic verification indicates that the joint motion sequence corresponding to the intermediate optimized trajectory segment is within the range of motion constraint characteristics, insert the intermediate optimized trajectory segment into the original robot motion posture sequence to replace the trajectory segment where the corresponding trajectory conflict feature point is located, and generate an incrementally corrected optimized trajectory sequence.

[0065] Among them, when performing inverse kinematic verification on the intermediate optimized trajectory segment, according to the robot's kinematic model, the D-H parameter method is used to calculate the position and posture information of the robot end effector corresponding to the intermediate optimized trajectory segment through the inverse kinematic algorithm to obtain the motion angle sequence of each joint. Then, the calculated joint motion angle sequence is compared with the joint motion angle limit range in the motion constraint characteristics. If the motion angles of all joints are within the limit range, it indicates that the joint motion sequence corresponding to the intermediate optimized trajectory segment is feasible. At this time, insert the intermediate optimized trajectory segment into the original robot motion posture sequence to replace the original trajectory segment where the trajectory conflict feature point is located, thereby generating an incrementally corrected optimized trajectory sequence.

[0066] In summary, through comprehensively obtaining data, accurately extracting features, generating an adaptive strategy, and finely correcting the trajectory, the embodiments of the present application significantly improve the trajectory planning and execution ability of the robot in a dynamic environment.

[0067] In an alternative embodiment, the training process of the deep reinforcement learning policy network includes:

[0068] Step 210: Obtain a simulation training data set containing various dynamic environment scenarios. Each training sample in the simulation training data set includes an environmental interaction state sequence, a robot motion posture sequence, and a corresponding ideal trajectory optimization strategy.

[0069] Among them, when obtaining the simulation training data set containing various dynamic environment scenarios, data is generated by constructing multiple different simulation environments. For example, create a warehouse environment where shelves with different layouts are set as obstacles, and handling vehicles with different speeds and motion trajectories are set as dynamic obstacles; then create an outdoor scene with different terrain undulations and moving pedestrians and other dynamic elements. In each simulation environment, run the robot and record its motion process. Record the environmental interaction state sequence at each timestamp, including information such as the position, size, and motion state of the obstacles; record the robot motion posture sequence, including information such as joint angles, speeds, and directions.

[0070] At the same time, according to the environmental conditions and the task requirements of the robot, corresponding ideal trajectory optimization strategies are set. These ideal trajectory optimization strategies can enable the robot to efficiently and safely complete tasks in this environment. Organize and combine these data to form a simulation training data set containing various dynamic environment scenarios, where each training sample includes an environmental interaction state sequence, a robot motion posture sequence, and a corresponding ideal trajectory optimization strategy.

[0071] Step 220: Input the training sample into the initial policy network for forward propagation calculation to generate a predicted trajectory optimization strategy.

[0072] In the embodiment of the present application, the initial policy network adopts a Deep Deterministic Policy Gradient network (DDPG). When inputting training samples into the initial policy network for forward propagation calculation, first, the environmental interaction state sequence and the robot motion pose sequence are preprocessed according to the network input requirements. For various information in the environmental interaction state sequence, such as the position and size of obstacles, numericalization and normalization processing are performed to enable it to enter the network as appropriate input data. The same preprocessing operations are also performed on information such as joint angles and speeds in the robot motion pose sequence. After the preprocessing is completed, these data are input into the input layer of the initial policy network (DDPG). The data propagates forward in the network according to the set network structure and weights. Taking DDPG as an example, the input layer receives the preprocessed data, performs operations with the weight matrix of the hidden layer, and undergoes a non-linear transformation through an activation function (such as the ReLU function) to transfer the information to the hidden layer. The hidden layer further processes the information, and finally, a predicted trajectory optimization strategy is obtained at the output layer. This predicted trajectory optimization strategy is a trajectory optimization scheme learned by the network based on the input training sample data.

[0073] Step 230: Calculate the policy difference loss value between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, and update the network parameters of the initial policy network through backpropagation based on the policy difference loss value; during the training process, in accordance with the preset stage training sequence, iteratively train using subsets of simulation training data with increasing obstacle movement speed and increasing environmental perturbation intensity in sequence; when the policy difference loss value of the initial policy network in the target complexity scenario falls within the preset convergence interval, terminate the training and generate a deep reinforcement learning policy network.

[0074] When calculating the policy difference loss value between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, first determine a suitable loss function, such as the Mean Squared Error loss function (MSE). For each corresponding element in the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, calculate the square of the difference between them, then sum up the squared differences of all elements, and divide by the total number of elements to obtain the policy difference loss value. This loss value reflects the degree of difference between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy. Based on this policy difference loss value, update the network parameters of the initial policy network (DDPG) through the backpropagation algorithm. The backpropagation algorithm starts from the output layer, calculates the error of the output layer according to the loss value, and then backpropagates the error to the previous layers. In each layer, adjust the weights and biases of this layer according to the error. For example, for the weight matrix of one layer, calculate the gradient of the weight according to the error and the input data of this layer, and then update the weight according to the gradient descent method, so that the weight changes in the direction of reducing the loss value.

[0075] During the training process, in accordance with the preset stage training sequence, iterative training is successively carried out using subsets of simulation training data with increasing obstacle movement speed and increasing environmental disturbance intensity. The preset stage training sequence is designed according to the training objectives and the characteristics of the data. First, the simulation training data set is divided into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameter and the environmental disturbance intensity parameter. Each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and the environmental disturbance intensity threshold of each stage data subset increase gradually with the increase of the stage number. For example, in the data subset of the first stage, the obstacle movement speed is slow and the environmental disturbance intensity is small; as the stage number increases, the obstacle movement speed in the subsequent stage data subsets gradually increases, and the environmental disturbance intensity also gradually increases. In the order from the lowest stage number to the highest, the stage data subsets are successively input into the initial policy network for stage training.

[0076] In each stage of training, based on the current stage data subset, multiple batches of policy predictions are made, and the stage matching degree between the predicted trajectory optimization policy and the corresponding ideal trajectory optimization policy is calculated. When the stage matching degree reaches the preset convergence criterion for the current stage, the training data is switched to the next stage data subset, and the network parameters completed in the current stage of training are used as the initialization parameters for the next stage. The operations of stage switching and parameter inheritance are repeatedly executed until the training of the highest stage data subset is completed.

[0077] In this way, the stability of the policy generation of the deep reinforcement learning policy network is gradually improved under the dynamically increasing interference intensity. When the policy difference loss value of the initial policy network in the target complexity scenario falls within the preset convergence interval, it indicates that the network has learned a policy that can adapt to the target complexity scenario. At this time, the training is terminated and the deep reinforcement learning policy network is generated. This deep reinforcement learning policy network can generate relatively effective trajectory optimization policies in dynamic environments with different complexities.

[0078] In a preferred embodiment, the iterative training using subsets of simulation training data with increasing obstacle movement speed and increasing environmental disturbance intensity in accordance with the preset stage training sequence includes:

[0079] Step 231: Divide the simulation training data set into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameter and the environmental disturbance intensity parameter; wherein, each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and the environmental disturbance intensity threshold of each stage data subset increase gradually with the increase of the stage number.

[0080] When dividing the simulation training data set according to the obstacle movement speed parameter and the environmental disturbance intensity parameter, first determine the value ranges of the obstacle movement speed and the environmental disturbance intensity. For example, the value range of the obstacle movement speed can be from low speed to high speed, and the environmental disturbance intensity can be from low intensity to high intensity. According to these value ranges, divide them into several intervals. For the obstacle movement speed, set different speed thresholds, such as speed threshold 1, speed threshold 2, etc., and classify the data set according to whether the obstacle movement speed is within these threshold ranges. For the environmental disturbance intensity, similarly set different intensity thresholds, such as intensity threshold 1, intensity threshold 2, etc., and classify the data set according to whether the environmental disturbance intensity is within these threshold ranges. Through the classification combination of these two parameters, divide the simulation training data set into stage data subsets corresponding to multiple training stages. Each stage data subset has a specific range of obstacle movement speed thresholds and environmental disturbance intensity thresholds, and as the stage serial number increases, these threshold ranges gradually increase to meet the requirement of gradually increasing difficulty during the training process.

[0081] Step 232: Input each stage data subset into the initial policy network in ascending order of the stage serial number for stage training.

[0082] When inputting each stage data subset into the initial policy network (DDPG) in ascending order of the stage serial number for stage training, start from the stage data subset with the lowest stage serial number. Input the training samples in this stage data subset into the initial policy network one by one, and perform training according to the method of forward propagation calculation and backpropagation to update network parameters described above. During the training process, record the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy. When the stage matching degree reaches the preset convergence criterion for the current stage, it indicates that the network has learned a better strategy on the data of the current stage. At this time, switch the training data to the next stage data subset, and use the network parameters trained in the current stage as the initialization parameters for the next stage. The purpose of doing this is to enable the network to further learn the strategy in a more complex environment based on the existing learning results. By performing the above training operations on each stage data subset in turn until the training of the highest stage data subset is completed, the network can adapt to dynamic environments with different difficulty levels.

[0083] In a preferred embodiment, each stage training process includes:

[0084] Step 2321: Perform multi-batch policy prediction based on the current stage data subset, and calculate the stage matching degree between the predicted trajectory optimization policy and the corresponding ideal trajectory optimization policy; when the stage matching degree reaches the convergence criterion preset for the current stage, switch the training data to the next stage data subset, and use the network parameters trained in the current stage as the initialization parameters for the next stage; loop through the stage switching and parameter inheritance operations until the training of the highest stage data subset is completed, so that the deep reinforcement learning policy network gradually improves the stability of policy generation under the dynamic interference of increasing intensity.

[0085] It can be understood that when performing multi-batch policy prediction based on the current stage data subset, the training samples in the current stage data subset are divided into multiple batches. For each batch of training samples, they are input into the initial policy network (DDPG) for forward propagation calculation to generate a predicted trajectory optimization policy. Then, for each predicted trajectory optimization policy, calculate the stage matching degree between it and the corresponding ideal trajectory optimization policy. There are various methods for calculating the stage matching degree, such as calculating the similarity index between the two, such as cosine similarity. Consider the predicted trajectory optimization policy and the ideal trajectory optimization policy as vectors, and measure their matching degree by calculating the cosine similarity between the vectors. When the calculated stage matching degree reaches the convergence criterion preset for the current stage, it means that the network has achieved a good learning effect on the data of the current stage. At this time, switch the training data to the next stage data subset, and at the same time pass the network parameters trained in the current stage to the next stage as initialization parameters. Through the above loop operation, each stage switch is based on the learning results of the previous stage, enabling the deep reinforcement learning policy network to continuously adjust and optimize its policy generation ability under the dynamic interference of increasing intensity, and gradually improve the stability of policy generation. As the training progresses, the network can better handle increasingly complex dynamic environments and generate more accurate and effective trajectory optimization policies.

[0086] After the above adjustments, the entire technical solution maintains consistency in the description of the deep reinforcement learning policy network and its training process.

[0087] In a non-limiting embodiment, after synchronizing the optimized trajectory sequence to the robot motion control system to trigger the trajectory execution operation, it further includes: collecting in real time the actual trajectory data set when the robot motion control system executes the optimized trajectory sequence; performing execution trajectory deviation analysis and processing on the actual trajectory data set to generate a dynamic trajectory deviation index set; based on the correlation between the dynamic trajectory deviation index set and the dynamic trajectory feature set, generating a trajectory dynamic compensation strategy through the deep reinforcement learning policy network; performing online trajectory parameter compensation processing on the optimized trajectory sequence according to the trajectory dynamic compensation strategy to generate a real-time compensation trajectory sequence and synchronizing it to the robot motion control system.

[0088] When collecting in real time the actual trajectory data set when the robot motion control system executes the optimized trajectory sequence, various sensors installed on the robot are used, such as position sensors, attitude sensors, etc. The position sensor can obtain the position information of the robot in space in real time, and the attitude sensor can obtain the attitude information of the robot, such as angles, directions, etc. During the process of the robot executing motion according to the optimized trajectory sequence, at regular time intervals, for example, every 0.1 second, the data of these sensors are collected, and these data are sorted and recorded to form the actual trajectory data set.

[0089] When performing execution trajectory deviation analysis and processing on the actual trajectory data set, first, the actual trajectory data is compared with the optimized trajectory sequence. For each point in the trajectory, the difference between the actual position and the corresponding point position in the optimized trajectory, as well as the difference between the actual attitude and the corresponding attitude in the optimized trajectory, are calculated. Based on these differences, a dynamic trajectory deviation index set is generated. For example, the dynamic trajectory deviation index set may include a position deviation index, that is, the distance deviation between the actual position and the optimized trajectory position; an attitude deviation index, that is, the angular deviation between the actual attitude and the optimized trajectory attitude, etc.

[0090] Based on the correlation between the dynamic trajectory deviation index set and the dynamic trajectory feature set, a trajectory dynamic compensation strategy is generated through the deep reinforcement learning policy network. Analyze the connection between the dynamic trajectory deviation index set and the dynamic trajectory feature set. For example, if the dynamic trajectory deviation index shows that the robot has a large deviation at a certain position, and the dynamic trajectory feature set indicates that there are specific types of obstacles or environmental interferences near this position, then this information is used as input and passed to the deep reinforcement learning policy network. Taking DDPG as an example, according to these input information, through its internal learning mechanism and calculation logic, after receiving the input, the network undergoes operations and processing of each layer, and through operations such as the evaluation of state and action values, generates a strategy that can compensate for the trajectory deviation, that is, the trajectory dynamic compensation strategy.

[0091] When performing online trajectory parameter compensation processing on the optimized trajectory sequence according to the trajectory dynamic compensation strategy, the trajectory parameters in the optimized trajectory sequence are adjusted according to the direction and adjustment amount indicated by the trajectory dynamic compensation strategy. For example, if the trajectory dynamic compensation strategy requires increasing the movement speed of the robot at a certain moment, then the speed parameter at that moment is correspondingly adjusted in the optimized trajectory sequence. Through the above adjustments, a real-time compensation trajectory sequence is generated and synchronized to the robot motion control system, enabling the robot to execute tasks more accurately according to the adjusted trajectory.

[0092] In a non-limiting embodiment, after synchronizing the optimized trajectory sequence to the robot motion control system to trigger a trajectory execution operation, it further includes: periodically obtaining an incremental environment interaction state sequence and an incremental motion posture sequence of the update cycle where the robot motion control system is located; performing incremental dynamic trajectory feature extraction processing on the incremental environment interaction state sequence and the incremental motion posture sequence to generate an incremental dynamic trajectory feature set; based on the superimposed relationship between the incremental dynamic trajectory feature set and the historical optimized trajectory sequence, generating an incremental trajectory optimization strategy through the deep reinforcement learning policy network; performing rolling horizon trajectory correction processing on the unexecuted part of the optimized trajectory sequence according to the incremental trajectory optimization strategy, generating a rolling updated trajectory sequence and synchronizing it to the robot motion control system.

[0093] When periodically obtaining the incremental environment interaction state sequence and the incremental motion posture sequence of the update cycle where the robot motion control system is located, a fixed time period is set, for example, every 1 minute as an update cycle. At the beginning of each update cycle, the data acquisition mechanism is started. For the incremental environment interaction state sequence, information on changes that occur in the environment during this cycle is obtained, such as whether new obstacles appear and whether the positions of existing obstacles have moved. For the incremental motion posture sequence, information on the posture changes of the robot relative to the previous cycle during this cycle is recorded, such as changes in joint angles and changes in the overall movement direction.

[0094] When performing incremental dynamic trajectory feature extraction on the incremental environmental interaction state sequence and the incremental motion posture sequence, a method similar to the previous dynamic trajectory feature extraction is adopted. For the incremental environmental interaction state sequence, extract the change features of the obstacle distribution therein, such as the positions and sizes of newly added obstacles, the new positions and moving directions of moving obstacles, etc., as the incremental obstacle distribution features. For the incremental motion posture sequence, extract the change features in the robot's motion constraints, such as whether the range of joint motion angle limits is adjusted, whether the speed mutation suppression interval is changed, etc., as the incremental motion constraint features. At the same time, analyze the temporal relationship between the incremental environmental interaction state sequence and the incremental motion posture sequence, detect whether there are changes in trajectory continuity, and generate incremental trajectory continuity features. Combine these incremental features together to form an incremental dynamic trajectory feature set.

[0095] Based on the superposition relationship between the incremental dynamic trajectory feature set and the historical optimized trajectory sequence, an incremental trajectory optimization strategy is generated through a deep reinforcement learning policy network. The incremental dynamic trajectory feature set and the historical optimized trajectory sequence are fused and analyzed, considering the situation during the execution of the historical optimized trajectory sequence and the impact of the current incremental environment and robot posture changes on the subsequent trajectory. Input this fused information into the deep reinforcement learning policy network (DDPG). The network will process this information according to its own structure and learning mechanism. For example, in DDPG, the input information undergoes computational processing through the input layer and hidden layer, and through the evaluation of states and actions and the calculation of Q values, an incremental trajectory optimization strategy that can adapt to new changes is finally generated.

[0096] When performing rolling horizon trajectory correction on the unexecuted part of the optimized trajectory sequence according to the incremental trajectory optimization strategy, starting from the current moment, adjust the unexecuted part of the optimized trajectory sequence. Modify the parameters of the unexecuted trajectory, such as the positions of trajectory points, the motion speed and direction of the robot, etc., according to the direction and adjustment amount indicated by the incremental trajectory optimization strategy. For example, if the incremental trajectory optimization strategy indicates that due to newly emerged obstacles, the robot needs to shift a certain distance to the left and appropriately reduce the speed in the next section of the trajectory, then adjust these parameters of the unexecuted part of the optimized trajectory sequence accordingly.

[0097] Through the above rolling horizon trajectory correction process, a rolling updated trajectory sequence is generated and synchronized to the robot motion control system. After receiving the rolling updated trajectory sequence, the robot motion control system controls the motion of the robot according to the new trajectory parameters, enabling the robot to dynamically adjust the subsequent motion trajectory according to the new environment and its own state changes, so as to better complete the task. This mechanism enables the robot to continuously optimize its motion trajectory when facing a constantly changing dynamic environment, improve the efficiency and safety of task execution, avoid collisions with newly emerging obstacles, and at the same time maintain the smoothness and continuity of motion.

[0098] When implementing the above technical solutions of the embodiments of the present application, those skilled in the art can, based on the prior art, reconstruct and optimize the implementation logic through deep reinforcement learning algorithms for dynamic environment modeling. Specifically, for the problem of mismatch with continuous trajectory control requirements, the DDPG (Deep Deterministic Policy Gradient) technology based on the Actor-Critic framework can be used to achieve continuous action output, and a multi-dimensional action space (such as curvature adjustment amount, speed change amount, and joint angle increment) is designed to adapt to the refined control of the robot motion posture. For the continuity contradiction caused by trajectory segmentation, the sliding window variance detection technology can be introduced to dynamically adjust the segmentation threshold, calculate the variance of the moving obstacle displacement direction (such as triggering segmentation when the angle standard deviation within the window exceeds 30°), and adaptively increase or decrease the segmentation sensitivity in combination with the historical trajectory repair success rate. In the virtual environment verification link, the dynamic parameter calibration can be achieved through the integration technology of a physical engine (such as Bullet), the actually collected obstacle motion data is fitted into a non-linear motion equation (such as a quadratic polynomial model), and the physical feasibility of the trajectory correction scheme is verified through the rigid body collision detection module, so as to eliminate the verification error caused by model simplification.

[0099] In terms of dimension unification and feature fusion, those skilled in the art can solve the multi-modal data compatibility problem through the classified Z-score standardization technology. For numerical features such as obstacle distance and joint angle, a normalization method of independently calculating the mean and standard deviation is adopted (such as using the average width of the warehouse passage 4m as the μ value for the obstacle distance), and the direction angle is converted into a two-dimensional vector and then normalized to the [-1, 1] interval to ensure the dimensional consistency of the neural network input layer. For the defect in the calculation of the safe avoidance distance, the kinematic braking formula reconstruction technology can be introduced, and the safe distance model (d_safe = v² / (2a_max) + size compensation term) is corrected based on the maximum deceleration parameter of the robot (such as the braking acceleration a_max = 3m / s² converted from the motor torque), and the dynamic correlation is established through the curvature-speed joint constraint equation (v_max = √(μg·r_min)), and the physical relationship between the ground friction coefficient μ and the minimum turning radius r_min is used to achieve the coordinated constraint of the trajectory curvature and the moving speed, avoiding the side slip risk during high-speed turning.

[0100] To further enhance the algorithm robustness, those skilled in the art can improve the adaptability to dynamic environments through the Extended Kalman Filter (EKF) and multi-modal training data fusion technology. The EKF technology is used to perform non-linear prediction on the motion trajectories of obstacles. A motion model is constructed through the state equation (including position and velocity components) and the observation equation, and the trajectory prediction accuracy is dynamically updated in combination with the covariance matrices of process noise and observation noise. During the training phase, the dataset can be extended through the multi-obstacle interaction scenario modeling technology. Typical scenarios such as intersections and narrow channels are constructed using path planning simulation tools (such as ROS Gazebo), and random moving obstacles (speed 1-3 m / s) and complex terrain data (such as 15° slopes and randomly raised ground) are injected. The training difficulty is gradually increased through the incremental curriculum learning strategy. For joint coupling conflicts, a linkage risk warning can be realized through a rigid body collision detection algorithm (such as GJK distance detection). The robot link is simplified into a cylinder model, and the minimum distance between it and the cube obstacle is calculated. When a potential collision is detected, a joint priority adjustment mechanism is triggered (such as the shoulder joint θ1 avoiding first), and the inverse kinematics verification technology is combined to ensure the mechanical reachability and dynamic safety of the corrected trajectory, and finally a complete optimized trajectory control system is formed.

[0101] The embodiments of this application can effectively improve the motion adaptability of the robot in dynamic environments. Specifically, by obtaining a real-time trajectory data set covering the environmental interaction state and motion posture sequence, the motion situation of the robot can be fully presented; through dynamic trajectory feature extraction and processing, key features can be accurately refined; an adaptive trajectory optimization strategy is generated based on a pre-trained deep reinforcement learning policy network, enabling the strategy to be flexibly adjusted according to dynamic features; incremental trajectory correction processing is performed on the robot motion posture sequence to gradually optimize the trajectory and generate an optimized trajectory sequence that meets the dynamic environment adaptability; the optimized trajectory sequence is synchronized to the motion control system, enabling the robot to quickly and accurately execute the optimized trajectory, significantly enhancing the motion planning ability and execution efficiency of the robot in complex dynamic environments.

[0102] The embodiments of this application provide a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the adaptive robot trajectory planning method based on deep reinforcement learning is implemented.

[0103] The embodiments of this application provide a processor, which is used to run a program. When the program runs, the adaptive robot trajectory planning method based on deep reinforcement learning is executed.

[0104] In the embodiments of this application, as Figure 2As shown, the adaptive robot trajectory planning system 100 includes at least one processor 101, at least one memory 102 connected to the processor 101, and a bus 103; wherein, the processor 101 and the memory 102 communicate with each other through the bus 103; the processor 101 is configured to call program instructions in the memory 102 to execute the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning.

[0105] This application is described with reference to the flowcharts and / or block diagrams of methods, adaptive robot trajectory planning systems (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for realizing the specified functions in one block or multiple blocks.

[0106] In a typical configuration, the adaptive robot trajectory planning system includes one or more processors (CPUs), a memory, and a bus. The adaptive robot trajectory planning system may also include an input / output interface, a network interface, etc.

[0107] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory includes at least one storage chip. The memory is an example of computer-readable media.

[0108] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage computer-readable storage media or any other non-transmission medium that can be used to store information that can be accessed by an adaptive robot trajectory planning system. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0109] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or computer-readable storage medium comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or computer-readable storage medium. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or computer-readable storage medium comprising the element.

[0110] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An adaptive robot trajectory planning method based on deep reinforcement learning, characterized in that Including: Obtaining a set of real-time trajectory data of a target robot in a dynamic environment, where the set of real-time trajectory data includes sequences of environmental interaction states and sequences of robot motion postures at multiple consecutive timestamps; Performing dynamic trajectory feature extraction processing on the set of real-time trajectory data to obtain a set of dynamic trajectory features; Based on a pre-trained deep reinforcement learning policy network, performing multi-dimensional trajectory optimization policy generation processing on the set of dynamic trajectory features to obtain an adaptive trajectory optimization policy; Performing incremental trajectory correction processing on the sequence of robot motion postures according to the adaptive trajectory optimization policy to generate an optimized trajectory sequence that meets the dynamic environment adaptability, and synchronizing the optimized trajectory sequence to the robot motion control system to trigger a trajectory execution operation; The performing incremental trajectory correction processing on the sequence of robot motion postures according to the adaptive trajectory optimization policy to generate an optimized trajectory sequence that meets the dynamic environment adaptability includes: Performing local trajectory interpolation processing on the trajectory conflict feature points in the sequence of robot motion postures according to the trajectory optimization direction in the adaptive trajectory optimization policy to generate a smooth transition trajectory segment; Performing dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment to generate an obstacle avoidance trajectory offset; Performing spatial superposition of the smooth transition trajectory segment and the obstacle avoidance trajectory offset to generate an intermediate optimized trajectory segment; Performing inverse kinematics verification on the intermediate optimized trajectory segment, and when the inverse kinematics verification indicates that the joint motion sequence corresponding to the intermediate optimized trajectory segment is within the range of motion constraint features, inserting the intermediate optimized trajectory segment into the original robot motion posture sequence to replace the trajectory segment where the corresponding trajectory conflict feature point is located, and generating an incrementally corrected optimized trajectory sequence.

2. The method according to claim 1, characterized in that, The set of dynamic trajectory features includes obstacle distribution features, motion constraint features, and trajectory continuity features. The performing dynamic trajectory feature extraction processing on the set of real-time trajectory data to obtain a set of dynamic trajectory features includes: Extracting obstacle distribution features from the sequence of environmental interaction states, where the obstacle distribution features include the boundaries of obstacle aggregation regions, the directions of obstacle movement trends, and the safe obstacle avoidance distances; Extracting motion constraint features from the sequence of robot motion postures, where the motion constraint features include the range of joint motion angle limits, the threshold of end path curvature limits, and the speed mutation suppression interval; Performing trajectory continuity detection on the temporal alignment relationship between the sequence of environmental interaction states and the sequence of robot motion postures to generate a trajectory continuity feature.

3. The method according to claim 2, characterized in that The performing trajectory continuity detection on the temporal alignment relationship between the sequence of environmental interaction states and the sequence of robot motion postures to generate a trajectory continuity feature includes: Aligning the sequence of environmental interaction states and the sequence of robot motion postures according to timestamps to generate a set of synchronized trajectory data segments; Perform trajectory segmentation processing on each synchronized trajectory data segment to obtain multiple trajectory segments, and perform trajectory conflict detection on each trajectory segment to generate trajectory conflict features; the trajectory conflict features include the conflict time window, conflict position area, and conflict type label where a dynamic obstacle intrudes into the trajectory segment; Evaluate the feasibility of smooth transition of the connection points between adjacent trajectory segments according to the trajectory conflict features to generate a trajectory interruption risk level and a trajectory direction mutation probability; Based on the trajectory interruption risk level and the trajectory direction mutation probability, generate the trajectory continuity feature, which is used to label the set of trajectory segments that need to be repaired for continuity in the optimized trajectory sequence.

4. The method according to claim 1, characterized in that, The multi-dimensional trajectory optimization strategy generation process for the dynamic trajectory feature set based on the pre-trained deep reinforcement learning policy network to obtain an adaptive trajectory optimization strategy includes: Input the dynamic trajectory feature set into the deep reinforcement learning policy network, and generate a joint optimization decision vector through a multi-level feature fusion module; Generate a trajectory optimization priority sequence according to the trajectory optimization weight allocation relationship in the joint optimization decision vector; Based on the trajectory optimization priority sequence, perform trajectory replanning processing on the trajectory conflict features in the dynamic trajectory feature set to generate a set of candidate trajectory correction plans; Perform dynamic environment adaptability verification on the set of candidate trajectory correction plans, screen out the optimized trajectory correction plans that meet the preset constraint conditions, and encode the optimized trajectory correction plans as the adaptive trajectory optimization strategy; Among them, the adaptive trajectory optimization strategy is used to synchronously indicate the trajectory smoothness optimization direction and the dynamic obstacle avoidance optimization direction.

5. The method according to claim 4, characterized in that, The dynamic environment adaptability verification of the set of candidate trajectory correction plans to screen out the optimized trajectory correction plans that meet the preset constraint conditions includes: Simulate the execution process of the candidate trajectory correction plan in the dynamic environment and collect virtual environment interaction response data; Perform trajectory safety analysis on the virtual environment interaction response data to generate trajectory safety evaluation indicators; the trajectory safety evaluation indicators include the achievement rate of the minimum obstacle avoidance distance, the number of times the joint motion angle exceeds the limit, and the compliance rate of the end path curvature; If any one of the trajectory safety evaluation indicators does not reach the corresponding preset safety threshold, perform trajectory parameter backtracking adjustment on the unqualified candidate trajectory correction plan to generate an adjusted candidate trajectory correction plan; Repeat the simulation and adjustment operations until the trajectory safety evaluation indicators of all candidate trajectory correction plans meet the corresponding preset safety thresholds to generate optimized trajectory correction plans that meet the preset constraint conditions.

6. The method according to claim 1, characterized in that, The training process of the deep reinforcement learning policy network includes: Obtain a simulation training data set containing various dynamic environment scenarios, and each training sample in the simulation training data set includes an environment interaction state sequence, a robot motion posture sequence, and a corresponding ideal trajectory optimization strategy; Input the training sample into the initial policy network for forward propagation calculation to generate a predicted trajectory optimization strategy; Calculate the policy difference loss value between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, and update the network parameters of the initial policy network through backpropagation based on the policy difference loss value; During the training process, in accordance with the preset stage training sequence, iteratively train using subsets of simulation training data with increasing obstacle movement speed and increasing environmental disturbance intensity in turn; When the policy difference loss value of the initial policy network in the target complexity scenario falls within the preset convergence interval, terminate the training and generate a deep reinforcement learning policy network.

7. The method according to claim 6, wherein The iterative training using subsets of simulation training data with increasing obstacle movement speed and increasing environmental disturbance intensity in turn according to the preset stage training sequence includes: Divide the simulation training data set into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameter and the environmental disturbance intensity parameter; wherein, each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and the environmental disturbance intensity threshold of each stage data subset increase step by step as the stage serial number increases; Input the stage data subsets into the initial policy network for stage training in the order of increasing stage serial numbers.

8. The method according to claim 7, wherein Each stage training process includes: Perform multi-batch policy prediction based on the current stage data subset, and calculate the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy; When the stage matching degree reaches the preset convergence criterion of the current stage, switch the training data to the next stage data subset, and use the network parameters completed in the current stage training as the initialization parameters for the next stage; Loop through the stage switching and parameter inheritance operations until the training of the highest stage data subset is completed, so that the deep reinforcement learning policy network gradually improves the policy generation stability under the dynamic interference of increasing intensity.

9. An adaptive robot trajectory planning system, characterized in that, It includes a processor, a memory and a bus connected to the processor; wherein, the processor and the memory complete communication with each other through the bus; the processor is used to call the program instructions in the memory to execute the adaptive robot trajectory planning method based on deep reinforcement learning according to any one of claims 1-8.

Citation Information

Patent Citations

  • Bionic robot control method and system based on AI learning

    CN119910648A

  • Multi-machine collaborative industrial robot intelligent scheduling system and application method

    CN119974019A

  • Method for controlling motions of quadruped robot based on reinforcement learning and position increment

    US20250021109A1