Adaptive robot trajectory planning method and system based on deep reinforcement learning
Through a method based on deep reinforcement learning, dynamic trajectory features are extracted and adaptive trajectory optimization strategies are generated, which solves the shortcomings of robot trajectory planning in the dynamic environment in the existing technology, and achieves more accurate and efficient trajectory planning and execution.
Patent Information
- Application Number
- CN202510582778.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing robot trajectory planning methods are difficult to fully obtain environmental interaction states and accurate motion postures in dynamic environments, resulting in a lack of comprehensiveness and adaptability in trajectory planning, making it difficult to generate accurate and efficient trajectories.
Adaptive robot trajectory planning method based on deep reinforcement learning is adopted, and dynamic trajectory feature extraction is performed by obtaining real-time trajectory data, adaptive trajectory optimization strategies are generated, and incremental trajectory correction is performed to generate an optimized trajectory sequence that conforms to the adaptability of the dynamic environment.
It significantly improves the motion adaptability and trajectory planning capabilities of the robot in dynamic environments, ensures the accuracy and efficiency of the trajectory, and enhances the execution efficiency of the robot in complex environments.
Smart Images

Figure CN120095834A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to robot path planning, and more specifically, to an adaptive robot trajectory planning method and system based on deep reinforcement learning. Background Art
[0002] In the field of robot trajectory planning, as robot application scenarios become more complex and diverse, the demand for its ability to autonomously plan and execute trajectories in dynamic environments continues to increase. In the early days, robot trajectory planning was mainly aimed at static environments, completing tasks by presetting fixed paths. With the development of technology, attempts to cope with dynamic environments have begun, but simple reactive strategies are mostly used.
[0003] However, when faced with complex dynamic environments, traditional methods find it difficult to fully obtain the robot's environmental interaction status and accurate motion posture, resulting in a lack of comprehensiveness in trajectory planning. Moreover, existing trajectory optimization strategies lack sufficient adaptive capabilities, making it difficult to adjust in real time according to dynamic changes, and unable to generate accurate and efficient trajectories. At the same time, traditional trajectory correction methods are relatively extensive and cannot achieve incremental fine adjustments, making it easy for robots to deviate when executing trajectories and difficult to meet the requirements of dynamic environment adaptability. In view of this, how to improve the robot's trajectory planning and execution capabilities in a dynamic environment is a technical problem that needs to be solved at present. Summary of the invention
[0004] In order to at least overcome the above-mentioned deficiencies in the prior art, one of the objectives of the present application is to provide an adaptive robot trajectory planning method and system based on deep reinforcement learning.
[0005] An embodiment of the present application provides an adaptive robot trajectory planning method based on deep reinforcement learning, comprising: obtaining a real-time trajectory data set of a target robot in a dynamic environment, the real-time trajectory data set comprising an environment interaction state sequence and a robot motion posture sequence under multiple continuous timestamps; performing dynamic trajectory feature extraction processing on the real-time trajectory data set to obtain a dynamic trajectory feature set; performing multi-dimensional trajectory optimization strategy generation processing on the dynamic trajectory feature set based on a pre-trained deep reinforcement learning strategy network to obtain an adaptive trajectory optimization strategy; performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that meets the adaptability of the dynamic environment, and synchronizing the optimized trajectory sequence to the robot motion control system to trigger a trajectory execution operation.
[0006] An embodiment of the present application also provides an adaptive robot trajectory planning system, comprising a processor and a memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning.
[0007] An embodiment of the present application also provides a computer-readable storage medium on which a program is stored, and when the program is executed by a processor, the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning is implemented.
[0008] The adaptive robot trajectory planning method and system based on deep reinforcement learning provided by the embodiment of the present application can effectively improve the robot's motion adaptability in a dynamic environment. In detail, by acquiring a real-time trajectory data set covering the environmental interaction state and motion posture sequence, the robot's motion situation can be fully presented; the key features can be accurately refined through dynamic trajectory feature extraction processing; an adaptive trajectory optimization strategy is generated based on a pre-trained deep reinforcement learning strategy network, so that the strategy can be flexibly adjusted according to dynamic features; the robot's motion posture sequence is incrementally corrected, and the trajectory can be gradually optimized to generate an optimized trajectory sequence that meets the adaptability of the dynamic environment; the optimized trajectory sequence is synchronized to the motion control system, so that the robot can quickly and accurately execute the optimized trajectory, significantly enhancing the robot's motion planning ability and execution efficiency in a complex dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0010] Figure 1 A flowchart of an adaptive robot trajectory planning method based on deep reinforcement learning provided in an embodiment of the present application.
[0011] Figure 2 A block diagram of an adaptive robot trajectory planning system provided in an embodiment of the present application.
[0012] icon:
[0013] 100-Adaptive robot trajectory planning system;
[0014] 101 - processor; 102 - memory; 103 - bus. DETAILED DESCRIPTION
[0015] The exemplary embodiments disclosed in the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0016] In order to better understand the above technical scheme, the technical scheme of the present application is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical scheme of the present application, rather than limitations on the technical scheme of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.
[0017] Figure 1 This is a flowchart of an adaptive robot trajectory planning method based on deep reinforcement learning provided according to an embodiment of the present application, which is applied to an adaptive robot trajectory planning system, including steps 110 to 140.
[0018] Step 110: Acquire a real-time trajectory data set of the target robot in a dynamic environment, wherein the real-time trajectory data set includes an environment interaction state sequence and a robot motion posture sequence at multiple continuous time stamps.
[0019] In the embodiment of the present application, for a dynamic environment, a logistics warehouse scene is taken as an example. The target robot is responsible for cargo handling in this environment. The acquisition process of the real-time trajectory data set is: in a continuous time process, each timestamp will record the corresponding data.
[0020] For the environmental interaction state sequence, it records the interaction between the robot and the surrounding environment in detail. For example, at one of the timestamps, relevant information about the obstacle is recorded, including the position information of the obstacle relative to the robot, such as the relative position of the front, back, left, and right, as well as the shape description of the obstacle, such as length, width, height and other dimensional information, and the movement state of the obstacle, such as the moving speed and moving direction. For the robot motion posture sequence, the angle information of each joint of the robot is recorded to describe the robot's limb posture, and the overall moving speed, moving direction and other information of the robot are recorded.
[0021] By continuously recording data at each timestamp, a sequence of environmental interaction states and a sequence of robot motion postures at multiple consecutive timestamps are eventually formed, which together constitute a real-time trajectory data set.
[0022] Step 120: Perform dynamic trajectory feature extraction processing on the real-time trajectory data set to obtain a dynamic trajectory feature set.
[0023] In this embodiment, the dynamic trajectory feature set includes the obstacle distribution feature, the motion constraint feature and the trajectory continuity feature. The dynamic trajectory feature extraction process is performed on the real-time trajectory data set to obtain the dynamic trajectory feature set, including:
[0024] Step 121: extracting obstacle distribution features from the environment interaction state sequence, wherein the obstacle distribution features include obstacle gathering area boundaries, obstacle movement trend directions, and obstacle safety avoidance distances.
[0025] When extracting obstacle distribution features from the environment interaction state sequence, the boundary of the obstacle clustering area is first determined, which requires analyzing the position information of each obstacle recorded in the sequence. The DBSCAN density clustering algorithm is used to regard adjacent and close obstacles as a clustering area. By calculating the distance between these obstacle positions, a distance threshold is set. If the distance between two obstacles is less than the threshold, they are considered to belong to the same clustering area. Then, by comprehensively analyzing the positions of all obstacles in the clustering area, the boundary range of the clustering area is determined, for example, the minimum enclosing rectangle is used to determine the boundary. For the direction of obstacle movement trend, the position change of the obstacle is observed at multiple timestamps, and its displacement vector is calculated. The direction of the displacement vector is the direction of obstacle movement trend. The determination of the obstacle safe avoidance distance needs to consider factors such as the size of the robot, the movement speed, and the possible consequences of collision. According to the size information of the robot and its movement speed, the distance at which the robot can safely brake at different speeds is calculated, and then the adaptive safety margin is comprehensively set to determine the obstacle safe avoidance distance.
[0026] Step 122: extracting motion constraint features from the robot motion posture sequence, wherein the motion constraint features include a joint motion angle limit range, a terminal path curvature limit threshold, and a speed mutation suppression interval.
[0027] When extracting motion constraint features from the robot motion posture sequence, the determination of the joint motion angle limit range is based on the robot's mechanical structure design. Different joints have different range limits. By consulting the robot's mechanical design documents, the minimum and maximum angles that each joint can move are obtained to determine the joint motion angle limit range. For the end path curvature limit threshold, the motion characteristics and working requirements of the robot's end effector are considered. When the end effector performs a task, the curvature of its motion path cannot be too large, otherwise it may cause task execution failure or damage the robot. By analyzing the motion of the robot's end effector in various typical tasks and combining existing empirical data, a suitable end path curvature limit threshold is determined. The determination of the speed mutation suppression interval is to ensure the stability of the robot's motion. For example, the Kalman filter algorithm is used to analyze the robot's motion stability under different speed changes. Considering the robot's dynamic characteristics, such as inertia, friction and other factors, a reasonable range of speed changes, namely the speed mutation suppression interval, is determined. Within this interval, the robot can change speed smoothly.
[0028] Step 123: Perform trajectory continuity detection on the time sequence alignment relationship between the environment interaction state sequence and the robot motion posture sequence to generate a trajectory continuity feature.
[0029] In this embodiment, the trajectory continuity detection is performed on the time sequence alignment relationship between the environment interaction state sequence and the robot motion posture sequence to generate a trajectory continuity feature, including:
[0030] Step 1231: Align the environment interaction state sequence and the robot motion posture sequence according to timestamps to generate a set of synchronous trajectory data segments.
[0031] Optionally, when aligning the environment interaction state sequence and the robot motion posture sequence by timestamp, the data in the environment interaction state sequence and the data in the robot motion posture sequence at the same timestamp are combined based on each timestamp. For example, at time t, the obstacle information, ambient lighting information, etc. at that time in the environment interaction state sequence are combined with the joint angle, speed, etc. at that time in the robot motion posture sequence to form a synchronization data unit. In this way, all timestamps are processed, and each synchronization data unit is arranged in sequence, thereby generating a set of synchronization trajectory data segments.
[0032] Step 1232: performing trajectory segmentation processing on each synchronized trajectory data segment to obtain multiple trajectory segments, and performing trajectory conflict detection on each trajectory segment to generate trajectory conflict features; the trajectory conflict features include a conflict time window, a conflict location area, and a conflict type label of a dynamic obstacle invading a trajectory segment.
[0033] When performing trajectory segmentation processing on each synchronous trajectory data segment, segmentation is performed according to the changes in the robot's motion posture and the changes in the environment interaction state. For example, when the robot's motion direction changes significantly, or a new obstacle appears in the environment that affects the robot's motion, the trajectory is segmented at these key points to obtain multiple trajectory segments. For each trajectory segment, trajectory conflict detection is performed. First, the motion trajectory of the dynamic obstacle is determined. By analyzing the obstacle position information in the environment interaction state sequence at multiple timestamps, a linear prediction algorithm is used to predict the obstacle's motion trajectory in the future. Then, the robot's trajectory segment is compared with the predicted trajectory of the dynamic obstacle to determine whether there is an intersection. If there is an intersection, the conflict time window is determined, that is, the time range from the robot entering the area where the conflict may occur to leaving the area; the conflict position area is determined, that is, the spatial range where the trajectory intersection is located; according to the specific situation of the conflict, such as whether it is a head-on collision, a side collision, etc., a label is added to the conflict type to generate a trajectory conflict feature.
[0034] Step 1233: Perform a smooth transition feasibility assessment on the connection points of adjacent trajectory segments according to the trajectory conflict characteristics, and generate a trajectory interruption risk level and a trajectory direction mutation probability.
[0035] When evaluating the feasibility of smooth transition of the connection points of adjacent trajectory segments based on trajectory conflict characteristics, the severity of the conflict and the location where the conflict occurs are considered for the assessment of the trajectory interruption risk level. If the conflict occurs at a critical position of the trajectory segment, such as a position close to the target point, or the conflict causes the robot to change its direction of movement significantly, then the trajectory interruption risk level is high; conversely, if the conflict occurs at a relatively unimportant position and the robot can avoid the conflict through minor adjustments, then the trajectory interruption risk level is low. For the assessment of the probability of trajectory direction mutation, the direction changes of adjacent trajectory segments at the connection point are analyzed. If the robot's direction of movement changes sharply at the connection point, such as the angle change exceeds the target threshold, the probability of trajectory direction mutation is high; if the direction change is relatively gentle, the probability of trajectory direction mutation is low. By comprehensively considering these factors, the trajectory interruption risk level and trajectory direction mutation probability are generated.
[0036] Step 1234: Based on the trajectory interruption risk level and the trajectory direction mutation probability, the trajectory continuity feature is generated, and the trajectory continuity feature is used to mark a set of trajectory segments that need to be repaired in continuity in the optimized trajectory sequence.
[0037] When generating trajectory continuity features based on the trajectory interruption risk level and the probability of trajectory direction mutation, set the corresponding evaluation criteria. For example, when the trajectory interruption risk level exceeds one of the set high risk thresholds and the probability of trajectory direction mutation also exceeds one of the set high probability thresholds, the corresponding trajectory segment is marked as a segment that needs to be repaired for continuity. By performing the above evaluation and marking on all trajectory segments, a trajectory continuity feature is finally generated. This feature can characterize which trajectory segments in the optimized trajectory sequence need to be repaired for continuity, so as to optimize the trajectory later and ensure the continuity and stability of the robot's motion trajectory.
[0038] Step 130: Based on the pre-trained deep reinforcement learning strategy network, a multi-dimensional trajectory optimization strategy generation process is performed on the dynamic trajectory feature set to obtain an adaptive trajectory optimization strategy.
[0039] In one implementation, the pre-trained deep reinforcement learning strategy network performs multi-dimensional trajectory optimization strategy generation processing on the dynamic trajectory feature set to obtain an adaptive trajectory optimization strategy, including:
[0040] Step 131: Input the dynamic trajectory feature set into the deep reinforcement learning strategy network, and generate a joint optimization decision vector through a multi-level feature fusion module.
[0041] The deep reinforcement learning policy network in the embodiment of the present application adopts a deep deterministic policy gradient network (DDPG). When the dynamic trajectory feature set is input into the deep reinforcement learning policy network, each feature in the dynamic trajectory feature set is first preprocessed. For the obstacle distribution feature, the information contained therein, such as the boundary of the obstacle gathering area, the direction of the obstacle movement trend, and the obstacle safe avoidance distance, is quantified and normalized to meet the requirements of the network input. For the motion constraint feature, the information such as the joint motion angle limit range, the terminal path curvature limit threshold, and the speed mutation suppression interval is also quantified and normalized. For the trajectory continuity feature, the information such as the trajectory interruption risk level and the trajectory direction mutation probability is adaptively converted and normalized. Then, the preprocessed dynamic trajectory feature set is input into the multi-level feature fusion module of the deep reinforcement learning policy network. In the multi-level feature fusion module, features at different levels are gradually fused. For example, the low-level features are firstly combined and weighted summed, and then the obtained results are further fused with the higher-level features. Through multiple fusion processes, a joint optimization decision vector is finally generated.
[0042] Step 132: Generate a trajectory optimization priority sequence according to the trajectory optimization weight distribution relationship in the joint optimization decision vector.
[0043] Among them, when generating the trajectory optimization priority sequence according to the trajectory optimization weight distribution relationship in the joint optimization decision vector, the trajectory optimization factors represented by each element in the joint optimization decision vector are first analyzed. For example, one of the elements in the vector can represent the weight of obstacle avoidance, and another element can represent the weight of trajectory smoothness, etc. Then, they are sorted according to the size relationship of these weights. The trajectory optimization tasks corresponding to the optimization factors with larger weights have higher priorities, and the trajectory optimization tasks corresponding to the optimization factors with smaller weights have lower priorities. Through the above sorting, the trajectory optimization priority sequence is generated to clarify the order of each task when performing trajectory optimization.
[0044] Step 133: Based on the trajectory optimization priority sequence, trajectory replanning processing is performed on the trajectory conflict features in the dynamic trajectory feature set to generate a set of candidate trajectory correction solutions.
[0045] Among them, when the trajectory conflict features in the dynamic trajectory feature set are replanned based on the trajectory optimization priority sequence, the trajectory conflict features are processed in sequence according to the order of the trajectory optimization priority sequence. For trajectory conflicts with higher priorities, the specific circumstances of the conflict are first analyzed, such as the conflict location, conflict type, etc. According to the analysis results, different trajectory adjustment methods are tried in combination with the robot's motion capabilities and environmental information. For example, if the conflict is caused by an obstacle in front, consider whether the robot can bypass the obstacle from the side, calculate new trajectory points, and generate a new trajectory segment to avoid the conflict. For each trajectory conflict, multiple possible trajectory correction schemes are generated in the above manner, and all these schemes are collected to form a set of candidate trajectory correction schemes.
[0046] Step 134: Verify the dynamic environment adaptability of the candidate trajectory correction scheme set, screen out the optimized trajectory correction scheme that meets the preset constraints, and encode the optimized trajectory correction scheme into the adaptive trajectory optimization strategy; the adaptive trajectory optimization strategy is used to synchronously indicate the trajectory smoothness optimization direction and the dynamic obstacle avoidance optimization direction.
[0047] In a preferred embodiment, the dynamic environment adaptability verification of the candidate trajectory correction scheme set is performed to select an optimized trajectory correction scheme that meets preset constraints, including:
[0048] Step 1341: Simulate the execution process of the candidate trajectory correction solution in a dynamic environment and collect virtual environment interactive response data.
[0049] When simulating the execution process of the candidate trajectory correction scheme in a dynamic environment, a virtual environment model similar to the actual dynamic environment is constructed. In this virtual environment model, the same obstacle distribution and environmental interference factors as the actual environment are set. The candidate trajectory correction scheme is input into the virtual environment model, and the robot motion trajectory and action specified in the scheme are simulated and executed. During the simulation execution process, the virtual environment interaction response data is collected, which includes the distance change between the robot and the obstacles in the virtual environment, the force of each joint of the robot, the speed and direction change of the robot, etc.
[0050] Step 1342: Perform trajectory safety analysis on the virtual environment interactive response data to generate trajectory safety evaluation indicators; the trajectory safety evaluation indicators include the obstacle minimum avoidance distance achievement rate, the number of joint motion angle exceeding the limit, and the terminal path curvature compliance rate.
[0051] It can be understood that when performing trajectory safety analysis on the interactive response data of the virtual environment, for the obstacle minimum avoidance distance achievement rate, first determine the minimum distance between the robot and each obstacle during the simulation execution. Then, compare these minimum distances with the preset obstacle safety avoidance distances, and calculate the proportion of the actual minimum avoidance distance that reaches the preset safety avoidance distance, that is, the obstacle minimum avoidance distance achievement rate. For the number of times the joint motion angle exceeds the limit, monitor the angle changes of each joint of the robot during the simulation execution, and count the number of times the joint motion angle exceeds its limit range. For the terminal path curvature compliance rate, calculate the curvature of the motion path of the robot's end effector during the simulation execution, compare it with the preset terminal path curvature limit threshold, and count the proportion of the path curvature that meets the threshold requirements, that is, the terminal path curvature compliance rate. Through these analyses, trajectory safety assessment indicators are generated.
[0052] Step 1343: If any one of the trajectory safety assessment indicators does not reach the preset safety threshold corresponding to the item, a trajectory parameter back-adjustment is performed on the candidate trajectory correction scheme that does not reach the standard to generate an adjusted candidate trajectory correction scheme.
[0053] If any of the trajectory safety assessment indicators does not reach the preset safety threshold corresponding to the item, for example, the obstacle minimum avoidance distance achievement rate is lower than the preset safety threshold, it means that the robot is too close to the obstacle during the simulation execution and there is a risk of collision. At this time, the trajectory parameters of the candidate trajectory correction scheme that does not meet the standard are retroactively adjusted. The reason for the close distance may be inaccurate prediction of the obstacle's motion during trajectory planning, or the robot's motion ability is not fully considered when adjusting the trajectory. According to the analysis results, the trajectory parameters are adjusted, such as changing the robot's steering angle, adjusting the speed, etc., to generate the adjusted candidate trajectory correction scheme.
[0054] Step 1344: Repeat the simulation and adjustment operations until the trajectory safety evaluation indicators of all candidate trajectory correction schemes meet the corresponding preset safety thresholds, and generate an optimized trajectory correction scheme that meets the preset constraints.
[0055] When repeating the simulation and adjustment operations, the adjusted candidate trajectory correction scheme is input into the virtual environment model again for simulation execution, new virtual environment interaction response data is collected, and the trajectory safety analysis is re-performed. If there are still indicators that do not meet the preset safety threshold, the trajectory parameters of the scheme are adjusted retroactively, and the simulation execution and analysis are performed again. This cycle is repeated until the trajectory safety evaluation indicators of all candidate trajectory correction schemes, namely the obstacle minimum avoidance distance achievement rate, the number of joint motion angle violations, and the terminal path curvature compliance rate, all meet the corresponding preset safety thresholds. At this time, an optimized trajectory correction scheme that meets the preset constraints is generated.
[0056] Step 140: performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that satisfies dynamic environment adaptability, and synchronizing the optimized trajectory sequence to the robot motion control system to trigger trajectory execution operations.
[0057] In an optional embodiment, performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that satisfies dynamic environment adaptability includes:
[0058] Step 141: According to the trajectory optimization direction in the adaptive trajectory optimization strategy, local trajectory interpolation processing is performed on the trajectory conflict feature points in the robot motion posture sequence to generate a smooth transition trajectory segment.
[0059] According to the trajectory optimization direction in the adaptive trajectory optimization strategy, when performing local trajectory interpolation processing on the trajectory conflict feature points in the robot motion posture sequence, the position of the trajectory conflict feature points is first determined. Then, according to the direction change trend indicated by the trajectory optimization direction, several reference points are selected within the preset range before and after the conflict feature points. Using the cubic spline interpolation algorithm, a trajectory segment that can smoothly transition at the conflict feature point is generated based on the position and posture information of these reference points. During the interpolation process, it is necessary to ensure that the generated trajectory segment is within the robot's motion capability and meets the requirements of the motion constraint characteristics.
[0060] Step 142: Perform dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment to generate an obstacle avoidance trajectory offset.
[0061] When performing dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment, the obstacle information in the dynamic environment is combined. First, determine whether there is a dynamic obstacle in the forward direction of the smooth transition trajectory segment. If so, analyze the movement trajectory and speed of the obstacle. According to the movement of the obstacle and the robot's safe avoidance distance requirements, calculate the direction and distance that the robot needs to offset, that is, generate the obstacle avoidance trajectory offset. During the calculation process, the robot's steering ability and speed adjustment ability must be considered to ensure that the offset is within a reasonable range.
[0062] Step 143: spatially superimpose the smooth transition trajectory segment and the obstacle avoidance trajectory offset to generate an intermediate optimized trajectory segment.
[0063] When the smooth transition trajectory segment is spatially superimposed with the obstacle avoidance trajectory offset, the obstacle avoidance trajectory offset is adjusted in spatial position based on the smooth transition trajectory segment according to its direction and size. For example, if the obstacle avoidance trajectory offset is a distance offset to the right, then each point on the smooth transition trajectory segment is spatially moved to the right according to this offset to generate an intermediate optimized trajectory segment. During the superposition process, the continuity and feasibility of the trajectory must be ensured to avoid unreasonable trajectory mutations.
[0064] Step 144: Perform kinematic inverse solution verification on the intermediate optimized trajectory segment. When the kinematic inverse solution verification indicates that the joint motion sequence corresponding to the intermediate optimized trajectory segment is within the limit range of the motion constraint feature, insert the intermediate optimized trajectory segment into the original robot motion posture sequence to replace the trajectory segment where the corresponding trajectory conflict feature point is located, and generate an incrementally corrected optimized trajectory sequence.
[0065] Among them, when the kinematic inverse solution is verified for the intermediate optimized trajectory segment, according to the robot's kinematic model, the DH parameter method is used to calculate the position and posture information of the robot end effector corresponding to the intermediate optimized trajectory segment, and the motion angle sequence of each joint is calculated through the kinematic inverse solution algorithm. Then, the calculated joint motion angle sequence is compared with the joint motion angle limit range in the motion constraint feature. If the motion angles of all joints are within the limit range, it means that the joint motion sequence corresponding to the intermediate optimized trajectory segment is feasible. At this time, the intermediate optimized trajectory segment is inserted into the original robot motion posture sequence, replacing the trajectory segment where the original trajectory conflict feature point is located, thereby generating an incrementally corrected optimized trajectory sequence.
[0066] In summary, the embodiments of the present application significantly improve the robot's trajectory planning and execution capabilities in a dynamic environment by comprehensively acquiring data, accurately extracting features, generating adaptive strategies, and finely correcting trajectories.
[0067] In an alternative embodiment, the training process of the deep reinforcement learning policy network includes:
[0068] Step 210: Acquire a simulation training data set containing multiple dynamic environment scenes, wherein each training sample in the simulation training data set includes an environment interaction state sequence, a robot motion posture sequence, and a corresponding ideal trajectory optimization strategy.
[0069] When obtaining a simulation training data set containing multiple dynamic environment scenes, data is generated by constructing multiple different simulation environments. For example, create a warehouse environment in which shelves with different layouts are set as obstacles, and transport vehicles with different speeds and motion trajectories are set as dynamic obstacles; then create an outdoor scene with dynamic elements such as different terrain undulations and moving pedestrians. In each simulation environment, run the robot and record its movement process. Record the sequence of environmental interaction states at each timestamp, including the location, size, motion state and other information of the obstacles; record the sequence of robot motion postures, including joint angles, speeds, directions and other information.
[0070] At the same time, according to the environmental conditions and the robot's task requirements, the corresponding ideal trajectory optimization strategies are set, which enable the robot to complete the task efficiently and safely in the environment. These data are sorted and combined to form a simulation training data set containing a variety of dynamic environment scenes, where each training sample contains an environmental interaction state sequence, a robot motion posture sequence and a corresponding ideal trajectory optimization strategy.
[0071] Step 220: Input the training samples into the initial strategy network for forward propagation calculation to generate a prediction trajectory optimization strategy.
[0072] The initial policy network in the embodiment of the present application adopts a deep deterministic policy gradient network (DDPG). When the training sample is input into the initial policy network for forward propagation calculation, the environment interaction state sequence and the robot motion posture sequence are first preprocessed according to the network input requirements. Various information in the environment interaction state sequence, such as the position and size of obstacles, are digitized and normalized so that they can enter the network as suitable input data. The same preprocessing operation is also performed for information such as joint angles and speeds in the robot motion posture sequence. After the preprocessing is completed, these data are input into the input layer of the initial policy network (DDPG). The data is forward propagated in the network according to the set network structure and weights. Taking DDPG as an example, the input layer receives the preprocessed data, calculates with the weight matrix of the hidden layer, and performs nonlinear transformation through the activation function (such as ReLU function) to pass the information to the hidden layer. The hidden layer further processes the information, and finally obtains the predicted trajectory optimization strategy in the output layer. This predicted trajectory optimization strategy is a trajectory optimization scheme learned by the network based on the input training sample data.
[0073] Step 230: Calculate the policy difference loss value between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, and update the network parameters of the initial policy network through back propagation based on the policy difference loss value; during the training process, according to the preset stage training sequence, use the simulation training data subsets with increasing obstacle movement speed and increasing environmental disturbance intensity to perform iterative training; when the policy difference loss value of the initial policy network in the target complexity scenario falls into the preset convergence interval, terminate the training and generate a deep reinforcement learning policy network.
[0074] When calculating the policy difference loss value between the predicted trajectory optimization policy and the ideal trajectory optimization policy, first determine a suitable loss function, such as the mean square error loss function (MSE). For each corresponding element in the predicted trajectory optimization policy and the ideal trajectory optimization policy, calculate the square of the difference between them, then sum the squares of the differences of all elements, and divide them by the total number of elements to obtain the policy difference loss value, which reflects the degree of difference between the predicted trajectory optimization policy and the ideal trajectory optimization policy. Based on this policy difference loss value, the network parameters of the initial policy network (DDPG) are updated by the back propagation algorithm. The back propagation algorithm starts from the output layer, calculates the error of the output layer according to the loss value, and then back propagates the error to the previous layers. In each layer, the weights and biases of the layer are adjusted according to the error. For example, for the weight matrix of one layer, the gradient of the weight is calculated according to the error and the input data of the layer, and then the weight is updated according to the gradient descent method so that the weight changes in the direction of reducing the loss value.
[0075] During the training process, according to the preset stage training sequence, the simulation training data subsets with increasing obstacle movement speed and increasing environmental disturbance intensity are used for iterative training. The preset stage training sequence is designed according to the training objectives and data characteristics. First, the simulation training data set is divided into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameters and environmental disturbance intensity parameters. Each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and environmental disturbance intensity threshold of each stage data subset are gradually increased as the stage number increases. For example, in the data subset of the first stage, the obstacle movement speed is slow and the environmental disturbance intensity is small; as the stage number increases, the obstacle movement speed in the subsequent stage data subset gradually increases, and the environmental disturbance intensity also gradually increases. According to the order of the stage number from low to high, the data subsets of each stage are input into the initial policy network for stage training.
[0076] In each stage of training, multiple batches of strategies are predicted based on the current stage data subset, and the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy is calculated. When the stage matching degree reaches the preset convergence standard of the current stage, the training data is switched to the next stage data subset, and the network parameters trained in the current stage are used as the initialization parameters of the next stage. The stage switching and parameter inheritance operations are executed cyclically until the training of the highest stage data subset is completed.
[0077] In this way, the deep reinforcement learning policy network gradually improves the stability of policy generation under the dynamic interference of increasing intensity. When the policy difference loss value of the initial policy network in the target complexity scenario falls into the preset convergence interval, it means that the network has learned a policy that can adapt to the target complexity scenario. At this time, the training is terminated and a deep reinforcement learning policy network is generated. This deep reinforcement learning policy network can generate a more effective trajectory optimization strategy in dynamic environments of different complexities.
[0078] In a preferred embodiment, the iterative training is performed in a preset stage training sequence by sequentially using simulation training data subsets with increasing obstacle movement speed and increasing environmental disturbance intensity, including:
[0079] Step 231: Divide the simulation training data set into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameter and the environmental disturbance intensity parameter; wherein each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and the environmental disturbance intensity threshold of each stage data subset are gradually increased as the stage number increases.
[0080] When dividing the simulation training data set according to the obstacle movement speed parameter and the environmental disturbance intensity parameter, first determine the value range of the obstacle movement speed and the environmental disturbance intensity. For example, the value range of the obstacle movement speed can be from low speed to high speed, and the environmental disturbance intensity can be from low intensity to high intensity. According to these value ranges, it is divided into several intervals. For the obstacle movement speed, different speed thresholds are set, such as speed threshold 1, speed threshold 2, etc., and the data set is classified according to whether the obstacle movement speed is within these threshold ranges. For the environmental disturbance intensity, different intensity thresholds are also set, such as intensity threshold 1, intensity threshold 2, etc., and the data set is classified according to whether the environmental disturbance intensity is within these threshold ranges. Through the classification combination of these two parameters, the simulation training data set is divided into stage data subsets corresponding to multiple training stages. Each stage data subset has a specific obstacle movement speed threshold range and environmental disturbance intensity threshold range, and as the stage number increases, these threshold ranges gradually increase to meet the needs of gradually increasing difficulty during training.
[0081] Step 232: In order of the stage numbers from low to high, the data subsets of each stage are sequentially input into the initial strategy network for stage training.
[0082] When the data subsets of each stage are input into the initial policy network (DDPG) in order from low to high stage numbers for stage training, start with the stage data subset with the lowest stage number. The training samples in the stage data subset are input into the initial policy network one by one, and training is performed according to the method of forward propagation calculation and back propagation updating network parameters described above. During the training process, the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy is recorded. When the stage matching degree reaches the preset convergence standard of the current stage, it means that the network has learned a better strategy on the data of the current stage. At this time, the training data is switched to the next stage data subset, and the network parameters trained in the current stage are used as the initialization parameters of the next stage. The purpose of this is to enable the network to further learn strategies in more complex environments based on the existing learning results. By performing the above training operations on each stage data subset in turn until the training of the highest stage data subset is completed, the network can adapt to dynamic environments of different difficulty levels.
[0083] In a preferred embodiment, each stage of the training process includes:
[0084] Step 2321: Perform multi-batch strategy prediction based on the current stage data subset, and calculate the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy; when the stage matching degree reaches the preset convergence standard of the current stage, switch the training data to the next stage data subset, and use the network parameters trained in the current stage as the initialization parameters of the next stage; loop the stage switching and parameter inheritance operations until the training of the highest stage data subset is completed, so that the deep reinforcement learning strategy network gradually improves the strategy generation stability under the dynamic interference of increasing intensity.
[0085] It can be understood that when multi-batch policy prediction is performed based on the current stage data subset, the training samples in the current stage data subset are divided into multiple batches. For each batch of training samples, they are input into the initial policy network (DDPG) for forward propagation calculation to generate a predicted trajectory optimization policy. Then, for each predicted trajectory optimization policy, the stage matching degree between it and the corresponding ideal trajectory optimization policy is calculated. The stage matching degree can be calculated in a variety of ways, such as calculating the similarity index between the two, such as cosine similarity. The predicted trajectory optimization policy and the ideal trajectory optimization policy are regarded as vectors, and their matching degree is measured by calculating the cosine similarity between the vectors. When the calculated stage matching degree reaches the preset convergence standard of the current stage, it means that the network has achieved a good learning effect on the data of the current stage. At this time, the training data is switched to the next stage data subset, and the network parameters trained in the current stage are passed to the next stage as initialization parameters. Through the above-mentioned cyclic operation, each stage switch is based on the learning results of the previous stage, so that the deep reinforcement learning policy network can continuously adjust and optimize its own strategy generation ability under the dynamic interference of increasing intensity, and gradually improve the stability of strategy generation. As training progresses, the network is able to better cope with increasingly complex dynamic environments and generate more accurate and effective trajectory optimization strategies.
[0086] After the above adjustments, the entire technical solution remains consistent in the description of the deep reinforcement learning policy network and its training process.
[0087] In a non-limiting embodiment, after synchronizing the optimized trajectory sequence to the robot motion control system to trigger the trajectory execution operation, it also includes: real-time acquisition of an actual trajectory data set when the robot motion control system executes the optimized trajectory sequence; performing trajectory deviation analysis processing on the actual trajectory data set to generate a dynamic trajectory deviation index set; based on the correlation between the dynamic trajectory deviation index set and the dynamic trajectory feature set, generating a trajectory dynamic compensation strategy through the deep reinforcement learning strategy network; performing online trajectory parameter compensation processing on the optimized trajectory sequence according to the trajectory dynamic compensation strategy, generating a real-time compensated trajectory sequence and synchronizing it to the robot motion control system.
[0088] When collecting the actual trajectory data set when the robot motion control system executes the optimized trajectory sequence in real time, various sensors installed on the robot are used, such as position sensors, attitude sensors, etc. The position sensor can obtain the robot's position information in space in real time, and the attitude sensor can obtain the robot's attitude information, such as angle, direction, etc. When the robot executes the movement according to the optimized trajectory sequence, the data of these sensors are collected at a certain time interval, such as 0.1 seconds, and these data are sorted and recorded to form an actual trajectory data set.
[0089] When performing trajectory deviation analysis on the actual trajectory data set, the actual trajectory data is first compared with the optimized trajectory sequence. For each point in the trajectory, the difference between the actual position and the position of the corresponding point in the optimized trajectory, as well as the difference between the actual posture and the corresponding posture in the optimized trajectory are calculated. Based on these differences, a dynamic trajectory deviation index set is generated. For example, the dynamic trajectory deviation index set can include a position deviation index, that is, the distance deviation between the actual position and the optimized trajectory position; a posture deviation index, that is, the angle deviation between the actual posture and the optimized trajectory posture, etc.
[0090] Based on the relationship between the dynamic trajectory deviation indicator set and the dynamic trajectory feature set, a trajectory dynamic compensation strategy is generated through a deep reinforcement learning policy network. The relationship between the dynamic trajectory deviation indicator set and the dynamic trajectory feature set is analyzed. For example, if the dynamic trajectory deviation indicator shows that the robot has a large deviation in one of its positions, and the dynamic trajectory feature set indicates that there are specific types of obstacles or environmental interference near the position, then this information is passed as input to the deep reinforcement learning policy network. Taking DDPG as an example, the deep reinforcement learning policy network, based on these input information, through its internal learning mechanism and calculation logic, after receiving the input, the network undergoes calculation processing at each layer, and generates a strategy that can compensate for the trajectory deviation based on operations such as the evaluation of state and action values, namely the trajectory dynamic compensation strategy.
[0091] When the optimized trajectory sequence is subjected to online trajectory parameter compensation processing according to the trajectory dynamic compensation strategy, the trajectory parameters in the optimized trajectory sequence are adjusted according to the direction and adjustment amount indicated by the trajectory dynamic compensation strategy. For example, if the trajectory dynamic compensation strategy requires increasing the robot's motion speed at one moment, then the speed parameter at that moment is adjusted accordingly in the optimized trajectory sequence. Through the above adjustment, a real-time compensation trajectory sequence is generated and synchronized to the robot motion control system, so that the robot can perform tasks more accurately according to the adjusted trajectory.
[0092] In a non-limiting embodiment, after synchronizing the optimized trajectory sequence to the robot motion control system to trigger the trajectory execution operation, it also includes: periodically obtaining the incremental environment interaction state sequence and the incremental motion posture sequence of the update period of the robot motion control system; performing incremental dynamic trajectory feature extraction processing on the incremental environment interaction state sequence and the incremental motion posture sequence to generate an incremental dynamic trajectory feature set; based on the superposition relationship between the incremental dynamic trajectory feature set and the historical optimized trajectory sequence, generating an incremental trajectory optimization strategy through the deep reinforcement learning strategy network; performing rolling time domain trajectory correction processing on the unexecuted part of the optimized trajectory sequence according to the incremental trajectory optimization strategy, generating a rolling update trajectory sequence and synchronizing it to the robot motion control system.
[0093] When periodically acquiring the incremental environment interaction state sequence and incremental motion posture sequence of the update cycle of the robot motion control system, a fixed time period is set, for example, every 1 minute as an update cycle. At the beginning of each update cycle, the data acquisition mechanism is started. For the incremental environment interaction state sequence, information about changes in the environment during this cycle is obtained, such as whether new obstacles appear, whether the positions of existing obstacles move, etc. For the incremental motion posture sequence, the posture change information of the robot in this cycle relative to the previous cycle is recorded, such as changes in joint angles, changes in the overall motion direction, etc.
[0094] When performing incremental dynamic trajectory feature extraction processing on the incremental environment interaction state sequence and the incremental motion posture sequence, a method similar to the previous dynamic trajectory feature extraction is used. For the incremental environment interaction state sequence, the change characteristics of the obstacle distribution are extracted, such as the position and size of the newly added obstacles, the new position and movement direction of the moving obstacles, etc., as the incremental obstacle distribution features. For the incremental motion posture sequence, the change characteristics of the robot motion constraints are extracted, such as whether the joint motion angle limit range is adjusted, whether the speed mutation suppression interval is changed, etc., as the incremental motion constraint features. At the same time, the temporal relationship between the incremental environment interaction state sequence and the incremental motion posture sequence is analyzed to detect whether there are changes in trajectory continuity and generate incremental trajectory continuity features. These incremental features are combined together to form an incremental dynamic trajectory feature set.
[0095] Based on the superposition relationship between the incremental dynamic trajectory feature set and the historical optimization trajectory sequence, the incremental trajectory optimization strategy is generated through the deep reinforcement learning policy network. The incremental dynamic trajectory feature set is fused and analyzed with the historical optimization trajectory sequence, considering the situation of the historical optimization trajectory sequence during execution and the impact of the current incremental environment and robot posture changes on the subsequent trajectory. The fused information is input into the deep reinforcement learning policy network (DDPG). The network will process this information according to its own structure and learning mechanism. For example, in DDPG, the input information is processed by the input layer and the hidden layer, and the incremental trajectory optimization strategy that can adapt to new changes is finally generated through the evaluation of the state and action and the calculation of the Q value.
[0096] When the rolling time domain trajectory correction processing is performed on the unexecuted part of the optimized trajectory sequence according to the incremental trajectory optimization strategy, the unexecuted part of the optimized trajectory sequence is adjusted starting from the current moment. The parameters of the unexecuted trajectory, such as the position of the trajectory point, the robot's movement speed and direction, are modified according to the direction and adjustment amount indicated by the incremental trajectory optimization strategy. For example, if the incremental trajectory optimization strategy indicates that the robot needs to shift a certain distance to the left and reduce the speed appropriately in the next section of the trajectory due to the emergence of a new obstacle, then these parameters of the unexecuted part of the optimized trajectory sequence are adjusted accordingly.
[0097] Through the above rolling time domain trajectory correction processing, a rolling update trajectory sequence is generated and synchronized to the robot motion control system. After receiving the rolling update trajectory sequence, the robot motion control system will control the robot's movement according to the new trajectory parameters, so that the robot can dynamically adjust the subsequent motion trajectory according to the new environment and its own state changes to better complete the task. This mechanism enables the robot to continuously optimize its own motion trajectory in the face of a constantly changing dynamic environment, improve the efficiency and safety of task execution, avoid collisions with newly emerging obstacles, and maintain the stability and continuity of movement.
[0098] When implementing the above technical solutions of the embodiments of the present application, those skilled in the art can realize logic correction through deep reinforcement learning algorithm reconstruction and dynamic environment modeling optimization based on the existing technology. In detail, in order to solve the problem of mismatch with the continuous trajectory control requirements, the DDPG (deep deterministic policy gradient) technology based on the Actor-Critic framework can be used to realize continuous action output, and the refined control of the robot's motion posture can be adapted by designing a multi-dimensional action space (such as curvature adjustment, speed change and joint angle increment). For the continuity contradiction caused by trajectory segmentation, the sliding window variance detection technology can be introduced to dynamically adjust the segmentation threshold, and the variance calculation of the displacement direction of the moving obstacle (such as the standard deviation of the angle in the window exceeds 30° to trigger segmentation) can be used to adaptively improve or reduce the segmentation sensitivity in combination with the historical trajectory repair success rate. In the virtual environment verification link, the dynamic parameter calibration can be realized through the integration technology of the physical engine (such as Bullet), the actual collected obstacle motion data can be fitted into a nonlinear motion equation (such as a quadratic polynomial model), and the physical feasibility of the trajectory correction scheme can be verified through the rigid body collision detection module, thereby eliminating the verification error caused by model simplification.
[0099] In terms of dimensional unification and feature fusion, technicians in this field can solve the compatibility problem of multimodal data through the type-based Z-score standardization technology. For numerical features such as obstacle distance and joint angle, a normalization method of independently calculating the mean and standard deviation is adopted (such as the obstacle distance is based on the average width of the warehouse channel 4m as the reference μ value), and the direction angle is converted into a two-dimensional vector and normalized to the interval [-1, 1] to ensure the dimensional consistency of the neural network input layer. In view of the defects in the calculation of safe avoidance distance, the kinematic braking formula reconstruction technology can be introduced to correct the safe distance model (d_safe=v² / (2a_max)+size compensation term) based on the robot's maximum deceleration parameter (such as the braking acceleration a_max=3m / s² converted from the motor torque), and establish a dynamic association through the curvature-speed joint constraint equation (v_max=√(μg·r_min)), and use the physical relationship between the ground friction coefficient μ and the minimum turning radius r_min to achieve the coordinated constraint of the trajectory curvature and the moving speed to avoid the risk of skidding during high-speed steering.
[0100] To further enhance the robustness of the algorithm, technicians in this field can improve the adaptability to dynamic environments through extended Kalman filtering (EKF) and multimodal training data fusion technology. The EKF technology is used to perform nonlinear prediction of the obstacle trajectory, and the motion model is constructed through the state equation (including position and velocity components) and the observation equation. The trajectory prediction accuracy is dynamically updated by combining the process noise and observation noise covariance matrix. During the training stage, the data set can be expanded through multi-obstacle interactive scene modeling technology, and path planning simulation tools (such as ROSGazebo) can be used to build typical scenes such as intersections and narrow passages. Random moving obstacles (speed 1-3m / s) and complex terrain data (such as 15° slopes and random raised ground) are injected, and the training difficulty is gradually increased through incremental course learning strategies. For joint coupling conflicts, the rigid body collision detection algorithm (such as GJK distance detection) can be used to realize linkage risk warning, simplify the robot link into a cylindrical model, calculate its minimum distance from the cubic obstacle, and trigger the joint priority adjustment mechanism (such as the upper arm joint θ1 priority avoidance) when a potential collision is detected. The kinematic inverse solution verification technology is combined to ensure the mechanical accessibility and dynamic safety of the corrected trajectory, and finally form a complete and optimized trajectory control system.
[0101] The embodiments of the present application can effectively improve the robot's motion adaptability in a dynamic environment. In detail, by acquiring a real-time trajectory data set covering the environmental interaction state and motion posture sequence, the robot's motion situation can be fully presented; the key features can be accurately refined through dynamic trajectory feature extraction processing; an adaptive trajectory optimization strategy is generated based on a pre-trained deep reinforcement learning strategy network, so that the strategy can be flexibly adjusted according to dynamic features; incremental trajectory correction processing is performed on the robot's motion posture sequence, which can gradually optimize the trajectory and generate an optimized trajectory sequence that meets the adaptability of the dynamic environment; the optimized trajectory sequence is synchronized to the motion control system, so that the robot can quickly and accurately execute the optimized trajectory, significantly enhancing the robot's motion planning ability and execution efficiency in a complex dynamic environment.
[0102] An embodiment of the present application provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the adaptive robot trajectory planning method based on deep reinforcement learning.
[0103] An embodiment of the present application provides a processor, which is used to run a program, wherein the program executes the adaptive robot trajectory planning method based on deep reinforcement learning when running.
[0104] In the present application embodiment, Figure 2As shown, the adaptive robot trajectory planning system 100 includes at least one processor 101, and at least one memory 102 and a bus 103 connected to the processor 101; wherein the processor 101 and the memory 102 communicate with each other through the bus 103; the processor 101 is used to call the program instructions in the memory 102 to execute the above-mentioned adaptive robot trajectory planning method based on deep reinforcement learning.
[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, adaptive robot trajectory planning systems (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] In a typical configuration, the adaptive robot trajectory planning system includes one or more processors (CPU), memory and bus. The adaptive robot trajectory planning system may also include input / output interface, network interface and the like.
[0107] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip. The memory is an example of a computer-readable medium.
[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage computer readable storage media or any other non-transmission media that can be used to store information that can be accessed by the adaptive robot trajectory planning system. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or computer-readable storage medium that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such process, method, commodity or computer-readable storage medium. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the presence of other identical elements in the process, method, commodity or computer-readable storage medium that includes the elements.
[0110] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0111] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. An adaptive robot trajectory planning method based on deep reinforcement learning, characterized in that: include: Acquire a real-time trajectory data set of a target robot in a dynamic environment, wherein the real-time trajectory data set includes an environment interaction state sequence and a robot motion posture sequence at multiple continuous time stamps; Performing dynamic trajectory feature extraction processing on the real-time trajectory data set to obtain a dynamic trajectory feature set; Based on the pre-trained deep reinforcement learning strategy network, a multi-dimensional trajectory optimization strategy generation process is performed on the dynamic trajectory feature set to obtain an adaptive trajectory optimization strategy; The robot motion posture sequence is incrementally corrected according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that meets the adaptability of the dynamic environment, and the optimized trajectory sequence is synchronized to the robot motion control system to trigger the trajectory execution operation.
2. The method according to claim 1, characterized in that The dynamic trajectory feature set includes obstacle distribution features, motion constraint features and trajectory continuity features. The real-time trajectory data set is subjected to dynamic trajectory feature extraction processing to obtain a dynamic trajectory feature set, including: Extracting obstacle distribution features from the environment interaction state sequence, wherein the obstacle distribution features include obstacle gathering area boundaries, obstacle movement trend directions, and obstacle safety avoidance distances; Extracting motion constraint features from the robot motion posture sequence, wherein the motion constraint features include a joint motion angle limit range, a terminal path curvature limit threshold, and a speed mutation suppression interval; A trajectory continuity detection is performed on the temporal alignment relationship between the environment interaction state sequence and the robot motion posture sequence to generate a trajectory continuity feature.
3. The method according to claim 2, characterized in that The performing trajectory continuity detection on the time sequence alignment relationship between the environment interaction state sequence and the robot motion posture sequence to generate a trajectory continuity feature includes: Aligning the environment interaction state sequence with the robot motion posture sequence according to timestamps to generate a set of synchronized trajectory data segments; Performing trajectory segmentation processing on each synchronous trajectory data segment to obtain multiple trajectory segments, and performing trajectory conflict detection on each trajectory segment to generate trajectory conflict features; the trajectory conflict features include a conflict time window, a conflict location area, and a conflict type label of a dynamic obstacle invading the trajectory segment; Performing a smooth transition feasibility assessment on the connection points of adjacent trajectory segments according to the trajectory conflict characteristics, and generating a trajectory interruption risk level and a trajectory direction mutation probability; The trajectory continuity feature is generated based on the trajectory interruption risk level and the trajectory direction mutation probability, and the trajectory continuity feature is used to mark a set of trajectory segments that need to be repaired in continuity in the optimized trajectory sequence.
4. The method according to claim 1, characterized in that The pre-trained deep reinforcement learning strategy network performs multi-dimensional trajectory optimization strategy generation processing on the dynamic trajectory feature set to obtain an adaptive trajectory optimization strategy, including: Inputting the dynamic trajectory feature set into the deep reinforcement learning strategy network, and generating a joint optimization decision vector through a multi-level feature fusion module; Generating a trajectory optimization priority sequence according to the trajectory optimization weight distribution relationship in the joint optimization decision vector; Based on the trajectory optimization priority sequence, performing trajectory replanning processing on the trajectory conflict features in the dynamic trajectory feature set to generate a set of candidate trajectory correction solutions; Performing dynamic environment adaptability verification on the candidate trajectory correction scheme set, screening out an optimized trajectory correction scheme that meets preset constraints, and encoding the optimized trajectory correction scheme as the adaptive trajectory optimization strategy; The adaptive trajectory optimization strategy is used to simultaneously indicate the trajectory smoothness optimization direction and the dynamic obstacle avoidance optimization direction.
5. The method according to claim 4, characterized in that The dynamic environment adaptability verification of the candidate trajectory correction scheme set is performed to select an optimized trajectory correction scheme that meets preset constraints, including: Simulating the execution process of the candidate trajectory correction scheme in a dynamic environment and collecting virtual environment interactive response data; Performing trajectory safety analysis on the virtual environment interactive response data to generate trajectory safety evaluation indicators; the trajectory safety evaluation indicators include the achievement rate of the minimum obstacle avoidance distance, the number of times the joint movement angle exceeds the limit, and the compliance rate of the terminal path curvature; If any one of the trajectory safety assessment indicators does not reach the preset safety threshold corresponding to the item, the trajectory parameters of the candidate trajectory correction scheme that does not meet the standard are retroactively adjusted to generate an adjusted candidate trajectory correction scheme; The simulation and adjustment operations are repeated until the trajectory safety evaluation indicators of all candidate trajectory correction schemes meet the corresponding preset safety thresholds, and an optimized trajectory correction scheme that meets the preset constraints is generated.
6. The method according to claim 1, characterized in that The step of performing incremental trajectory correction processing on the robot motion posture sequence according to the adaptive trajectory optimization strategy to generate an optimized trajectory sequence that satisfies dynamic environment adaptability includes: According to the trajectory optimization direction in the adaptive trajectory optimization strategy, local trajectory interpolation processing is performed on the trajectory conflict feature points in the robot motion posture sequence to generate a smooth transition trajectory segment; Performing dynamic obstacle avoidance path offset processing on the smooth transition trajectory segment to generate an obstacle avoidance trajectory offset; Spatially superimposing the smooth transition trajectory segment and the obstacle avoidance trajectory offset to generate an intermediate optimized trajectory segment; The intermediate optimized trajectory segment is subjected to kinematic inverse solution verification. When the kinematic inverse solution verification indicates that the joint motion sequence corresponding to the intermediate optimized trajectory segment is within the limit range of the motion constraint feature, the intermediate optimized trajectory segment is inserted into the original robot motion posture sequence to replace the trajectory segment where the corresponding trajectory conflict feature point is located, so as to generate an incrementally corrected optimized trajectory sequence.
7. The method according to claim 1, characterized in that The training process of the deep reinforcement learning strategy network includes: Acquire a simulation training data set containing multiple dynamic environment scenes, wherein each training sample in the simulation training data set includes an environment interaction state sequence, a robot motion posture sequence, and a corresponding ideal trajectory optimization strategy; Inputting the training samples into the initial strategy network for forward propagation calculation to generate a prediction trajectory optimization strategy; Calculating a strategy difference loss value between the predicted trajectory optimization strategy and the ideal trajectory optimization strategy, and updating network parameters of the initial strategy network through back propagation based on the strategy difference loss value; During the training process, according to the preset stage training sequence, iterative training is performed using the subsets of simulation training data with increasing obstacle movement speed and increasing environmental disturbance intensity; When the policy difference loss value of the initial policy network in the target complexity scenario falls into the preset convergence interval, the training is terminated and a deep reinforcement learning policy network is generated.
8. The method according to claim 7, characterized in that The iterative training is performed in accordance with a preset stage training sequence, using simulation training data subsets with increasing obstacle movement speed and increasing environmental disturbance intensity, including: The simulation training data set is divided into stage data subsets corresponding to multiple training stages according to the obstacle movement speed parameter and the environmental disturbance intensity parameter; wherein each training stage corresponds to a stage data subset, and the obstacle movement speed threshold and the environmental disturbance intensity threshold of each stage data subset are gradually increased as the stage sequence number increases; In order from low to high stage numbers, the data subsets of each stage are input into the initial strategy network in sequence for stage training.
9. The method according to claim 8, characterized in that Each stage of training process includes: Perform multi-batch strategy prediction based on the current stage data subset, and calculate the stage matching degree between the predicted trajectory optimization strategy and the corresponding ideal trajectory optimization strategy; When the stage matching degree reaches the preset convergence standard of the current stage, the training data is switched to the data subset of the next stage, and the network parameters trained in the current stage are used as the initialization parameters of the next stage; The phase switching and parameter inheritance operations are executed cyclically until the training of the highest phase data subset is completed, so that the deep reinforcement learning policy network gradually improves the policy generation stability under the dynamic interference of increasing intensity.
10. An adaptive robot trajectory planning system, characterized in that: It includes a processor and a memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call the program instructions in the memory to execute the adaptive robot trajectory planning method based on deep reinforcement learning as described in any one of claims 1-9.
Citation Information
Patent Citations
Robot motion control method, robot and system
CN114952821A
Mechanical arm obstacle avoidance path planning method and system based on reinforcement learning
CN118386252A
Remote operation space manipulator trajectory planning method based on deep reinforcement learning
CN119115953A
Bionic robot control method and system based on AI learning
CN119910648A
Multi-machine collaborative industrial robot intelligent scheduling system and application method
CN119974019A
Cited By
Humanoid robot welding pose adjusting method and system
CN120326636A
Robot control method and system combining fuzzy control and neural network
CN120347774A
Robot control method and system combining fuzzy control and neural network
CN120347774B
Robot joint module control method and system based on self-learning strategy
CN120461455A
Robot motion generation method and robot control system
CN120697036A