Wheel-legged robot trajectory planning method and system based on diffusion strategy
By constructing a trajectory planning model with multi-objective constraints and energy optimization, the problems of high-dimensional coupling characteristics and weak terrain response in the trajectory planning of wheeled and legged robots are solved, and stable movement and efficient operation of wheeled and legged robots in complex environments are realized.
Patent Information
- Application Number
- CN202610632611.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-25
AI Technical Summary
Existing trajectory planning algorithms cannot effectively match the high-dimensional coupling characteristics of wheeled and legged robots' wheel speed-leg joint angle. The diffusion strategy has not established an adaptation mechanism with the wheeled and legged composite motion space, resulting in conflicts between trajectory and hardware execution, making it difficult to meet the requirements of composite motion coupling, terrain dynamic response, and multi-objective optimization.
By employing the method of 'observation, perception and fusion - constraint and optimization function construction - diffusion strategy model adaptation - multi-time step trajectory generation', and combining multi-objective constraint functions and energy optimization functions, a trajectory planning model is constructed to generate a trajectory planning method and system that conforms to the robot's motion characteristics, thereby achieving accurate matching of multimodal perception fusion and wheel-leg coordinated actions.
It improves the rationality and motion economy of trajectory generation, enhances the robot's terrain adaptability and motion stability, solves the problem of wheel-leg coordination adaptation in unstructured terrain, strengthens the comprehensiveness of environmental perception and dynamic trajectory correction capabilities, and improves the robustness of planning.
Smart Images

Figure CN122632826A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of robot application technology, specifically relating to a trajectory planning method and system for wheeled and legged robots based on a diffusion strategy. Background Technology
[0002] Wheeled-legged robots combine the efficiency of wheeled mobility with the terrain adaptability of legged mobility, making them core equipment for operations in complex environments. The rationality of their trajectory planning directly determines motion stability, task completion rate, and real-time response capability. Existing trajectory planning algorithms are mostly designed for single wheeled or legged robots and cannot match the high-dimensional coupling characteristics of wheeled-legged robots' "wheel speed-leg joint angle". Existing diffusion strategies are mainly aimed at single-modal joint control of robotic arms and have not established an adaptation mechanism with the composite motion space of wheeled-legged robots. This makes it difficult to meet the core operational requirements of wheeled-legged robots, and direct migration can easily lead to conflicts between trajectory and hardware execution, making practical applications impossible.
[0003] Therefore, there is an urgent need for a trajectory planning method and system that achieves deep adaptation between diffusion strategies and wheeled robots through task-oriented function customization and input / output standardization design, while meeting the requirements of complex motion coupling, terrain dynamic response, and multi-objective optimization, so as to promote the engineering implementation of diffusion strategies in this field. Summary of the Invention
[0004] The purpose of this application is to provide a trajectory planning method and system for wheeled and legged robots based on a diffusion strategy. By "observation perception and fusion - constraint and optimization function construction - diffusion strategy model adaptation - multi-time step trajectory generation", it solves the problems of difficult adaptation of composite actions, weak terrain response, insufficient integration of multiple constraints, and difficulty in engineering implementation in the prior art.
[0005] This application discloses a trajectory planning method for wheeled-legged robots based on a diffusion strategy. The method is applied to wheeled-legged robots with wheels and legs, and includes: Based on the hardware configuration and target task requirements of the wheeled robot, the core observation objects and observation dimensions are determined, and the core state variables of the core observation objects are collected as observation information. The observation information is then fused to generate multimodal perception fusion data. The core state variables are environmental feature data and robot body state data that match the observation dimensions. Based on the requirements of the target task, a multi-objective constraint function and an energy optimization function are dynamically constructed, and the constraint weights and optimization index ratios are adjusted according to the task priority, providing constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation; Based on the target task requirements and the wheeled robot hardware configuration, a trajectory planning model is constructed with a diffusion strategy as the core. The input-output mapping relationship of the trajectory planning model is defined, and multimodal perception fusion data and corresponding adapted wheeled-leg cooperative motion data are selected as training samples. Under the constraints of the multi-objective constraint function and the optimization guidance of the energy optimization function, the trajectory planning model is trained until the model outputs cooperative motion commands that fit the target task requirements and robot motion characteristics. The multimodal perception fusion data is input into the adapted trajectory planning model. Under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, the trajectory planning model generates a continuous trajectory action sequence containing multiple time steps, which adapts to the motion requirements of the wheeled robot in terrains of different complexity.
[0006] Furthermore, the hardware configuration includes the number of joints and wheel diameter parameters of the wheeled robot; the target task requirements include inspection path constraints and rescue scenario priorities; the core observation objects include obstacles, terrain, robot body state, and task objectives, and the observation dimensions include obstacle position parameters, terrain geometric features, robot posture and motion parameters, and task objective parameters, including target position, preset path, task priority, and expected motion indicators; the environmental feature data includes terrain type, height difference, roughness, slope gradient, and obstacle position and size; and the robot body state data includes joint angles, wheel speed, and chassis posture.
[0007] Furthermore, the method also includes: combining the target task requirements and the corresponding task environment characteristics to construct a terrain complexity evaluation function to quantify the terrain complexity level, and adaptively adjusting the observation time window and action time window of the trajectory planning model according to the complexity level result: the higher the terrain complexity, the longer the observation time window and the shorter the action time window; the lower the terrain complexity, the shorter the observation time window and the longer the action time window.
[0008] Furthermore, the formula for the terrain complexity evaluation function is as follows: , In the formula, The terrain complexity level, These are the weighting coefficients. The root mean square of the terrain elevation difference. This refers to terrain roughness.
[0009] Furthermore, the trajectory planning model constructed with the diffusion strategy as its core includes: The multimodal perception fusion data is encoded into guiding features by a conditional feature encoder, which serve as the conditional features of the diffusion strategy model to constrain the terrain and task adaptability of subsequent trajectory generation. The initial wheel-leg coordinated motion sequence is processed through a forward noise addition process. Gaussian noise is gradually added to obtain a noisy action sequence. Uncertainty in simulated trajectory movements: With noisy action sequences Time step A noise prediction network is constructed using conditional features as input and a reverse denoising backbone. Obtain the current prediction noise Generate trajectory movements that fit the constraints and optimization objectives; During model training, a composite loss function that integrates noise prediction error, constraints, and energy optimization is used to train the trajectory planning model, and the total loss value of the composite loss function is obtained. ; Based on total loss value By updating the model parameters through backpropagation and iteratively optimizing the noise prediction network, the final output is a wheel-leg coordinated action sequence that meets the objectives of multi-time step, multi-constraint, and energy optimization.
[0010] Furthermore, the composite loss function that integrates noise prediction error, constraints, and energy optimization is as follows: , In the formula, The total loss value for training the trajectory planning model; Predict loss for noise; , This is the weighted penalty coefficient; For energy optimization loss term, The set of coordinated motion instructions for "wheel speed - leg joint angle" predicted by the model; To constrain the penalty for violations, hinge loss quantification is used, as shown in the following formula: , in, The number of constraint terms in a multi-objective constraint function; To predict action commands The corresponding number Actual values of the constraint indicators; For the first The preset threshold of the constraint index is used when an action command violates the constraint function. This generates a penalty value to force the model to conform to the constraints.
[0011] Furthermore, the multimodal perception fusion data is formed by weighted fusion through deep learning and attention mechanisms, including terrain feature vectors, robot body state parameters, and task target parameters; the model output is a wheel speed-leg joint angle coordinated action instruction set adapted to the robot hardware configuration, and the instruction set corresponds to the rotational speed of each wheel and the angle of each leg joint.
[0012] Furthermore, the core constraints integrated by the multi-objective constraint function include: kinematic constraints: wheel speed + joint angle constraints, used to limit the hardware physical boundaries such as wheel speed range and leg joint angle travel; dynamic constraints: driving torque constraints, used to match motor load capacity and chassis load limit; obstacle avoidance constraints: dynamic safety distance constraints, used to ensure a safe distance between the trajectory and obstacles; wheel-leg cooperative adaptive constraints: terrain-related penalty term, used to control chassis attitude deviation and foot force distribution balance; the core indicators of the energy optimization function include the estimated energy consumption of wheel speed motor and joint motor, chassis vibration amplitude, and trajectory completion time. This function quantifies the trajectory optimization objective by weighting the core indicators.
[0013] Furthermore, the method also includes: splitting the multi-time-step trajectory action sequence output by the trajectory generation module into independent action commands for the wheels and legs, and performing multi-objective constraint function constraint verification and energy optimization function optimization verification on the decoupled action commands step by step; when both verifications pass, issuing compliant action commands to the actuators of the wheels and legs of the adapted wheel-legged robot; if they fail, feeding back to the trajectory generation module to regenerate the trajectory.
[0014] This application also discloses a trajectory planning system for a wheeled and legged robot based on a diffusion strategy, used to implement the method described in any of the preceding claims, wherein the system includes a scene perception module, a prediction module, a diffusion strategy adaptation module, and a trajectory generation module; The scene perception module includes a multimodal sensor and a data processing unit. The multimodal sensor is used to collect the core state variables of the core observation object as observation information. The data processing unit is used to clarify the observation object and observation dimension based on the wheeled robot hardware configuration and target task requirements, and to perform weighted fusion of the observation information to form and output multimodal perception fusion data. The prediction module, as the core adaptation unit of the system, has a built-in function construction module. The function construction module is used to dynamically construct multi-objective constraint functions and energy optimization functions by combining the target task requirements and terrain conditions. It can also adjust the constraint weights and optimization index ratios according to task priorities, providing constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation. The diffusion strategy adaptation module includes a mapping rule unit and a model training unit. The mapping rule unit is used to define the input-output mapping relationship of the trajectory planning model according to the hardware configuration of the wheel-legged robot and the target task requirements, convert the multimodal perception fusion data into a standardized input sequence that the model can recognize, and parse the model output into a set of wheel speed-leg joint angle coordinated action instructions. The model training unit is used to use the multimodal perception fusion data and the corresponding adapted wheel-leg coordinated action data as training samples, and complete the model training under the dual guidance of multi-objective constraint functions and energy optimization functions, so that the model outputs coordinated action instructions that conform to the target task requirements and robot motion characteristics. The trajectory generation module is used to receive the adapted multimodal fusion data and, under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, generate a continuous trajectory action sequence containing multiple time steps to adapt to the motion requirements of the wheeled robot in terrains of different complexity.
[0015] This application discloses a trajectory planning method and system for wheeled and legged robots based on a diffusion strategy. It employs a dual-drive design with multi-objective constraints and energy optimization to precisely define the robot's motion boundaries, while simultaneously optimizing energy consumption, vibration, and time consumption, significantly improving the rationality of trajectory generation and the economy of motion. Furthermore, by integrating the temporal generation advantages of the diffusion strategy with terrain-adaptive wheel-leg cooperative constraints, it can dynamically match the requirements of terrains with varying complexity, solving the challenge of wheel-leg cooperative adaptation in unstructured terrain and improving the robot's terrain adaptability and motion stability. The application also employs multimodal perception fusion and time window adaptive adjustment, combined with closed-loop feedback control, to enhance the comprehensiveness of environmental perception and the ability to dynamically correct trajectories, effectively avoiding trajectory failures caused by sudden environmental changes and perception biases, thus significantly enhancing the robustness of the planning. Attached Figure Description
[0016] To make the contents of this application easier to understand, the invention will be further described in detail below with reference to specific implementation examples and accompanying drawings: Figure 1 A flowchart illustrating a trajectory planning method for a wheeled legged robot based on a diffusion strategy, provided in an embodiment of this application; Figure 2 Flowchart of another trajectory planning method for a wheeled legged robot based on a diffusion strategy provided in this application embodiment; Figure 3 This is a diffusion strategy architecture diagram provided in the embodiments of this application; Figure 4 This is a structural block diagram of a wheeled robot trajectory planning system based on a diffusion strategy, provided in an embodiment of this application. Detailed Implementation
[0017] The following example, using a bipedal wheeled robot reaching a target location along a preset trajectory, illustrates the method and system of this application in detail, verifying their effectiveness and engineering applicability. This scenario covers the requirements of "flat advancement, obstacle adaptation, and target approach," adapting to the complex motion characteristics of bipedal wheeled robots. It also verifies the practical effects of multi-objective constraints and energy optimization, and is applicable to the generation and dynamic adjustment of wheel-leg cooperative trajectories for wheeled robots in unstructured terrain (such as steps, gravel roads, and slopes). It can be widely applied to complex scenarios such as outdoor inspection, emergency rescue, field exploration, and special operations, realizing the engineering implementation of diffusion strategies in this field.
[0018] Figure 1 This document presents a flowchart illustrating a trajectory planning method for a wheeled legged robot based on a diffusion strategy, as provided in an embodiment of this application. Figure 1 As shown, this embodiment discloses a trajectory planning method for a wheeled legged robot based on a diffusion strategy. The core process is "observation perception and fusion - constraint and optimization function construction - diffusion strategy model adaptation - multi-time step trajectory generation," specifically including the following steps: Step S1: Based on the hardware configuration and target task requirements of the wheeled robot, determine the core observation objects and observation dimensions, and collect the core state variables of the core observation objects as observation information. The observation information is then fused to generate multimodal perception fusion data. The core state variables are environmental feature data and robot body state data that match the observation dimensions.
[0019] The hardware configuration includes the number of joints and wheel diameter parameters of the wheeled robot; the target task requirements include inspection path constraints and rescue scenario priorities; the core observation objects include obstacles, terrain, and robot body state; and the observation dimensions include obstacle position parameters, terrain geometric features, robot posture and motion parameters, and task target parameters.
[0020] In one implementation, such as Figure 2 As shown, the system first constructs a scene perception module. Based on the wheeled robot's hardware structure and target task requirements, the observation objects are clearly defined as obstacles, terrain, and robot body state. The observation dimensions are defined to include at least terrain geometric features, obstacle position parameters, robot posture and motion parameters, and task target parameters (including target position, preset path, task priority, and expected motion indicators). Core state variables that can fully characterize the robot's motion state, environmental state, and task target are collected. These core state variables include environmental feature data and robot body state data that match the observation dimensions, such as terrain type, height difference, roughness, slope gradient, obstacle position and size, robot body joint angles, and wheel speed. These core state variables are then weighted and fused to form multimodal perception fusion data.
[0021] Furthermore, in this embodiment, the multimodal sensors integrated into the scene perception module, such as cameras, lidar, IMU, and foot force sensors, simultaneously collect environmental feature data and robot body state data that match the observation dimensions. After standardized preprocessing such as noise reduction, normalization, and timestamp alignment, the data is weighted and fused through deep learning and attention mechanisms to form multimodal perception fusion data. This enables comprehensive perception of the environment and the robot body state, and dynamic weighting of multimodal features, providing comprehensive and accurate data support for subsequent trajectory planning.
[0022] Specifically, the collected multi-source data is first preprocessed sequentially, including feature encoding, dimensionality unification, and standardization (noise reduction, normalization, and timestamp alignment). Then, it is weighted and fused using deep learning and attention mechanisms to form the multimodal perception fusion data required for trajectory planning of the wheel-legged robot, which serves as the input observation sequence for the trajectory planning model. The weighted fusion formula for the attention mechanism is as follows: ,in, For multimodal sensing fusion feature vectors, For the number of modes, For the first Attention weights for each modality and satisfying , For the first The feature vectors of each mode after preprocessing.
[0023] In practical applications, the complexity of unstructured terrain (such as slopes and gravel roads) changes dynamically. Existing algorithms lack a dynamic adjustment mechanism for "observation time-terrain adaptation." In complex terrain, insufficient perception leads to trajectory instability, while in flat terrain, planning redundancy reduces motion efficiency, making it difficult to balance real-time response and stability. To address this imbalance between dynamic terrain response and real-time performance, this embodiment further considers adaptive adjustment of the trajectory planning model's time window. Combining the target task requirements and task environment characteristics, a terrain complexity evaluation function quantifies the terrain complexity level. Based on the complexity result, the observation and action time windows of the trajectory planning model are adaptively adjusted: higher terrain complexity results in a longer observation time window and a shorter action time window, ensuring sufficient terrain perception and trajectory stability; lower terrain complexity results in a shorter observation time window and a longer action time window, improving trajectory planning efficiency and motion real-time performance. The terrain complexity evaluation function formula is as follows: , In the formula, The terrain complexity level, These are the weighting coefficients. The root mean square of the terrain elevation difference. This refers to terrain roughness.
[0024] Step S2: Combine the target task requirements to dynamically construct multi-objective constraint functions and energy optimization functions, and adjust the constraint weights and optimization index ratios according to task priorities to provide constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation.
[0025] like Figure 2 As shown, in the method of this embodiment, a prediction module can be constructed by combining target task requirements or constraints, such as inspection path constraints and rescue scenario priorities. A built-in function construction module dynamically constructs multi-objective constraint functions and energy optimization functions to predict the wheeled robot's traversal capability, time consumption, and energy consumption on target terrain. The weights of constraint terms and the ratio of optimization indicators can be flexibly adjusted to adapt to different operational scenarios. The constructed multi-objective constraint functions include kinematic constraints, dynamic constraints, obstacle avoidance constraints, and wheel-leg cooperative adaptive constraints, used to limit the robot's physical hardware boundaries, motion safety boundaries, and task requirements, and also provide optimization guidance for subsequent trajectory generation. The constructed energy optimization function uses energy consumption (estimated energy consumption of wheel speed motors and joint motors), vibration (chassis vibration amplitude), and time consumption (trajectory completion time) as core indicators. Through weight allocation, the trajectory optimization objectives are quantified, guiding the trajectory towards optimization towards "low energy consumption, low vibration, and high efficiency."
[0026] Furthermore, among the four core constraints integrated by the multi-objective constraint function, the kinematic constraint specifically comprises wheel speed + joint angle constraints, used to limit the hardware physical boundaries such as wheel speed range and leg joint angle travel; the dynamic constraint specifically comprises driving torque constraints, used to match motor load capacity and chassis load-bearing limits; the obstacle avoidance constraint specifically comprises dynamic safety distance constraints, used to ensure a safe distance between the trajectory and obstacles; and the wheel-leg cooperative adaptive constraint specifically comprises terrain-related penalty terms, used to control chassis attitude deviation and foot force distribution balance. Attitude stability constraints are ensured collaboratively by the above four types of constraints, guaranteeing stable robot motion attitude.
[0027] Specifically, each constraint function is represented using a higher-order quantization formula: 1. Kinematic constraints: , in, For the first The rotational speed of each wheel; , These are the minimum and maximum threshold values for wheel speed, respectively; For the first The angle of each leg joint; , These are the minimum and maximum mechanical limit thresholds for the corresponding joint rotation angles (determined by the robot hardware parameters).
[0028] 2. Dynamic constraints: , in, For the first Instantaneous output torque of each driving joint (including inertia, Coriolis, gravity terms, and higher-order dynamics model). Here is the joint inertia matrix; Joint angular velocity; Joint angular acceleration; The matrix represents the Coriolis force and the centrifugal force. This is the gravity load term; , This sets the upper and lower limits for torque (to prevent hardware overload and adapt to dynamic characteristics).
[0029] 3. Obstacle avoidance constraints: , in, For the robot The instantaneous distance from each feature point to the obstacle. The equivalent radius of the robot body. For the robot's instantaneous acceleration, The attenuation coefficient is... The first is the baseline safety distance; the second is the dynamic safety distance (the greater the acceleration, the greater the adaptive increase in safety distance to avoid sudden stop collisions and reflect dynamic adaptability).
[0030] Wheel-leg cooperative adaptive constraints: , in, This is the penalty value for the leg joint angle; This is the terrain-related penalty coefficient; The difference in terrain elevation (characterizing terrain complexity); This is the current joint angle; The reference joint angle is adapted to flat terrain.
[0031] Furthermore, the wheel-leg cooperative adaptive constraint can adjust the penalty coefficient by combining terrain conditions and task priorities. Achieving dynamic modulation: The higher the task priority, The larger the value, the stronger the constraint; the lower the task priority. The smaller the value, the more relaxed the constraint. By applying terrain-related penalties and constraint relaxation to the leg joint angles, the model is guided to generate wheel-leg cooperative trajectories adapted to unstructured terrain.
[0032] Furthermore, the energy optimization function uses energy consumption (estimated energy consumption of wheel speed motors and joint motors), vibration (chassis vibration amplitude), and time consumption (trajectory completion time) as core indicators. It quantifies the trajectory optimization target through weight allocation, guiding the trajectory towards optimization in the direction of "low energy consumption, low vibration, and high efficiency".
[0033] Specifically, the energy optimization function formula is as follows: , Among them, standardized indicator items Defined as: (Energy consumption term, related dynamic constraints: torque and angular velocity); (Vibration term, integral of the deviation between joint angular velocity and angular velocity reference value); (Time consumed, total time to complete the trajectory); Where J is the comprehensive energy optimization target value; (Dynamic weighting coefficients); To optimize the target threshold (determined by working backward from the constraint boundary to ensure that the optimization does not exceed the constraints); This represents the total duration of the trajectory.
[0034] While existing technologies can construct basic constraint systems, they mostly employ fixed-weight constraint patterns, failing to achieve dynamic collaborative adaptation of kinematics, dynamics, obstacle avoidance, and posture stability. This leads to problems such as wheel-leg movement interference and posture deviation in unstructured terrain. Furthermore, energy optimization often focuses on a single energy consumption metric, neglecting crucial aspects of wheel-legged robot operations such as vibration suppression (affecting hardware lifespan) and time-consuming adaptation (affecting task timeliness), resulting in a disconnect between optimization goals and actual operational requirements. To address the poor coordination of multiple constraints and the imbalance of optimization goals in existing technologies, this embodiment dynamically adjusts the weights of each constraint term in the multi-objective constraint function based on task priority. For example, energy consumption optimization weights are increased in inspection scenarios, while obstacle avoidance constraint weights are increased in rescue scenarios. The multi-objective energy optimization function uses energy consumption, vibration, and time consumption as core indicators, quantifying trajectory optimization goals through weight allocation to guide trajectory optimization towards "low energy consumption, low vibration, and high efficiency."
[0035] Step S3: Based on the target task requirements and the wheeled robot hardware configuration, a trajectory planning model is constructed with a diffusion strategy as the core. The input-output mapping relationship of the trajectory planning model is defined, and multimodal perception fusion data and corresponding adapted wheeled robot cooperative motion data are selected as training samples. Under the constraints of the multi-objective constraint function and the optimization guidance of the energy optimization function, the trajectory planning model is trained until the model outputs cooperative motion commands that fit the target task requirements and robot motion characteristics.
[0036] In the method of this embodiment, a trajectory planning model is first constructed with a diffusion policy as the core, and the input-output mapping relationship of the trajectory planning model is defined based on the mapping rule unit to realize the connection between the diffusion policy and the wheel-legged robot task. Furthermore, the model is trained using multimodal perception fusion data and corresponding wheel-legged coordinated action data as training samples to achieve accurate mapping from multimodal perception fusion data to wheel-legged coordinated action commands and model training.
[0037] In one specific implementation, a trajectory planning model can be constructed based on the target task requirements and the wheeled robot hardware configuration, with a diffusion strategy as the core: First, the input-output mapping relationship of the trajectory planning model is defined based on the mapping rule unit, that is, the data type, format and dimension of the model's input data, as well as the type, format and dimension of the model's output instructions are clarified, and a one-to-one correspondence is established from multimodal perception fusion data to the wheel speed-leg joint angle coordinated action instruction set, so that the model output can be directly adapted to the robot's actuator, realizing the connection between the diffusion strategy and the wheeled robot task.
[0038] Specifically, the model input consists of multimodal perception fusion data generated by the scene perception module. This data includes terrain feature vectors (such as height difference and roughness), robot body state parameters (such as joint angles, wheel speeds, and chassis posture), and task target parameters (such as path constraints and scene priorities). It has been converted into a standardized input sequence that the model can recognize. The model output is a set of wheel speed-leg joint angle coordinated action instructions adapted to the robot's hardware configuration. The instruction set corresponds to the rotational speed of each wheel and the angle of each leg joint. The output dimension can be flexibly adjusted according to the robot's configuration (such as bipedal, quadrupedal, and hexapedal). There is no need to modify the core algorithm of the diffusion strategy. It can be adapted to different configurations of wheeled and legged robots simply by adjusting the mapping rules, thereby improving compatibility and ensuring accurate matching with the robot's hardware configuration.
[0039] During the model training phase, multimodal perception fusion data and corresponding wheel-leg coordinated motion data are used as training samples to train the model, achieving accurate mapping from multimodal perception fusion data to wheel-leg coordinated motion commands and model training. During model training, multimodal perception fusion data is selected as input samples, and corresponding wheel-leg coordinated motion data is selected as output label samples, forming a complete training sample pair. Under the dual guidance of the multi-objective constraint function and energy optimization function constructed in step S2, the trajectory planning model is trained to learn the mapping rules from multimodal perception fusion data to wheel-leg coordinated motion commands. During training, the constraint function ensures that the model's output motion commands conform to the robot's hardware physical boundaries and motion safety boundaries, while the energy optimization function guides the model to generate motion commands with low energy consumption, low vibration, and high efficiency. Ultimately, this ensures that the coordinated motion commands output by the trained model not only meet the target task requirements but also match the actual motion characteristics of the wheel-legged robot.
[0040] Furthermore, the trajectory planning model is constructed with a diffusion strategy as its core. This diffusion strategy is based on a two-stage architecture of forward diffusion and reverse denoising, which is adapted to the temporal generation requirements of trajectory planning and realizes the generation of wheel-leg coordinated trajectories from noisy action sequences to those that meet the constraints and energy optimization objectives.
[0041] Figure 3 This is a diagram illustrating the diffusion strategy architecture provided in this embodiment. Figure 3 As shown, the steps for constructing a trajectory planning model based on a diffusion strategy are as follows: S31. Conditional Guided Input Stage: Multimodal perception fusion data (such as terrain, attitude, obstacles, etc.) are encoded into guiding features by a conditional feature encoder, which serve as the conditional features of the diffusion strategy model to constrain the terrain and task adaptability of subsequent trajectory generation.
[0042] Specifically, the conditional input of the diffusion strategy model is multimodal perception fusion data generated by the scene perception module, which is encoded into guiding features by a conditional feature encoder. This data covers terrain feature vectors (such as height difference and roughness), robot body state parameters (such as joint angles, wheel speeds, and chassis posture), and task objective parameters (such as path constraints and scene priorities), and has been converted into a standardized input sequence that the model can recognize. The model output is a set of wheel speed-leg joint angle coordinated action instructions adapted to the robot's hardware configuration. The instruction set corresponds to the rotational speed of each wheel and the angle of each leg joint. The output dimension can be flexibly adjusted according to the robot's configuration (such as bipedal, quadrupedal, and hexapedal). There is no need to modify the core algorithm of the diffusion strategy. It can be adapted to different configurations of wheeled and legged robots simply by adjusting the mapping rules, thereby improving compatibility and ensuring accurate matching with the robot's hardware configuration.
[0043] S32, Forward Diffusion Noise Addition Stage: The initial wheel-leg coordinated motion sequence is denoised through a forward diffusion process. (Noise-free baseline motion) Gaussian noise is gradually added to obtain a noisy motion sequence. (Multi-step wheel speed + joint angle), simulating the uncertainty of trajectory motion: , In the formula, for The sequence of joint commands for the coordinated movement of the wheel and leg at any given moment corresponds to the single-moment trajectory movement command; for The sequence of action instructions with noise added at each step; for Real-time noise attenuation coefficient (controls the intensity of noise addition and dynamically adjusts to adapt to terrain complexity). for Gaussian noise that always follows a standard normal distribution (simulating random disturbances in unstructured terrain); It follows a standard normal distribution. It is an identity matrix.
[0044] To simplify calculations, a cumulative product coefficient is introduced. Then the forward diffusion process can be simplified as follows: , In the formula, This is the initial wheel-leg coordinated motion instruction sequence (noise-free, baseline motion adapted to flat terrain). for Step cumulative noise attenuation coefficient, The larger, The smaller, The stronger the noise, the greater the uncertainty of the simulation in complex terrain.
[0045] S33, Reverse Denoising Generation Stage: Using noisy action sequences Time step A noise prediction network constructed using conditional features as input and a reverse denoising backbone (U-Net1D). Obtain the current prediction noise Generate trajectory actions that fit the constraints and optimization objectives: , In the formula, For inverse denoising distribution (by diffusion strategy model parameters) Modeling); The mean of the denoised action command sequence (predicted by the prediction model, conforming to multi-objective constraints and energy optimization objectives); To reduce the variance (which can be fixed or adaptively predicted by the prediction model, control the smoothness of the denoising and ensure the continuity of the trajectory).
[0046] The core formula for noise prediction in the diffusion strategy model is used for inverse denoising: , In the formula, The prediction noise at time t is the model's prediction. The backbone network for the diffusion strategy model (adapted to trajectory temporal characteristics, with input being a noisy action sequence) Time step Multimodal sensing fusion data ); The multimodal perception fusion data generated in step S1 is used to guide the model to generate action commands that are adapted to the current terrain and task.
[0047] S34. Dual-Objective Loss Constraint Training Phase: During model training, a composite loss function that integrates noise prediction error, constraints, and energy optimization is used to train the trajectory planning model. The total loss is: , In the formula, The total loss value for training the trajectory planning model; For noise prediction loss (the deviation between the predicted value and the actual value of the action command); The penalty coefficient is... To constrain and optimize energy losses (including wheel-leg coordination, joint limits, driving force constraints, and obstacle avoidance constraints).
[0048] Furthermore, the weighting formula for balancing fitting accuracy, constraint compliance, and energy optimization objectives can be dynamically adjusted according to task requirements, as follows: , In the formula, For energy optimization loss term, ( The energy optimization function in step S2 is directly associated with the model's predicted "wheel speed-leg joint angle" coordinated motion instruction set. ; To constrain violations and penalties, , Here, is the weighted penalty coefficient, where, To quantify hinge loss, the formula is as follows: , in, The number of constraint terms (including kinematic, dynamic, obstacle avoidance, and wheel-leg cooperative adaptive constraints) in the multi-objective constraint function of step S2. To predict action commands The corresponding number Actual values of the constraint indicators; For the first The preset threshold for the constraint indicator.
[0049] S35. Model Update and Output Stage: Based on Total Loss By updating the model parameters through backpropagation and iteratively optimizing the noise prediction network, the final output is a wheel-leg coordinated action sequence that meets the objectives of multi-time step, multi-constraint, and energy optimization.
[0050] When the action instruction violates the constraint function of step S2 This generates penalty values to force the model to conform to constraints, ensuring that the output instructions not only meet the boundaries of the multi-objective constraint function but also approximate the energy optimization objective. This allows the model to learn and generate trajectory action sequences that comprehensively optimize energy consumption, vibration, and time consumption, thus achieving model training guided by dual functions as constraints and optimization.
[0051] Step S4: Input the multimodal perception fusion data into the adapted trajectory planning model. Under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, the trajectory planning model generates a continuous trajectory action sequence containing multiple time steps at one time, adapting to the motion requirements of the wheeled robot under terrains of different complexity.
[0052] like Figure 2 As shown, in the method of this embodiment, the multimodal perception fusion data processed by the scene perception module can be input into the trained trajectory planning model. Combined with the adaptively adjusted observation time window and action time window, the trajectory planning model generates continuous wheel-leg coordinated action instructions step by step under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, forming a trajectory action sequence containing multiple time steps. This sequence fully describes the wheel speed change and leg joint angle adjustment of the robot within a certain period of time (e.g., 0.18s), adapting to the motion requirements of the wheel-leg robot under terrains of different complexity, and ensuring the continuity and consistency of wheel-leg coordinated motion.
[0053] Furthermore, regarding trajectory planning strategies, this embodiment adapts different trajectory planning models to different scenarios. For example, in conventional task scenarios, a balanced planning path is used, employing a general trajectory scheme that balances obstacle avoidance safety, motion stability, and traffic efficiency. For highly complex environments with dense obstacles, a safety-first planning path is used, generating a safe trajectory with a large detour range by strengthening wheel-leg coordination constraints and extending the observation time window, ensuring motion stability in complex terrain. For energy-first task scenarios, an energy- and efficiency-first planning path is used, taking the energy optimization function as the core guide under the premise of satisfying multi-objective constraints. By adjusting various energy loss terms and weights, the path is planned along low-lying, barrier-free areas, resulting in shorter paths, smoother motion, and ensuring the lowest overall energy consumption and highest traffic efficiency.
[0054] Furthermore, depending on the complexity of the terrain, trajectory generation can be divided into different stages. Each stage achieves dynamic adaptation of wheel speed and leg joint angle by adjusting the wheel-leg coordination constraint weight, observation time window and action time window, ensuring motion stability and achieving optimization goals under different terrains.
[0055] This embodiment adapts to the high-dimensional coupling characteristics of the "wheel-leg" composite motion space. By integrating multi-objective constraint functions, it ensures that the trajectory conforms to the hardware boundaries and motion stability. It constructs a terrain adaptive adjustment mechanism that adjusts the observation time and fixes the observation dimension to balance real-time performance and stability under unstructured terrain. At the same time, it optimizes the overall trajectory performance with energy consumption, vibration, and time consumption as the core objectives, adapting to diverse and complex operation scenarios. It is the first to realize the engineering implementation of the diffusion strategy in unstructured terrain trajectory planning for wheel-legged robots, without modifying the core algorithm of the diffusion strategy, thus reducing development costs and implementation threshold.
[0056] Furthermore, in this embodiment, the multi-time-step trajectory action sequence output by the trajectory generation module can be split into independent action commands for the wheels and legs, and the decoupled action commands can be checked for multi-objective constraint functions and energy optimization functions step by step. When both checks pass, compliant action commands are sent to the actuators of the wheels and legs of the adapted wheel-legged robot. If they fail, the feedback is sent back to the trajectory generation module to regenerate the trajectory.
[0057] In one implementation, a motion decoupling and verification module can be constructed to decompose the multi-time-step trajectory motion sequence output by the trajectory generation module into independent motion commands for the wheels and legs. This module adapts to the actuator interfaces of the wheeled and legged robot, ensuring that the commands can directly drive the actuators. Simultaneously, constraint verification and optimization verification are performed on the decoupled motion commands step-by-step. Constraint verification checks whether indicators such as wheel speed, joint angle, driving torque, and obstacle avoidance distance meet the preset requirements of the multi-objective constraint function, ensuring no constraint violations. Optimization verification checks whether the comprehensive energy optimization target value J at each time step is lower than a preset threshold. The system checks whether the energy consumption, vibration, and time consumption indicators match the weight allocation requirements set for the task priority. Only when both checks pass can the instruction be sent to the execution layer. If it fails, it is fed back to the trajectory generation module to regenerate the trajectory. This dual screening ensures the validity of the instruction.
[0058] Furthermore, this embodiment can also achieve dynamic trajectory optimization by constructing a closed-loop execution module. The closed-loop execution module can consist of an execution mechanism and a feedback acquisition unit. The execution mechanism includes wheel speed motors and leg joint motors, used to receive compliant action commands to drive the robot's movement. The feedback acquisition unit collects and transmits the executed data in real time, updates the multimodal perception fusion observation information, and corrects the robot's position and posture deviations. If there are sudden changes in the environment or an increase in terrain complexity, the time window parameters and wheel-leg cooperative constraint weights will be readjusted, and the trajectory generation module will dynamically correct the trajectory action sequence, thereby constructing a closed-loop control link. Specifically: 1. Command execution: The actuator receives compliant action commands from the action decoupling and verification module, and drives the wheel and leg actuators of the wheel-legged robot to coordinate actions in stages, completing terrain adaptation and task advancement according to the generated trajectory action sequence.
[0059] 2. Real-time feedback: The feedback acquisition unit collects the robot's actual motion data and environmental change data in real time through the multimodal sensors of the scene perception module, compares the deviation between the actual actions and the predicted instructions, verifies whether the actual indicators such as energy consumption, vibration, and time consumption meet the energy optimization target, and monitors the dynamic changes of terrain complexity and environmental characteristics in real time.
[0060] 3. Closed-loop update: The feedback acquisition unit transmits real-time feedback data back to the scene perception module and trajectory generation module to update the multimodal perception fusion observation information and correct the robot's position and posture deviations. If there are sudden changes in the environment or a significant increase in terrain complexity, the time window parameters and wheel-leg cooperative constraint weights will be readjusted, and the trajectory generation module will dynamically correct the trajectory action sequence to enter the next round of closed-loop control of "perception-prediction-modeling-generation-execution-feedback" to ensure that the robot continuously conforms to constraints and optimization goals and stably completes the target task.
[0061] Based on the diffusion-strategy-based trajectory planning method for wheeled and legged robots disclosed in the above embodiments, this embodiment correspondingly discloses a diffusion-strategy-based trajectory planning system for wheeled and legged robots. This system serves as the hardware and software implementation carrier for the above embodiments, with each functional module working collaboratively to ensure efficient operation throughout the entire process. Figure 4 As shown, the wheeled robot trajectory planning system based on diffusion strategy includes: a scene perception module, a prediction module, a diffusion strategy adaptation module, and a trajectory generation module.
[0062] The scene perception module includes a multimodal sensor and a data processing unit. The multimodal sensor is used to collect the core state variables of the core observation object as observation information. The data processing unit is used to clarify the observation object and observation dimension based on the wheeled robot hardware configuration and target task requirements, and to perform weighted fusion of the observation information to form and output multimodal perception fusion data.
[0063] The prediction module, as the core adaptation unit of the system, has a built-in function construction module. The function construction module is used to dynamically construct multi-objective constraint functions and energy optimization functions by combining the target task requirements and terrain conditions. It can also adjust the constraint weights and optimization index ratios according to task priorities, providing constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation.
[0064] The diffusion strategy adaptation module includes a mapping rule unit and a model training unit. The mapping rule unit is used to define the input-output mapping relationship of the trajectory planning model according to the hardware configuration of the wheel-legged robot and the target task requirements, convert the multimodal perception fusion data into a standardized input sequence that the model can recognize, and parse the model output into a set of wheel speed-leg joint angle coordinated action instructions. The model training unit is used to use the multimodal perception fusion data and the corresponding adapted wheel-leg coordinated action data as training samples, and complete the model training under the dual guidance of multi-objective constraint functions and energy optimization functions, so that the model outputs coordinated action instructions that conform to the target task requirements and robot motion characteristics.
[0065] The trajectory generation module is used to receive the adapted multimodal fusion data and, under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, generate a continuous trajectory action sequence containing multiple time steps to adapt to the motion requirements of the wheeled robot in terrains of different complexity.
[0066] Furthermore, the system in this embodiment also includes an action decoupling and verification module, which consists of a decoupling module, a constraint verification module, and an optimization verification module. The decoupling module is used to decompose the multi-time-step trajectory action sequence output by the trajectory generation module into independent action commands for the wheels and legs. The constraint verification module is used to perform constraint verification of the decoupled action commands using multi-objective constraint functions. The optimization verification module is used to perform optimization verification of the decoupled action commands using energy optimization functions. The constraint verification module and the optimization verification module are also used to, after successful verification, issue compliant action commands to the actuators of the wheels and legs of the adapted wheel-legged robot; and if the verification fails, feedback is sent to the trajectory generation module to regenerate the trajectory.
[0067] Furthermore, the system in this embodiment also includes a closed-loop execution module, which consists of an execution mechanism and a feedback acquisition unit. The execution mechanism includes wheel speed motors and leg joint motors, used to receive commands to drive the robot to move. The feedback acquisition unit is used to collect and transmit data after execution in real time, update multimodal perception fusion observation information, and correct robot position and posture deviations. It is also used to readjust time window parameters and wheel-leg cooperative constraint weights when the environment changes suddenly or the terrain complexity increases, and the trajectory generation module dynamically corrects the trajectory action sequence to construct a closed-loop control link.
[0068] The method and system of this application adopt a six-level architecture consisting of "scene perception module - prediction module - diffusion strategy adaptation module - trajectory generation module - action decoupling and verification module - closed-loop execution module". With task-oriented function customization and standardized input and output design as the core, the process is highly compatible with the hardware configuration of wheeled and legged robots. The action decoupling and dual verification links can be directly connected to the actuator, and the parameters can be flexibly adjusted according to the task / hardware without redesigning the core logic. The engineering implementation cost is low and the versatility is strong, realizing the engineering implementation of the diffusion strategy.
Claims
1. A trajectory planning method for a wheeled legged robot based on a diffusion strategy, characterized in that, The method includes: Based on the hardware configuration and target task requirements of the wheeled robot, the core observation objects and observation dimensions are determined, and the core state variables of the core observation objects are collected as observation information. The observation information is then fused to generate multimodal perception fusion data. The core state variables are environmental feature data and robot body state data that match the observation dimensions. Based on the requirements of the target task, a multi-objective constraint function and an energy optimization function are dynamically constructed, and the constraint weights and optimization index ratios are adjusted according to the task priority, providing constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation; Based on the target task requirements and the wheeled robot hardware configuration, a trajectory planning model is constructed with a diffusion strategy as the core. The input-output mapping relationship of the trajectory planning model is defined, and multimodal perception fusion data and corresponding adapted wheeled-leg cooperative motion data are selected as training samples. Under the constraints of the multi-objective constraint function and the optimization guidance of the energy optimization function, the trajectory planning model is trained until the model outputs cooperative motion commands that fit the target task requirements and robot motion characteristics. The multimodal perception fusion data is input into the adapted trajectory planning model. Under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, the trajectory planning model generates a continuous trajectory action sequence containing multiple time steps, which adapts to the motion requirements of the wheeled robot in terrains of different complexity.
2. The trajectory planning method according to claim 1, characterized in that, The hardware configuration includes the number of joints and wheel diameter parameters of the wheeled robot; the target task requirements include inspection path constraints and rescue scenario priorities; the core observation objects include obstacles, terrain, robot body state, and task objectives; the observation dimensions include obstacle position parameters, terrain geometric features, robot posture and motion parameters, and task objective parameters; the task objective parameters include target position, preset path, task priority, and expected motion indicators; the environmental feature data includes terrain type, height difference, roughness, slope gradient, and obstacle position and size; the robot body state data includes joint angles, wheel speed, and chassis posture.
3. The trajectory planning method according to claim 1, characterized in that, The method further includes: combining the target task requirements and the corresponding task environment characteristics to construct a terrain complexity evaluation function to quantify the terrain complexity level, and adaptively adjusting the observation time window and action time window of the trajectory planning model according to the complexity level result: the higher the terrain complexity, the longer the observation time window and the shorter the action time window; the lower the terrain complexity, the shorter the observation time window and the longer the action time window.
4. The trajectory planning method according to claim 4, characterized in that, The formula for the terrain complexity evaluation function is as follows: , In the formula, The terrain complexity level, These are the weighting coefficients. The root mean square of the terrain elevation difference. This refers to terrain roughness.
5. The trajectory planning method according to claim 1, characterized in that, The trajectory planning model constructed with the diffusion strategy as its core includes: The multimodal perception fusion data is encoded into guiding features by a conditional feature encoder, which serve as the conditional features of the diffusion strategy model to constrain the terrain and task adaptability of subsequent trajectory generation. The initial wheel-leg coordinated motion sequence is processed through a forward noise addition process. Gaussian noise is gradually added to obtain a noisy action sequence. Uncertainty in simulated trajectory movements: With noisy action sequences Time step A noise prediction network is constructed using conditional features as input and a reverse denoising backbone. Obtain the current prediction noise Generate trajectory movements that fit the constraints and optimization objectives; During model training, a composite loss function that integrates noise prediction error, constraints, and energy optimization is used to train the trajectory planning model, and the total loss value of the composite loss function is obtained. ; Based on total loss value By updating the model parameters through backpropagation and iteratively optimizing the noise prediction network, the final output is a wheel-leg coordinated action sequence that meets the objectives of multi-time step, multi-constraint, and energy optimization.
6. The trajectory planning method according to claim 5, characterized in that, The composite loss function, which integrates noise prediction error, constraints, and energy optimization, is as follows: , In the formula, The total loss value for training the trajectory planning model; Predict loss for noise; , This is the weighted penalty coefficient; For energy optimization loss term, The set of coordinated motion instructions for "wheel speed - leg joint angle" predicted by the model; To constrain the penalty for violations, hinge loss quantification is used, as shown in the following formula: , in, The number of constraint terms in a multi-objective constraint function; To predict action commands The corresponding number Actual values of the constraint indicators; For the first The preset threshold of the constraint index is used when an action command violates the constraint function. This generates a penalty value to force the model to conform to the constraints.
7. The trajectory planning method according to claim 1, characterized in that, The multimodal perception fusion data is formed by weighted fusion through deep learning and attention mechanisms, including terrain feature vectors, robot body state parameters, and task target parameters; the model output is a set of wheel speed-leg joint angle coordinated action instructions adapted to the robot hardware configuration, and the instruction set corresponds to the rotational speed of each wheel and the angle of each leg joint.
8. The trajectory planning method according to claim 1, characterized in that, The core constraints integrated by the multi-objective constraint function include: kinematic constraints: wheel speed + joint angle constraints, used to limit the hardware physical boundaries such as wheel speed range and leg joint angle travel; dynamic constraints: driving torque constraints, used to match motor load capacity and chassis load limit; obstacle avoidance constraints: dynamic safety distance constraints, used to ensure the safe distance between the trajectory and obstacles; wheel-leg cooperative adaptive constraints: terrain-related penalty term, used to control chassis attitude deviation and foot force distribution balance; The core indicators of the energy optimization function include the estimated energy consumption of the wheel speed motor and the joint motor, the chassis vibration amplitude, and the trajectory completion time. The function quantifies the trajectory optimization target by assigning weights to the core indicators.
9. The trajectory planning method according to claim 1, characterized in that, The method further includes: splitting the multi-time-step trajectory action sequence output by the trajectory generation module into independent action commands for the wheels and legs, and performing multi-objective constraint function constraint verification and energy optimization function optimization verification on the decoupled action commands step by step; when both verifications pass, issuing compliant action commands to the actuators of the wheels and legs of the adapted wheel-legged robot; if they fail, feeding back to the trajectory generation module to regenerate the trajectory.
10. A trajectory planning system for a wheeled legged robot based on a diffusion strategy, characterized in that, The system is used to implement the method according to any one of claims 1 to 9, and includes a scene perception module, a prediction module, a diffusion strategy adaptation module, and a trajectory generation module; The scene perception module includes a multimodal sensor and a data processing unit. The multimodal sensor is used to collect the core state variables of the core observation object as observation information. The data processing unit is used to clarify the observation object and observation dimension based on the wheeled robot hardware configuration and target task requirements, and to perform weighted fusion of the observation information to form and output multimodal perception fusion data. The prediction module, as the core adaptation unit of the system, has a built-in function construction module. The function construction module is used to dynamically construct multi-objective constraint functions and energy optimization functions by combining the target task requirements and terrain conditions. It can also adjust the constraint weights and optimization index ratios according to task priorities, providing constraint boundaries and optimization guidance for trajectory planning model training and trajectory generation. The diffusion strategy adaptation module includes a mapping rule unit and a model training unit. The mapping rule unit is used to define the input-output mapping relationship of the trajectory planning model according to the hardware configuration of the wheel-legged robot and the target task requirements, convert the multimodal perception fusion data into a standardized input sequence that the model can recognize, and parse the model output into a set of wheel speed-leg joint angle coordinated action instructions. The model training unit is used to use the multimodal perception fusion data and the corresponding adapted wheel-leg coordinated action data as training samples, and complete the model training under the dual guidance of multi-objective constraint functions and energy optimization functions, so that the model outputs coordinated action instructions that conform to the target task requirements and robot motion characteristics. The trajectory generation module is used to receive the adapted multimodal fusion data and, under the boundary constraints of the multi-objective constraint function and the objective guidance of the energy optimization function, generate a continuous trajectory action sequence containing multiple time steps to adapt to the motion requirements of the wheeled robot in terrains of different complexity.