Robot text space joint control method, device and storage medium

By employing a joint control method based on robot text space, which comprehensively considers hardware, dynamics, tracking accuracy, and semantic matching degree, a stable robot motion trajectory is generated. This solves the problem of low control stability in existing technologies and enables stable execution of the robot on a real machine.

CN122449965APending Publication Date: 2026-07-24SUZHOU TONGYUAN SOFT CONTROL INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU TONGYUAN SOFT CONTROL INFORMATION TECH CO LTD
Filing Date
2026-06-25
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing robot control methods, while focusing on the tracking accuracy of kinematic trajectories, fail to effectively address the hardware and environmental mechanical influences experienced by robots during actual movement. This results in unstable trajectory execution on the actual robot, leading to problems such as falls or joint over-limits, and low control stability.

Method used

A joint text-space control method for robots is adopted. By acquiring text semantic instructions and spatial control constraint data, a motion trajectory generation model is used for comprehensive verification. Hardware constraints, dynamic constraints, tracking accuracy constraints and semantic matching degree constraints are introduced. Preference sample pairs are constructed and a motion diffusion model is trained to generate executable trajectories. A physical compliance penalty term is introduced to ensure the stability of the trajectory.

Benefits of technology

It improves the stability of robot control, avoids falls and joint over-limit issues, ensures stable execution of trajectories on the actual machine, and takes into account text matching accuracy, spatial position requirements, and operational compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122449965A_ABST
    Figure CN122449965A_ABST
Patent Text Reader

Abstract

The application discloses a robot text space joint control method, equipment and a storage medium, relates to the technical field of humanoid robot multi-modal motion control, and the robot text space joint control method comprises the following steps: obtaining a text semantic instruction and space control constraint data of a robot, and inputting the text semantic instruction and the space control constraint data into a motion trajectory generation model to obtain an executable trajectory, wherein the motion trajectory generation model is obtained by training a motion diffusion model based on a preference sample pair, the preference sample pair is constructed based on scores of the motion trajectory of the robot on hardware, dynamics, tracking accuracy and semantic matching degree constraints, a training loss value comprises a physical compliance penalty term determined based on the hardware constraint and the dynamics constraint, and the robot is controlled based on the executable trajectory. Through the setting of the model training sample pair and the loss value, the semantic text matching degree, the space position requirement and the operation compliance of the robot control are improved, and the stability of the control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal motion control technology for humanoid robots, and in particular to a method, device and storage medium for joint control of robot text space. Background Technology

[0002] Current humanoid robot motion generation technology has achieved two independent functions: text-driven and spatial control. However, in actual application scenarios, robots often need to simultaneously meet text semantic instructions and hard spatial constraints. That is, the controlled robot needs to simultaneously meet the requirements of instructions and spatial position.

[0003] Current robot control methods focus solely on the tracking accuracy of kinematic trajectories. However, during actual robot movement, the robot is also subject to the mechanical influences of its interaction with the environment, as well as the limitations of its own hardware. This means that although the generated trajectory meets the spatial position requirements, it cannot be stably executed on a real machine, easily leading to problems such as falls and joint over-limits. Consequently, the control stability of current robot control methods is relatively low. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, and storage medium for joint control of robot text space, aiming to solve the technical problem of low control stability in current robot control methods.

[0005] To achieve the above objectives, this application proposes a robot text space joint control method, the method comprising: Acquire the robot's textual semantic instructions and spatial control constraint data; The text semantic instructions and the spatial control constraint data are input into the motion trajectory generation model to obtain the executable trajectory of the robot. The motion trajectory generation model is trained on a preset motion diffusion model based on preference sample pairs. The preference sample pairs are constructed based on the scores of the robot's motion trajectory on hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints. The training loss value includes a physical compliance penalty term, which is determined based on the hardware constraints and the dynamic constraints. The robot is controlled based on the executable trajectory.

[0006] In one embodiment, before the step of inputting the text semantic instructions and the spatial control constraint data into the motion trajectory generation model, the method further includes: Obtain text semantic instruction samples and spatial control constraint samples, and extract features from the text semantic instruction samples and the spatial control constraint samples respectively to obtain text semantic features and spatial control features; The text semantic features and spatial control features are fused to obtain fused features, which are then input into the motion diffusion model to obtain multiple candidate motion sequences for the robot. The candidate motion sequences are subjected to hardware constraint verification, and the compliant motion sequences obtained from the verification are input into a preset simulation model to obtain full simulation parameters. The simulation model is constructed based on the physical parameters of the robot. Based on the full simulation parameters and the candidate motion sequences, the preference sample pairs are constructed; Based on the preferred sample pairs, the network weight parameters of the motion diffusion model are adjusted to obtain the motion trajectory generation model.

[0007] In one embodiment, the hardware constraint verification includes kinematic constraints. The step of performing hardware constraint verification on the candidate motion sequence and inputting the verified compliant motion sequence into a preset simulation model to obtain full simulation parameters includes: Based on the candidate motion sequence, determine the kinematic parameters of the robot when executing the candidate motion sequence; Based on the kinematic parameters, invalid motion sequences that do not conform to the kinematic constraints are identified from each of the candidate motion sequences, and the invalid motion sequences are deleted from the candidate motion sequences to obtain compliant motion sequences. The full set of simulation parameters is obtained by inputting the compliant motion sequence into the simulation model.

[0008] In one embodiment, the step of constructing the preference sample pairs based on the full simulation parameters and the candidate motion sequences includes: Determine the number of non-compliance items in the candidate motion sequence where the kinematic parameters do not conform to the kinematic constraints, and calculate the hardware constraint score based on the number of non-compliance items; Based on the zero-moment point duration, zero-moment point offset distance, centroid height fluctuation range, centroid acceleration change rate, and foot-ground slippage in the full simulation parameters, the dynamic constraint score is calculated. Calculate the spatial mean square error between the spatial trajectory in the full simulation parameters and the preset ideal spatial trajectory, and calculate the tracking accuracy constraint score based on the mean square error; Extract the features of the spatial trajectory to obtain spatial trajectory features, calculate the cosine similarity between the spatial trajectory features and the text semantic features, and calculate the semantic matching degree constraint score based on the cosine similarity. The preference sample pairs are constructed based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score.

[0009] In one embodiment, the step of constructing the preference sample pair based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score includes: The target task of the robot is determined based on textual semantic instructions; Based on the target task, determine the reward score weights corresponding to the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score, respectively. The hardware constraint score, the dynamics constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score are multiplied by their respective corresponding reward score weights to obtain the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score. The preference sample pairs are constructed based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score.

[0010] In one embodiment, the step of constructing the preference sample pair based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score includes: The hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score are added together to obtain the comprehensive reward score; Based on the comprehensive reward score, each of the compliant sports sequences is sorted to obtain the compliant sports sequence ranking; The physical compliance score is obtained by adding the hardware reward score and the dynamics reward score. Based on the ranking of the compliant motion sequences, motion sequences with physical compliance scores higher than a preset winning compliance threshold are identified as winning samples. From the invalid motion sequences, motion sequences with physical compliance scores lower than a preset elimination compliance threshold are identified as elimination samples. Based on the winning samples and the eliminated samples, the preference sample pairs are constructed.

[0011] In one embodiment, the step of adjusting the network weight parameters of the motion diffusion model based on the preference sample pairs to obtain the motion trajectory generation model includes: Based on the preferred sample pairs, the winning log probability of the winning sample and the elimination log probability of the losing sample in the motion diffusion model are calculated respectively. Based on the winning log probability, the elimination log probability and the preset preference optimization loss function, the loss value is calculated. Based on the preset ideal zero-moment point duration, ideal zero-moment point offset distance, ideal centroid height fluctuation range, ideal centroid acceleration change rate, and ideal foot-to-ground slip, the ideal dynamic constraint score is calculated, and the difference between the ideal dynamic constraint score and the dynamic constraint score is calculated to obtain the dynamic constraint score deviation. The physical compliance penalty item is obtained by adding the dynamic constraint score deviation and the hardware constraint score. Based on the physical compliance penalty and the loss value, the network weight parameters of the motion diffusion model are adjusted to obtain an iterative motion diffusion model; Determine whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets the preset iteration termination condition. If it does, then the iterative motion diffusion model is used as the motion trajectory generation model.

[0012] In one embodiment, the step of determining whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets a preset iteration termination condition, and if so, using the iterative motion diffusion model as the motion trajectory generation model, includes: The iterative candidate motion sequence is used as the candidate motion sequence. The process of performing hardware constraint verification on the candidate motion sequence is returned. The compliant motion sequence obtained from the verification is input into the preset simulation model to obtain the full set of simulation parameters. The invalid motion sequence, the tracking accuracy constraint score and the semantic matching degree constraint score are obtained. If the number of invalid motion sequences is lower than a preset number threshold, the tracking accuracy constraint score is higher than a preset tracking accuracy constraint score threshold, and the semantic matching degree constraint score is higher than a preset semantic matching degree score threshold, then the iterative candidate motion sequence is determined to meet the iteration termination condition, and the iterative motion diffusion model is used as the motion trajectory generation model.

[0013] Furthermore, to achieve the above objectives, this application also proposes a robot text-space joint control device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the robot text-space joint control method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the robot text space joint control method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: The robot's textual semantic instructions and spatial control constraint data are acquired and input into a motion trajectory generation model to obtain the robot's executable trajectory. The motion trajectory generation model is trained on a preset motion diffusion model based on preference sample pairs. These preference sample pairs are constructed based on the robot's motion trajectory scores across hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching constraints. The training loss value includes a physical compliance penalty term, which is determined based on the hardware and dynamic constraints. The robot is then controlled based on the executable trajectory.

[0016] Compared to current methods that focus solely on tracking accuracy of kinematic trajectories, resulting in trajectories that, while meeting spatial requirements, cannot be stably executed on real machines and are prone to problems such as falls and joint overruns, leading to low control stability, this application simultaneously addresses semantic text matching, spatial position requirements, and operational compliance in robot control, thereby improving control stability. Specifically, this application trains the model based on preference data through preference optimization. Because the preference data is constructed based on scores for hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching constraints, the final motion trajectory generation model generates motion trajectories with high scores across all constraint dimensions. Furthermore, during model training, this application constructs additional penalty terms based on hardware and dynamic constraints, ensuring that the generated trajectories better adhere to physical compliance. Therefore, overall, this application can obtain robot motion trajectories that simultaneously consider text matching, spatial position requirements, and operational compliance, thereby improving the stability of robot control. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the robot text space joint control method of this application. Figure 2 A schematic diagram of the overall process of the technical solution provided in Embodiment 1 of the robot text space joint control method of this application; Figure 3This is a flowchart illustrating Embodiment 2 of the robot text space joint control method of this application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the robot text space joint control method in the embodiments of this application; Figure 5 This is a schematic diagram illustrating the data acquisition consent process involved in the robot text space joint control method in this application embodiment.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or robot text-space joint control device capable of performing the above functions. The following description uses a robot text-space joint control device as an example to illustrate this embodiment and the subsequent embodiments.

[0024] Current humanoid robot motion generation technology has achieved two independent functions: text-driven and spatial control. However, in actual application scenarios, robots often need to simultaneously meet text semantic instructions and hard spatial constraints. That is, the controlled robot needs to simultaneously meet the requirements of instructions and spatial position.

[0025] Current robot control methods focus solely on the tracking accuracy of kinematic trajectories. However, during actual robot movement, the robot is also subject to the mechanical influences of its interaction with the environment, as well as the limitations of its own hardware. This means that although the generated trajectory meets the spatial position requirements, it cannot be stably executed on a real machine, easily leading to problems such as falls and joint over-limits. Consequently, the control stability of current robot control methods is relatively low.

[0026] Furthermore, the current solution is not specifically optimized for text-space joint control scenarios, and cannot resolve the conflict between semantic matching, spatial accuracy, and physical feasibility, resulting in poor generalization and practicality.

[0027] Based on this, embodiments of this application provide a method for joint control of a robot's text space, referring to... Figure 1 , Figure 1This is a flowchart illustrating the first embodiment of the robot text space joint control method of this application.

[0028] In this embodiment, the robot text space joint control method includes steps S10~S30: Step S10: Obtain the robot's textual semantic instructions and spatial control constraint data; It should be noted that text semantic instructions are instruction data that use natural language to describe the robot's intention to move; spatial control constraint data is data that describes the hard geometric constraints such as the position, trajectory, and attitude of the robot's joints or end effectors in space.

[0029] Step S20: Input the text semantic instructions and the spatial control constraint data into the motion trajectory generation model to obtain the executable trajectory of the robot. The motion trajectory generation model is trained on a preset motion diffusion model based on preference sample pairs. The preference sample pairs are constructed based on the scores of the robot's motion trajectory on hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints. The training loss value includes a physical compliance penalty term, which is determined based on the hardware constraints and the dynamic constraints. It should be noted that the motion trajectory generation model is a neural network model that receives input conditions and outputs executable motion trajectories for the robot; the preference sample pair is a sample pair consisting of a winning motion trajectory and a losing motion trajectory, used for preference optimization training; the motion diffusion model is a generative model that generates motion sequences through a progressive denoising method. Hardware constraints are physical hardware boundary conditions such as robot joint limits, torque overload, and link collisions; dynamic constraints are physical conditions related to dynamic equilibrium during robot motion, such as center of mass trajectory, foot contact force, and ZMP (zero moment point) stability; tracking accuracy constraints are the position or posture error requirements between the generated trajectory and the input space control constraints; semantic matching degree constraints are the similarity requirements between the action semantics of the generated trajectory and the semantic instructions of the input text; physical compliance penalty terms are additional terms in the loss function used to penalize violations of hardware or dynamic constraints.

[0030] It is understood that the motion trajectory generation model used in this embodiment introduces a physical compliance penalty term based on hardware constraints and dynamic constraints during training. Furthermore, the preferred sample pairs simultaneously consider the scores of four dimensions: hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints. This ensures that the executable trajectory output by the model not only meets the requirements in terms of semantic and spatial accuracy but also conforms to the physical limits and dynamic stability of the robot, thus avoiding problems such as joint over-limits or falls when the generated trajectory is executed on a real machine.

[0031] Step S30: Control the robot based on the executable trajectory.

[0032] It is understood that this embodiment directly controls the robot based on an executable trajectory. This executable trajectory has undergone comprehensive verification of hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints during generation and training. Furthermore, a physical compliance penalty term is introduced into the loss function, thereby enabling the robot to stably reproduce the trajectory requirements during actual execution and avoiding problems such as falls, joint over-limits, or execution failures caused by the trajectory violating physical limits. The overall process of the text space joint control technical solution in this embodiment can be referred to... Figure 2 .

[0033] In one feasible implementation, the specific implementation prior to the step of inputting the text semantic instructions and the spatial control constraint data into the motion trajectory generation model may also be: Obtain text semantic instruction samples and spatial control constraint samples, and extract features from the text semantic instruction samples and the spatial control constraint samples respectively to obtain text semantic features and spatial control features; The text semantic features and spatial control features are fused to obtain fused features, which are then input into the motion diffusion model to obtain multiple candidate motion sequences for the robot. The candidate motion sequences are subjected to hardware constraint verification, and the compliant motion sequences obtained from the verification are input into a preset simulation model to obtain full simulation parameters. The simulation model is constructed based on the physical parameters of the robot. Based on the full simulation parameters and the candidate motion sequences, the preference sample pairs are constructed; Based on the preferred sample pairs, the network weight parameters of the motion diffusion model are adjusted to obtain the motion trajectory generation model.

[0034] It should be noted that the text semantic instruction sample is the natural language description part of the motion generation input data used for training; the spatial control constraint sample is the spatial geometric constraint part of the motion generation input data used for training; the text semantic feature is the vectorized representation extracted from the text semantic instruction sample, and the spatial control feature is the vectorized representation extracted from the spatial control constraint sample; the fusion feature is the unified vector representation after combining the text semantic feature and the spatial control feature through an attention mechanism.

[0035] Candidate motion sequences are multiple possible motion trajectories generated by the motion diffusion model for sampling the same fusion feature; compliant motion sequences are candidate motion sequences that have been retained after hardware constraint verification and have not triggered hardware violations; simulation models are digital twin models built based on robot object parameters to simulate the physical dynamics of the robot; full simulation parameters are complete physical quantity data including joint angles, torques, foot contact forces, and center of mass trajectories output after the simulation model runs; physical parameters are physical property parameters such as the robot's geometric dimensions, mass distribution, joint limits, motor characteristics, and friction coefficient.

[0036] It should also be noted that the specific method and network structure used in this embodiment for encoding text semantics and spatial hard constraints into a unified representation space are as follows: A three-level architecture of independent feature encoding, unified dimension, and cross-fusion is adopted to achieve a unified representation of the two types of constraints; the network structure uses a combination architecture of BERT (a bidirectional encoder representation based on transformers), 1D-CNN (a one-dimensional convolutional neural network), cross attention, and fully connected layers. Specifically, the text semantic encoding uses the BERT-base model (a 12-layer Transformer encoder with a hidden layer dimension of 768). After word segmentation and embedding of the text instructions, the feature vectors of the classification label positions are extracted as text semantic features (768 dimensions). Spatial hard constraint encoding uses a 1D-CNN network (3 convolutional layers with kernel sizes of 3, 5, and 7, stride of 1, padding=same (filling method to keep the size unchanged), activation function is ReLU (linear rectified unit)) to serialize and encode parameters such as trajectory points, joint angles, and obstacle avoidance areas of spatial constraints, and outputs a spatial feature vector with a dimension of 768, achieving dimensionality unification with text features; Subsequently, an 8-head cross-attention mechanism (attention head dimension 96) is used to achieve bidirectional interactive fusion of text features and spatial features. After fusion, a fully connected layer (output dimension 512) is used for feature compression, and finally a joint control embedding vector with dimension 512 is generated to complete the mapping of the two types of constraints to a unified representation space, ensuring that the fused features retain both the semantic integrity of the text and the accuracy of the spatial constraints.

[0037] Furthermore, in this embodiment, a 1:1 digital twin multibody dynamics model of the target humanoid robot is constructed based on the Modelica language in the MWORKS (a scientific computing and system modeling and simulation platform) platform. The model covers all physical characteristics of rigid body dynamics, foot-ground contact, servo drive, and hardware constraints.

[0038] It is understood that this implementation method sequentially performs hardware constraint verification and simulation model verification during training, respectively filtering out candidate sequences that violate hardware boundaries and obtaining all dynamic parameters for constraint scoring. At the same time, the text semantic features and spatial control features are fused and input into the diffusion model to ensure the joint encoding of multimodal conditions. This makes the final trained motion trajectory generation model have hardware compliance and dynamic feasibility when generating trajectories, effectively avoiding the trial and error cost of generating invalid trajectories in the subsequent inference stage, and improving the model training efficiency and generation quality.

[0039] In one feasible implementation, the hardware constraint verification includes kinematic constraints. A further implementation of performing hardware constraint verification on the candidate motion sequence and inputting the verified compliant motion sequence into a preset simulation model to obtain full simulation parameters may be: Based on the candidate motion sequence, determine the kinematic parameters of the robot when executing the candidate motion sequence; Based on the kinematic parameters, invalid motion sequences that do not conform to the kinematic constraints are identified from each of the candidate motion sequences, and the invalid motion sequences are deleted from the candidate motion sequences to obtain compliant motion sequences. The full set of simulation parameters is obtained by inputting the compliant motion sequence into the simulation model.

[0040] It should be noted that kinematic constraints are boundary conditions related to geometric motion, such as robot joint angle limits and link interference distances, but not involving forces and mass; kinematic parameters are numerical values ​​describing the robot's joint angles, angular velocities, end-effector positions, link spacing, and other motion geometric quantities; invalid motion sequences are candidate motion sequences that are determined to be unexecutable after triggering any kinematic constraint.

[0041] It is understood that, before inputting candidate motion sequences into the simulation model, this implementation method first filters and deletes invalid motion sequences that do not conform to kinematic constraints based on kinematic parameters, and only sends compliant motion sequences into the simulation model. This avoids the simulation model performing unnecessary dynamic calculations on invalid sequences that obviously violate joint limits or cause collisions, effectively reducing the consumption of simulation resources and improving the execution efficiency of the overall verification process.

[0042] In one feasible implementation, the specific implementation of constructing the preference sample pairs based on the full simulation parameters and the candidate motion sequences can also be: Determine the number of non-compliance items in the candidate motion sequence where the kinematic parameters do not conform to the kinematic constraints, and calculate the hardware constraint score based on the number of non-compliance items; Based on the zero-moment point duration, zero-moment point offset distance, centroid height fluctuation range, centroid acceleration change rate, and foot-ground slippage in the full simulation parameters, the dynamic constraint score is calculated. Calculate the spatial mean square error between the spatial trajectory in the full simulation parameters and the preset ideal spatial trajectory, and calculate the tracking accuracy constraint score based on the mean square error; Extract the features of the spatial trajectory to obtain spatial trajectory features, calculate the cosine similarity between the spatial trajectory features and the text semantic features, and calculate the semantic matching degree constraint score based on the cosine similarity. The preference sample pairs are constructed based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score.

[0043] It should be noted that the number of non-compliance items refers to the specific number of items in the candidate motion sequence that violate kinematic constraints. The hardware constraint score is a rating value calculated based on the number of non-compliance items, used to quantify the candidate sequence's performance in terms of hardware safety. The zero-moment duration is the length of time during the simulation when the robot's zero-moment point falls inside the foot support polygon. The zero-moment offset distance is the maximum or average spatial offset between the zero-moment point and the center of the foot support polygon during the simulation.

[0044] The range of centroid height fluctuation is the maximum variation of the robot's centroid in the vertical direction during the simulation. The rate of change of centroid acceleration is the rate at which the robot's centroid acceleration changes with time during the simulation.

[0045] Foot-to-ground slip is the relative sliding displacement between the robot's foot and the ground contact point during the simulation.

[0046] The dynamic constraint score is a score calculated based on the zero-moment point duration, zero-moment point offset distance, centroid height fluctuation range, centroid acceleration change rate, and foot-to-ground slip in the full simulation parameters. It is used to quantify the performance of candidate sequences in terms of dynamic stability.

[0047] Spatial mean square error (SMI) is the square root of the mean squared errors at each point between the spatial trajectory of the candidate motion sequence and the preset ideal spatial trajectory. Tracking accuracy constraint score is a score calculated based on SMI and used to quantify the candidate sequence's performance in spatial trajectory tracking accuracy. Spatial trajectory features are vectorized representations obtained after feature extraction from the spatial trajectory of the candidate motion sequence. Cosine similarity is the cosine of the angle between spatial trajectory features and text semantic features, used to measure the similarity between the two feature vectors. Semantic matching constraint score is a score calculated based on cosine similarity and used to quantify the candidate sequence's performance in text semantic matching.

[0048] It should also be noted that the specific scoring method for each dimension in this embodiment is as follows: 1. Spatial Tracking Accuracy Score: Calculated based on the mean square error (MSE) between the generated trajectory and the target spatial constraint trajectory. The smaller the error, the closer the score is to 100.

[0049] 2. Text semantic matching score: Calculated based on cosine similarity, the similarity value (0~1) is mapped to a percentage system (0~100 points).

[0050] 3. Second-level dynamic constraint score: The score adopts a mechanism of full score minus deviation penalty. For example, if the full score is set to 100 points, 10 points are deducted if the ZMP deviates from the center by 1cm, and 5 points are deducted if the foot slips by 1mm. The more serious the deviation from the ideal physical state, the lower the score.

[0051] The core physical quantities to be verified include ZMP support stability (the percentage of time the ZMP point falls within the foot support polygon, and the offset distance between the ZMP point and the center of the support polygon, with an offset distance ≤2cm); centroid trajectory rationality (centroid height fluctuation range, centroid acceleration change rate, such as centroid height fluctuation ≤5cm, acceleration change rate ≤10m / s²); and foot slippage (the relative slippage between the foot and the ground, with a slippage amount ≤2mm). The above physical quantities are output in real time through Modelica simulation and compared with preset thresholds to complete the dynamic compliance verification and quantitative scoring.

[0052] 4. Hard constraint compliance score: Calculated based on the degree of violation. Specifically, if there is no collision or exceeding the limit, 100 points are awarded; if a collision or exceeding the limit occurs, points will be significantly deducted based on the number of joints exceeding the limit and the extent of the exceedance, or even recorded as 0 points.

[0053] The core physical quantities to be verified include: joint angles (the actual angle of each joint is compared with the preset limit range, such as the hip joint limit of ±90° and the elbow joint limit of 0°-150°); joint torques (the output torque of each joint is compared with the peak torque of the motor, such as 80% of the peak torque of the servo motor as the overload threshold); and link collisions (the collision detection module of the Modelica model is used to verify the distance between the links of the robot and between the links and the environment. A distance ≤3mm is considered a collision).

[0054] It is understood that this implementation method simultaneously calculates hardware constraint scores, dynamic constraint scores, tracking accuracy constraint scores, and semantic matching degree constraint scores, and uses these four dimensions as the basis for constructing preferred sample pairs. This ensures that the winning and losing samples in the preferred sample pairs have distinguishable differences in physical feasibility (hardware and dynamics), task accuracy (spatial tracking), and semantic understanding (text matching). As a result, the trained motion trajectory generation model can optimize four objectives simultaneously, avoiding the problem of performance degradation in other dimensions caused by single-dimensional optimization in traditional schemes.

[0055] In one feasible implementation, the specific implementation of constructing the preference sample pair based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score can also be: The target task of the robot is determined based on textual semantic instructions; Based on the target task, determine the reward score weights corresponding to the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score, respectively. The hardware constraint score, the dynamics constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score are multiplied by their respective corresponding reward score weights to obtain the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score. The preference sample pairs are constructed based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score.

[0056] It should be noted that the target task is the specific operation that the robot needs to perform, such as grasping, walking, or raising its hand, which is parsed from the text semantic instructions. The reward score weights are coefficients assigned to the hardware constraint score, dynamics constraint score, tracking accuracy constraint score, and semantic matching degree constraint score, and are used to adjust the importance of different constraint dimensions in the final score.

[0057] The hardware reward score is a weighted score obtained by multiplying the hardware constraint score by its corresponding reward score weight. The dynamics reward score is a weighted score obtained by multiplying the dynamics constraint score by its corresponding reward score weight. The tracking accuracy reward score is a weighted score obtained by multiplying the tracking accuracy constraint score by its corresponding reward score weight. The semantic matching score is a weighted score obtained by multiplying the semantic matching constraint score by its corresponding reward score weight.

[0058] It should also be noted that the specific method for adjusting the weights based on the target task in this embodiment is as follows: First, the basic weight allocation is defined as follows: spatial tracking accuracy 30%, text semantic matching degree 25%, secondary dynamic constraints 25%, and primary hard constraint compliance 20% (the sum of the weights is 100%, which meets the requirement that the sum of the weights of spatial tracking accuracy and text semantic matching degree is not less than 50%). In fine operation scenarios (such as grasping and placing): increase the weight of spatial tracking accuracy to 35% and decrease the weight of secondary dynamic constraints to 20%; in dynamic motion scenarios (such as walking and turning): increase the weight of secondary dynamic constraints to 30% and decrease the weight of spatial tracking accuracy to 25%; in complex semantic scenarios (such as multiple command combinations): increase the weight of text semantic matching degree to 30% and decrease the weight of primary hard constraint compliance to 15%; all weights can be flexibly adjusted through configuration files to ensure that multi-target optimization adapts to different application scenarios.

[0059] It is understood that this implementation determines the target task based on the text semantic instructions, and then dynamically allocates the reward score weights of each constraint dimension according to the target task, so that the judgment criteria of the preference sample pairs in different scenarios are adjusted accordingly. Specifically, the tracking accuracy weight is increased during fine operation and the dynamic constraint weight is increased during dynamic motion. The preference sample pairs constructed can guide the model to optimize for specific task characteristics, thereby improving the adaptability of the motion trajectory generation model in multiple scenarios.

[0060] In one feasible implementation, the specific implementation of constructing the preference sample pair based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score can also be: The hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score are added together to obtain the comprehensive reward score; Based on the comprehensive reward score, each of the compliant sports sequences is sorted to obtain the compliant sports sequence ranking; The physical compliance score is obtained by adding the hardware reward score and the dynamics reward score. Based on the ranking of the compliant motion sequences, motion sequences with physical compliance scores higher than a preset winning compliance threshold are identified as winning samples. From the invalid motion sequences, motion sequences with physical compliance scores lower than a preset elimination compliance threshold are identified as elimination samples. Based on the winning samples and the eliminated samples, the preference sample pairs are constructed.

[0061] It should be noted that the comprehensive reward score is the total score obtained by adding the hardware reward score, dynamics reward score, tracking accuracy reward score, and semantic matching score. The compliant motion sequence ranking is the order in which compliant motion sequences are arranged from highest to lowest according to the comprehensive reward score. The physical compliance score is the score obtained by adding the hardware reward score and the dynamics reward score, and is specifically used to measure the comprehensive performance of candidate sequences in terms of hardware security and dynamic stability.

[0062] The winning compliance threshold is a pre-defined score limit used to determine the minimum physical compliance requirements for a compliant motion sequence to be considered a winning sample. Winning samples are motion sequences selected from compliant motion sequences whose physical compliance scores are higher than the winning compliance threshold. The losing compliance threshold is a pre-defined score limit used to determine the maximum physical compliance requirements for an invalid motion sequence to be considered a losing sample. Losing samples are motion sequences selected from invalid motion sequences whose physical compliance scores are lower than the losing compliance threshold.

[0063] Understandably, this implementation sorts compliant motion sequences based on a comprehensive reward score that includes scores from all dimensions, and also specifically calculates a physical compliance score that reflects only hardware and dynamic performance. It sets winning compliance thresholds and eliminating compliance thresholds to filter winning and eliminating samples, thereby ensuring that the positive and negative samples in the preferred sample pairs not only differ in overall task performance, but also have a clear and significant distinction in the key dimension of physical feasibility. As a result, the trained motion trajectory generation model can more reliably prioritize physically safe and stable trajectories, effectively reducing the risk of falls or damage during real device deployment.

[0064] In one embodiment, the step of constructing the preference sample pairs based on the full simulation parameters and the candidate motion sequences further includes: The difference between the physical compliance score of each winning sample and the physical compliance score of each losing sample is calculated to obtain the physical compliance difference matrix. Based on the physical compliance difference matrix, all winning and losing samples are combined and divided into high difference sample pairs, medium difference sample pairs, and low difference sample pairs according to the difference value. A first number of sample pairs are sampled from the set of high-discrepancy sample pairs, a second number of sample pairs are sampled from the set of medium-discrepancy sample pairs, and a third number of sample pairs are sampled from the set of low-discrepancy sample pairs. The three sets of sampling results are combined to form the final preference sample pairs, where the first number is greater than the second number, and the second number is greater than the third number.

[0065] It should be noted that the physical compliance difference matrix is ​​a two-dimensional array with winning samples as rows and eliminated samples as columns. Each element is a value of the difference in physical compliance scores between the corresponding winning and eliminated samples. The high difference sample pair set is the set of all winning and eliminated sample combinations whose physical compliance difference is higher than a preset high threshold. The medium difference sample pair set is the set of all winning and eliminated sample combinations whose physical compliance difference is between a preset low and high threshold. The low difference sample pair set is the set of all winning and eliminated sample combinations whose physical compliance difference is lower than a preset low threshold.

[0066] The first quantity is the number of sample pairs sampled from the set of highly dissimilar sample pairs, and its value is greater than the second and third quantities. The second quantity is the number of sample pairs sampled from the set of moderately dissimilar sample pairs, and its value is between the first and third quantities. The third quantity is the number of sample pairs sampled from the set of lowly dissimilar sample pairs, and its value is less than the first and third quantities.

[0067] It is understood that in this embodiment, a difference matrix is ​​constructed based on the difference in physical compliance scores and sampling is performed in layers according to the difference magnitude. Sample pairs with high difference are given higher sampling weights, which significantly increases the proportion of sample pairs with distinct physical compliance in the training dataset, while reducing the proportion of ambiguous sample pairs with similar physical compliance. This enhances the ability of the finally trained motion trajectory generation model to discriminate physical constraint boundaries.

[0068] In one embodiment, the step of dividing all winning and losing samples into sets of high-difference, medium-difference, and low-difference samples based on the physical compliance difference matrix further includes: For each substandard sample, extract the peak foot slip, the proportion of maximum joint torque exceeding the limit, and the minimum stability margin at the zero moment point from the corresponding full simulation parameters, and calculate the physical risk coefficient of the substandard sample. Based on the distribution of physical risk coefficients of all inferior samples, the high difference threshold and low difference threshold are dynamically calculated. The high difference threshold is equal to the mean of the difference between the physical compliance of all superior and inferior samples plus the product of the standard deviation of the physical risk coefficient and the first coefficient. The low difference threshold is equal to the mean minus the product of the standard deviation of the physical risk coefficient and the second coefficient. Sample pairs with physical compliance differences greater than the high difference threshold are assigned to the high difference sample pair set; sample pairs with differences between the low and high difference thresholds are assigned to the medium difference sample pair set; and sample pairs with differences less than the low difference threshold are assigned to the low difference sample pair set.

[0069] It should be noted that the peak foot slippage is the maximum relative sliding displacement between the foot and the ground during the simulation of the eliminated sample. The maximum joint torque over-limit ratio is the maximum ratio of the magnitude by which the joint output torque exceeds the peak motor torque to the peak torque in the eliminated sample. The minimum stability margin at the zero-moment point is the shortest distance from the zero-moment point to the boundary of the foot support polygon during the simulation of the eliminated sample. The physical risk coefficient is a comprehensive indicator characterizing the degree of danger of the eliminated sample, calculated by weighting the peak foot slippage, the maximum joint torque over-limit ratio, and the minimum stability margin at the zero-moment point.

[0070] The high difference threshold is a critical value dynamically calculated based on the mean of the physical compliance differences and the standard deviation of the physical risk coefficient. A difference higher than this value is considered high difference. The low difference threshold is a critical value obtained by subtracting the standard deviation of the physical risk coefficient from the mean of the physical compliance differences and multiplying it by a second coefficient. A difference lower than this value is considered low difference. The first coefficient is a preset positive parameter used to adjust the deviation of the high difference threshold relative to the mean. The second coefficient is a preset positive parameter used to adjust the deviation of the low difference threshold relative to the mean.

[0071] It is understandable that this embodiment uses the physical risk coefficient to dynamically adjust the difference threshold, so that inferior samples with high physical risk can be classified into the high difference set even if the difference between their physical compliance scores and those of the winning samples is not particularly large. As a result, the model will be forced to prioritize learning these high-risk edge cases during training, which significantly improves the model's ability to identify and avoid high-risk trajectories, and there is no need for manual labeling of high-risk samples.

[0072] In summary, this embodiment acquires the robot's textual semantic instructions and spatial control constraint data, inputs these data into a motion trajectory generation model, and obtains the robot's executable trajectory. The motion trajectory generation model is trained on a preset motion diffusion model based on preference sample pairs. These preference sample pairs are constructed based on the robot's motion trajectory scores across hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching constraints. The training loss value includes a physical compliance penalty term, which is determined based on the hardware and dynamic constraints. Based on the executable trajectory, the robot is controlled.

[0073] Compared to current methods that only focus on the tracking accuracy of kinematic trajectories, resulting in trajectories that meet spatial requirements but cannot be stably executed on real machines, easily leading to problems such as falls and joint over-limits, thus causing low control stability, this embodiment simultaneously considers the semantic text matching degree, spatial position requirements, and operational compliance of robot control, thereby improving control stability. Specifically, this embodiment trains the model based on preference data through preference optimization. Since the preference data is constructed based on scores of hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints, the final motion trajectory generation model generates motion trajectories with high scores in each constraint dimension. Furthermore, during model training, this embodiment also constructs additional penalty terms based on hardware and dynamic constraints, making the generated trajectory better reflect physical compliance. Therefore, overall, this embodiment can obtain robot motion trajectories that simultaneously consider text matching degree, spatial position requirements, and operational compliance, thereby improving the stability of robot control.

[0074] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 The robot text space joint control method further includes steps S100-S500: Based on the preference sample pairs, the steps adjust the network weight parameters of the motion diffusion model to obtain the motion trajectory generation model. Step S100: Based on the preferred sample pairs, calculate the winning log probability of the winning sample and the elimination log probability of the losing sample in the motion diffusion model respectively. Based on the winning log probability, the elimination log probability and the preset preference optimization loss function, calculate the loss value. It should be noted that the winning log probability is a value obtained by logarithmically transforming the output probability distribution when the motion diffusion model generates winning samples. It is used to quantify the model's tendency to generate winning samples.

[0075] The logarithmic probability of elimination is the value obtained by logarithmically transforming the output probability distribution when the motion diffusion model generates elimination samples. It is used to quantify the model's tendency to generate elimination samples. The preference optimization loss function is a mathematical expression designed based on preferred sample pairs, calculating the loss value by comparing the difference between the logarithmic probabilities of winning and elimination. The loss value is a scalar result calculated by the preference optimization loss function, used to measure the gap between the current motion diffusion model's performance on preferred sample pairs and the expected performance.

[0076] It should also be noted that the comprehensive loss function in this embodiment is: L_DPO = L_base + λ×L_phys Where L_base is the standard DPO (Preference Optimization) loss, L_phys is the physical compliance penalty term (calculated by weighting the number of hard constraint violations and the dynamic constraint score deviation), and λ is the penalty coefficient (ranging from 0.3 to 0.8, preferably 0.5). The worse the physical compliance, the larger L_phys is, and the higher the loss value, forcing the model to learn physically compliant motion sequences.

[0077] Understandably, this embodiment calculates the log probabilities of the model generating winning samples and losing samples respectively, and calculates the loss value based on the difference between the two through a preference optimization loss function. This allows the loss value to directly reflect whether the model tends to generate better samples (i.e., a higher log probability of winning and a lower log probability of losing). As a result, subsequent parameter adjustments can specifically widen the probability gap between the model's generation of winning and losing samples, guiding the model to learn a motion generation strategy that conforms to its preferences.

[0078] Step S200: Based on the preset ideal zero-moment point duration, ideal zero-moment point offset distance, ideal centroid height fluctuation range, ideal centroid acceleration change rate, and ideal foot-to-ground slip, calculate the ideal dynamic constraint score, and calculate the difference between the ideal dynamic constraint score and the dynamic constraint score to obtain the dynamic constraint score deviation; It should be noted that the ideal zero-moment point duration is a preset value, representing the optimal duration for the robot's zero-moment point to fall within the foot support polygon under ideal dynamic conditions. The ideal zero-moment point offset distance is a preset value, representing the optimal offset distance between the robot's zero-moment point and the center of the support polygon under ideal dynamic conditions. The ideal center of mass height fluctuation range is a preset value, representing the optimal range for the vertical fluctuation of the robot's center of mass under ideal dynamic conditions. The ideal center of mass acceleration change rate is a preset value, representing the optimal rate of change of the robot's center of mass acceleration under ideal dynamic conditions.

[0079] The ideal foot-to-ground slip is a preset value representing the optimal allowable slip between the robot's foot and the ground under ideal dynamic conditions. The ideal dynamic constraint score is a reference score calculated based on preset parameters such as the ideal zero-moment point duration, ideal zero-moment point offset distance, ideal center-of-mass height fluctuation range, ideal center-of-mass acceleration rate of change, and ideal foot-to-ground slip. The dynamic constraint score deviation is the difference between the ideal dynamic constraint score and the actual dynamic constraint score, used to quantify the gap between the actual motion sequence and the ideal dynamic state.

[0080] It is understood that this embodiment uses preset ideal dynamic parameters and calculates the deviation between the ideal dynamic constraint score and the actual score, so that the deviation of the dynamic constraint score can intuitively reflect the degree of deficiency of the candidate motion sequence in ZMP stability, center of mass control, foot slip, etc., thereby providing a quantifiable basis for the calculation of subsequent physical compliance penalty items, and thus accurately penalizing the generated results with poor dynamic performance.

[0081] Step S300: Add the dynamic constraint score deviation and the hardware constraint score to obtain the physical compliance penalty item; It is understood that in this embodiment, the dynamic constraint score deviation, which reflects the degree of deviation from the dynamic ideal, is added to the hardware constraint score, which reflects the hardware safety performance, to obtain a physical compliance penalty term. This penalty term covers both the punishment for dynamic instability and the punishment for hardware constraint violation. Therefore, in subsequent training, the model must reduce the dynamic deviation and avoid hardware violations to minimize the penalty term, thereby achieving overall constraint on the physical integrity of the robot's motion.

[0082] Step S400: Based on the physical compliance penalty term and the loss value, adjust the network weight parameters of the motion diffusion model to obtain an iterative motion diffusion model; It should be noted that the iterative motion diffusion model is the motion diffusion model obtained after one adjustment of the network weight parameters and is used for the next round of optimization iteration.

[0083] It is understandable that this embodiment adjusts the network weight parameters based on both the loss value and the physical compliance penalty term. This means that during the optimization process, the model not only needs to learn to distinguish between winning and losing samples, but also must actively reduce its tendency to violate physical constraints. As a result, the iterative motion diffusion model obtained maintains semantic and spatial accuracy while possessing stronger physical feasibility.

[0084] Step S500: Determine whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets the preset iteration termination condition. If it does, then use the iterative motion diffusion model as the motion trajectory generation model.

[0085] It should be noted that the iterative candidate motion sequence is a candidate motion sequence generated by the iterative motion diffusion model under the current network weight parameters. The iteration termination condition is a pre-defined set of conditions used to determine whether the model optimization is complete, such as spatial trajectory tracking error being lower than a threshold, text semantic similarity being higher than a threshold, and physical constraint verification passing rate reaching 100%.

[0086] In one feasible implementation, the step of determining whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets a preset iteration termination condition, and if so, using the iterative motion diffusion model as a specific implementation of the motion trajectory generation model, can also be: The iterative candidate motion sequence is used as the candidate motion sequence. The process of performing hardware constraint verification on the candidate motion sequence is returned. The compliant motion sequence obtained from the verification is input into the preset simulation model to obtain the full set of simulation parameters. The invalid motion sequence, the tracking accuracy constraint score and the semantic matching degree constraint score are obtained. If the number of invalid motion sequences is lower than a preset number threshold, the tracking accuracy constraint score is higher than a preset tracking accuracy constraint score threshold, and the semantic matching degree constraint score is higher than a preset semantic matching degree score threshold, then the iterative candidate motion sequence is determined to meet the iteration termination condition, and the iterative motion diffusion model is used as the motion trajectory generation model.

[0087] It should be noted that the quantity threshold is a preset critical value used to determine whether the number of invalid motion sequences is acceptable. The tracking accuracy constraint score threshold is a preset minimum score used to determine whether the tracking accuracy constraint score reaches a qualified level. The semantic matching score threshold is a preset minimum score used to determine whether the semantic matching constraint score reaches a qualified level.

[0088] It is understandable that this implementation method uses three dimensions—the number of invalid motion sequences, the tracking accuracy constraint score, and the semantic matching degree constraint score—as the criteria for terminating the iteration. It also requires that the number of invalid sequences be below a threshold, the tracking accuracy score be above a threshold, and the semantic matching degree score be above a threshold. This ensures that the final output motion trajectory generation model meets the preset qualification standards in the three core dimensions of physical feasibility, spatial accuracy, and semantic understanding. This avoids the erroneous output of a suboptimal model that meets the standard for only one indicator but not the other.

[0089] In summary, this embodiment, based on the standard preference optimization loss value, uses a physical compliance penalty term composed of dynamic constraint score deviation and hardware constraint score. This subjectes the model to dual constraints of preference difference and physical feasibility during parameter updates, thereby improving the semantic matching and spatial accuracy of the generated motion trajectory while also specifically enhancing physical compliance. Furthermore, the iteration termination condition simultaneously assesses three dimensions: the number of invalid motion sequences, tracking accuracy score, and semantic matching score. This avoids suboptimal model output caused by terminating upon achieving a single metric. As a result, the final motion trajectory generation model meets the preset qualification standards in terms of physical feasibility, spatial tracking accuracy, and semantic understanding ability, and can be directly applied to the stable control of real robots.

[0090] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the robot text space joint control method of this application. Any simple transformations based on this technical concept are all within the protection scope of this application.

[0091] This application provides a robot text-space joint control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the robot text-space joint control method in the above embodiment 1.

[0092] The following is for reference. Figure 4 The diagram illustrates a structural schematic suitable for implementing the robot text-space joint control device in the embodiments of this application. The robot text-space joint control device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, tablets, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The robot text space joint control device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0093] like Figure 4As shown, the robot text-space joint control device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the robot text-space joint control device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the robot text-space joint control device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows robot text-space joint control devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0094] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0095] The robot text-space joint control device provided in this application, employing the robot text-space joint control method in the above embodiments, can solve the technical problem of low control stability in current robot control methods. Compared with the prior art, the beneficial effects of the robot text-space joint control device provided in this application are the same as those of the robot text-space joint control method provided in the above embodiments, and other technical features in this robot text-space joint control device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0096] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0097] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0098] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the robot text space joint control method in the above embodiments.

[0099] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0100] The aforementioned computer-readable storage medium may be included in the robot text space joint control device; or it may exist independently and not be assembled into the robot text space joint control device.

[0101] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the robot text space joint control device, cause the robot text space joint control device to execute the aforementioned robot text space joint control method.

[0102] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0104] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0105] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described robot text-space joint control method, which can solve the technical problem of low control stability in current robot control methods. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the robot text-space joint control method provided in the above embodiments, and will not be repeated here.

[0106] All user-related data involved in this application was obtained with the user's permission or consent, as per [reference]. Figure 5 In other words, when this application is applied to a specific product or technology, user permission is required to acquire and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.

[0107] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for joint text-space control of a robot, characterized in that, The method includes: Acquire the robot's textual semantic instructions and spatial control constraint data; The text semantic instructions and the spatial control constraint data are input into the motion trajectory generation model to obtain the executable trajectory of the robot. The motion trajectory generation model is trained on a preset motion diffusion model based on preference sample pairs. The preference sample pairs are constructed based on the scores of the robot's motion trajectory on hardware constraints, dynamic constraints, tracking accuracy constraints, and semantic matching degree constraints. The training loss value includes a physical compliance penalty term, which is determined based on the hardware constraints and the dynamic constraints. The robot is controlled based on the executable trajectory.

2. The method as described in claim 1, characterized in that, Before the step of inputting the text semantic instructions and the spatial control constraint data into the motion trajectory generation model, the method further includes: Obtain text semantic instruction samples and spatial control constraint samples, and extract features from the text semantic instruction samples and the spatial control constraint samples respectively to obtain text semantic features and spatial control features; The text semantic features and spatial control features are fused to obtain fused features, which are then input into the motion diffusion model to obtain multiple candidate motion sequences for the robot. The candidate motion sequences are subjected to hardware constraint verification, and the compliant motion sequences obtained from the verification are input into a preset simulation model to obtain full simulation parameters. The simulation model is constructed based on the physical parameters of the robot. Based on the full simulation parameters and the candidate motion sequences, the preference sample pairs are constructed; Based on the preferred sample pairs, the network weight parameters of the motion diffusion model are adjusted to obtain the motion trajectory generation model.

3. The method as described in claim 2, characterized in that, The hardware constraint verification includes kinematic constraints. The step of performing hardware constraint verification on the candidate motion sequence and inputting the verified compliant motion sequence into a preset simulation model to obtain full simulation parameters includes: Based on the candidate motion sequence, determine the kinematic parameters of the robot when executing the candidate motion sequence; Based on the kinematic parameters, invalid motion sequences that do not conform to the kinematic constraints are identified from each of the candidate motion sequences, and the invalid motion sequences are deleted from the candidate motion sequences to obtain compliant motion sequences. The full set of simulation parameters is obtained by inputting the compliant motion sequence into the simulation model.

4. The method as described in claim 3, characterized in that, The step of constructing the preference sample pairs based on the full simulation parameters and the candidate motion sequences includes: Determine the number of non-compliance items in the candidate motion sequence where the kinematic parameters do not conform to the kinematic constraints, and calculate the hardware constraint score based on the number of non-compliance items; Based on the zero-moment point duration, zero-moment point offset distance, centroid height fluctuation range, centroid acceleration change rate, and foot-ground slippage in the full simulation parameters, the dynamic constraint score is calculated. Calculate the spatial mean square error between the spatial trajectory in the full simulation parameters and the preset ideal spatial trajectory, and calculate the tracking accuracy constraint score based on the mean square error; Extract the features of the spatial trajectory to obtain spatial trajectory features, calculate the cosine similarity between the spatial trajectory features and the text semantic features, and calculate the semantic matching degree constraint score based on the cosine similarity. The preference sample pairs are constructed based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score.

5. The method as described in claim 4, characterized in that, The step of constructing the preference sample pair based on the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score includes: The target task of the robot is determined based on textual semantic instructions; Based on the target task, determine the reward score weights corresponding to the hardware constraint score, the dynamic constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score, respectively. The hardware constraint score, the dynamics constraint score, the tracking accuracy constraint score, and the semantic matching degree constraint score are multiplied by their respective corresponding reward score weights to obtain the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score. The preference sample pairs are constructed based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score.

6. The method as described in claim 5, characterized in that, The step of constructing the preference sample pair based on the hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score includes: The hardware reward score, the dynamics reward score, the tracking accuracy reward score, and the semantic matching degree reward score are added together to obtain the comprehensive reward score; Based on the comprehensive reward score, each of the compliant sports sequences is sorted to obtain the compliant sports sequence ranking; The physical compliance score is obtained by adding the hardware reward score and the dynamics reward score. Based on the ranking of the compliant motion sequences, motion sequences with physical compliance scores higher than a preset winning compliance threshold are identified as winning samples. From the invalid motion sequences, motion sequences with physical compliance scores lower than a preset elimination compliance threshold are identified as elimination samples. Based on the winning samples and the eliminated samples, the preference sample pairs are constructed.

7. The method as described in claim 6, characterized in that, The step of adjusting the network weight parameters of the motion diffusion model based on the preference sample pairs to obtain the motion trajectory generation model includes: Based on the preferred sample pairs, the winning log probability of the winning sample and the elimination log probability of the losing sample in the motion diffusion model are calculated respectively. Based on the winning log probability, the elimination log probability and the preset preference optimization loss function, the loss value is calculated. Based on the preset ideal zero-moment point duration, ideal zero-moment point offset distance, ideal centroid height fluctuation range, ideal centroid acceleration change rate, and ideal foot-to-ground slip, the ideal dynamic constraint score is calculated, and the difference between the ideal dynamic constraint score and the dynamic constraint score is calculated to obtain the dynamic constraint score deviation. The physical compliance penalty item is obtained by adding the dynamic constraint score deviation and the hardware constraint score. Based on the physical compliance penalty and the loss value, the network weight parameters of the motion diffusion model are adjusted to obtain an iterative motion diffusion model; Determine whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets the preset iteration termination condition. If it does, then the iterative motion diffusion model is used as the motion trajectory generation model.

8. The method as described in claim 7, characterized in that, The step of determining whether the iterative candidate motion sequence generated by the iterative motion diffusion model meets the preset iteration termination condition, and if so, using the iterative motion diffusion model as the motion trajectory generation model, includes: The iterative candidate motion sequence is used as the candidate motion sequence. The process of performing hardware constraint verification on the candidate motion sequence is returned. The compliant motion sequence obtained from the verification is input into the preset simulation model to obtain the full set of simulation parameters. The invalid motion sequence, the tracking accuracy constraint score and the semantic matching degree constraint score are obtained. If the number of invalid motion sequences is lower than a preset number threshold, the tracking accuracy constraint score is higher than a preset tracking accuracy constraint score threshold, and the semantic matching degree constraint score is higher than a preset semantic matching degree score threshold, then the iterative candidate motion sequence is determined to meet the iteration termination condition, and the iterative motion diffusion model is used as the motion trajectory generation model.

9. A robot text-space joint control device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the robot text space joint control method as claimed in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the robot text space joint control method as described in any one of claims 1 to 8.