Control method, training method, control device, movable device, and storage medium

CN122525892APending Publication Date: 2026-08-07INTELLIGENT BODY TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTELLIGENT BODY TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-04-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明实施例的一个目的是提供一种可移动设备的新的控制技术方案,以解决相关技术中存在的末端执行器在复杂工况下控制精度低的问题

Benefits of technology

[0016] One beneficial effect of this invention is that it first obtains the state control torque pair sequence of the end effector within the most recent preset window length, and extracts conditional information containing current operating state information, current state information, and control torque from the sequence. This allows the model to perceive historical motion trends and changes in operating conditions, thus avoiding ambiguity caused by relying solely on instantaneous states. Subsequently, a generative model is used, guided by this conditional information, to model the complete conditional distribution of the comprehensive residual vector. The predicted residual value is then output through inverse denoising. This not only overcomes the limitations of single-point prediction in traditional methods but also preserves the multimodal characteristics of the residual distribution, significantly improving the modeling accuracy for nonlinear and non-stationary disturbances. Finally, the high-precision residual prediction value is combined with adaptive gain to determine the current control torque, achieving real-time compensation for residual prediction errors. Through repeated iterations, a continuously optimized closed-loop control is formed. This series of mechanisms works together to enable the end effector to maintain high-precision trajectory tracking and stable control performance even under complex operating conditions such as sudden load changes, configuration switching, and speed variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122525892A_ABST
    Figure CN122525892A_ABST
Patent Text Reader

Abstract

The application discloses a control method, a training method, a control device, a movable device and a storage medium, wherein the method comprises the following steps: acquiring a state control torque pair sequence of an end effector of a movable device in a recent preset window length; determining condition information according to the state control torque pair sequence; modeling a conditional distribution of a comprehensive residual error vector of the end effector by using a generation model and taking the condition information as a guide, and determining a predicted residual error value of the end effector by reverse denoising; determining a current control torque of the end effector according to reference state information, current state information of a current moment of the end effector, the predicted residual error value and an adaptive gain of a previous moment; controlling the end effector to move according to the current control torque; and repeatedly performing the above steps until the end effector reaches a preset terminal position and completes a preset operation action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic control technology, and more specifically, to a control method, a training method, a control device, a movable device, and a computer-readable storage medium. Background Technology

[0002] Industrial robots mainly include mobile robotic arms and fixed-base industrial robots (such as multi-joint serial robotic arms), and are widely used in complex operation scenarios such as inspection, contact operation, and load transportation. During operation, the end effector of such robots needs to cope with rapid changes in variable loads (such as grasping and releasing objects) and variable configurations (robotic arm joint movements), which causes abrupt changes in inertial parameters, centroid shift, and strong inertial coupling forces, making the dynamic model exhibit obvious nonlinear and non-stationary characteristics.

[0003] To address the aforementioned issues, traditional analytical modeling methods construct nominal models based on the Euler-Lagrange equations, but these methods struggle to accurately characterize non-modeled disturbances such as contact impacts and inertial abrupt changes, leading to a sharp decline in model fidelity as operating conditions change. To compensate for the shortcomings of analytical models, existing data-driven methods have been introduced: deep neural networks only learn deterministic mappings and cannot handle the distribution characteristics of residual forces, resulting in poor prediction reliability under unknown loads or configurations. Furthermore, Gaussian processes, constrained by smoothness assumptions, struggle to capture transient nonlinear interactions.

[0004] In summary, existing technologies cannot meet the control accuracy requirements of end effectors under complex working conditions such as variable loads and configurations. Summary of the Invention

[0005] One objective of this invention is to provide a new control technology solution for mobile devices to solve the problem of low control accuracy of end effectors under complex working conditions in related technologies.

[0006] According to a first aspect of the present invention, a control method is provided, comprising: Obtain the sequence of state control torque pairs of the end effector of the mobile device within the most recent preset window length; Based on the state control torque pair sequence, condition information is determined; wherein, the condition information includes the current operating state information and current state information of the end effector, as well as the control torque at the previous moment; By generating a model guided by the conditional information, the conditional distribution of the comprehensive residual vector of the end effector is modeled, and the predicted residual value of the end effector is determined by inverse denoising. The current control torque of the end effector is determined based on the current reference state information, current state information, predicted residual value, and adaptive gain of the previous moment. The end effector is moved according to the current control torque; Repeat the above steps until the end effector reaches the preset endpoint position and completes the preset operation.

[0007] Optionally, determining the current control torque of the end effector based on the current reference state information, current state information, predicted residual value, and adaptive gain from the previous moment includes: Based on the current state information and the reference state information at the current moment, the tracking error and error variables are determined; wherein, the reference state information at the current moment includes the desired position, desired velocity, and desired acceleration; The adaptive gain at the previous time step is updated based on the error variable to obtain the adaptive gain at the current time step. The current control torque is determined based on the nominal inertia matrix of the end effector, the desired acceleration, the error variable, the predicted residual value, and the adaptive gain at the current moment.

[0008] Optionally, before obtaining the sequence of state control torque pairs of the end effector of the mobile device at the most recent preset window length, the method further includes: Initialize the control torque and adaptive gain of the end effector, and clear the history buffer. Obtain the current state information of the end effector, and store the current state information and the currently used control torque as a state control torque pair in the history buffer; Determine whether the number of state control torque pairs stored in the historical buffer reaches the preset window length; If not, the predicted residual value of the end effector is set to zero, and the current control torque of the end effector is determined based on the current state information and the reference state information at the current moment. In this case, the step of obtaining the state control torque pair sequence of the end effector of the mobile device at the most recent preset window length is performed.

[0009] According to a second aspect of this application, a training method is also provided, comprising: Obtain a training sample set; where each training sample includes a sequence of sample state control torque pairs and sample residual values; Based on the sample state control torque pair sequence, sample condition information is determined; wherein, the condition information includes sample operation state information, sample state information, and sample control torque; Noise is added to the sample residuals of the training samples to obtain a dataset with added noise; The sample condition information corresponding to the training samples and the dataset with added noise are input into the generative model, so that the generative model predicts the noise added to the dataset based on the sample condition information, and obtains the predicted noise value. A loss function is constructed using the loss of the predicted noise value relative to the true noise value, and the model parameters of the generated model are updated.

[0010] Optionally, obtaining the training sample set includes: Obtain the dataset; the dataset includes a sequence of state control torque pairs of the end effector under different operating conditions; The dataset is divided into multiple samples according to a preset window length; each sample includes a sequence of sample state control torque pairs. For any sample, the sample residual value is determined based on the sample state control torque pair sequence and the preset comprehensive residual vector calculation formula. The sample state control torque pair sequence and sample residual value of a sample are used as a training sample to obtain the training sample set corresponding to the dataset.

[0011] Optionally, determining the sample condition information based on the sample state control torque pair sequence includes: The sample state control torque pair sequence of the sample is input into the feature extraction model to obtain the sample operation state information of the sample; The step of constructing a loss function based on the loss between the predicted noise value and the true noise value, and updating the model parameters of the generated model, includes: A loss function is constructed using the loss of the predicted noise value relative to the true noise value, and the model parameters of the generative model and the parameters of the feature extraction model are updated.

[0012] Optionally, the preset formula for calculating the comprehensive residual vector is determined through the following steps: The original dynamic model of the end effector is established based on the Euler-Lagrange equations. The original dynamic model includes inertial force terms, Coriolis force and centripetal force terms, gravity terms, and aerodynamic disturbance terms. Based on the nominal inertia matrix, the inertial force term in the original dynamic model is decomposed into a nominal inertia term and a comprehensive residual term; Based on the definition of the comprehensive residual term, the non-nominal inertial force, Coriolis force, centripetal force, gravity, and aerodynamic disturbance in the original dynamic model are integrated into the comprehensive residual vector, and the preset comprehensive residual vector calculation formula is obtained.

[0013] According to a third aspect of the invention, a control device is also provided, comprising a memory and a processor, the memory being configured to store executable instructions; the processor being configured to operate under the control of the instructions to perform the methods described in the first and second aspects.

[0014] According to a fourth aspect of the invention, a mobile device is also provided, including the control device as described in the second aspect.

[0015] According to a fifth aspect of the invention, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the methods described in the first and second aspects.

[0016] One beneficial effect of this invention is that it first obtains the state control torque pair sequence of the end effector within the most recent preset window length, and extracts conditional information containing current operating state information, current state information, and control torque from the sequence. This allows the model to perceive historical motion trends and changes in operating conditions, thus avoiding ambiguity caused by relying solely on instantaneous states. Subsequently, a generative model is used, guided by this conditional information, to model the complete conditional distribution of the comprehensive residual vector. The predicted residual value is then output through inverse denoising. This not only overcomes the limitations of single-point prediction in traditional methods but also preserves the multimodal characteristics of the residual distribution, significantly improving the modeling accuracy for nonlinear and non-stationary disturbances. Finally, the high-precision residual prediction value is combined with adaptive gain to determine the current control torque, achieving real-time compensation for residual prediction errors. Through repeated iterations, a continuously optimized closed-loop control is formed. This series of mechanisms works together to enable the end effector to maintain high-precision trajectory tracking and stable control performance even under complex operating conditions such as sudden load changes, configuration switching, and speed variations. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0018] Figure 1 This is a schematic diagram of the structure of a mobile device according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a control method according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating a training method according to an embodiment of the present invention. Detailed Implementation

[0019] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0020] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0021] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0022] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0023] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0024] Figure 1 A schematic diagram of the structure of a mobile device according to one embodiment is shown.

[0025] Mobile devices can be industrial robots, quadruped robots, bipedal humanoid robots, exoskeleton robots, CNC machine tools, large lifting equipment, and other electromechanical equipment that have problems with variable configuration, variable load, and non-stationary disturbances; there are no restrictions here.

[0026] For example, the mobile device can be a wheeled mobile robot platform equipped with a 2R serial robotic arm, while also being compatible with ground mobile platforms such as quadruped robot dogs, comprehensively covering complex industrial operation scenarios such as ground inspection, load transportation, contact operation, and material grasping and release, and adapting to working environments with frequent non-stationary dynamic disturbances.

[0027] like Figure 1 As shown, the mobile device 10 includes a data acquisition device 110 and a control device 120.

[0028] The acquisition device 110 is used to acquire the current state information and state control torque pair sequence of the end effector. It is the core sensing unit for realizing state perception, residual calculation and closed-loop control. The sampling frequency of all acquisition modules is consistent with the control frequency to ensure data synchronization and meet the real-time control requirements.

[0029] The data acquisition device includes the following sub-modules: A joint encoder is used to collect motion state information of the end effector. This motion state information includes the joint position, rotation angle, and rotational speed of the end effector.

[0030] Continuing with the example of the mobile robot platform above, the joint encoder uses the high-precision encoder that comes with the Dynamixel XM430-W210-T torque control motor. It collects data on the two degrees of freedom of the 2R serial robotic arm, and provides real-time feedback on the angular displacement, angular velocity and angular acceleration of the robotic arm joints. The sampling frequency is set to 100Hz, and the collected data is directly transmitted to the control device to provide basic motion data for state coding and residual prediction.

[0031] For multi-degree-of-freedom industrial robotic arms, the number of encoders can be increased accordingly to adapt to different degree-of-freedom configurations.

[0032] An inertial measurement unit (IMU) is used to acquire attitude information of the end effector. Attitude information includes body attitude, acceleration, and angular velocity data.

[0033] Continuing with the mobile robot platform example above, the IMU module is integrated into the mobile robot chassis control board and used in conjunction with a high-performance motion controller to collect the robot's three-axis acceleration, three-axis angular velocity, and attitude angle data in real time. Combined with the 120fps positioning data from the OptiTrack motion capture system, it achieves accurate fusion of state information under generalized coordinates, effectively compensating for the drift error of a single IMU and accurately representing attitude changes during non-stationary motion.

[0034] The storage module is used to store the sequence of state control torque pairs with the most recent preset window length.

[0035] Continuing with the mobile robot platform example above, the storage module uses the high-speed RAM of the onboard Jetson Orin Nano Super to build a rolling history buffer. The preset window length L is set to 10 control cycles. The buffer only retains the state control torque pair sequence of the most recent 10 control cycles. It is updated in real time with new data and expired data is removed to avoid data redundancy. At the same time, it ensures the integrity of the historical time sequence features required for encoding the current operation state information. The buffer data read and write latency is less than 1ms, which meets the real-time control requirements.

[0036] In some examples, the acquisition device also includes force sensors to acquire end-contact forces and external disturbance load signals for use in the calculation of the composite residual vector.

[0037] Continuing with the mobile robot platform example above, a force sensor is installed at the end effector of the robotic arm to collect the contact impact force and the magnitude of the external disturbance load (within the range of 0g-500g) during the grasping process in real time. The collected force signal is converted into an electrical signal and transmitted to the control device to assist in the accurate calculation of the comprehensive residual vector including the external disturbance, thereby further improving the modeling accuracy of non-modeled disturbances.

[0038] In one example, such as Figure 1 As shown, the control device 120 includes a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, etc. The entire control device is divided into two parts: the airborne end side and the ground end side, which work together to realize model reasoning and control execution.

[0039] The processor 1100 can be a mobile processor.

[0040] Continuing with the mobile robot platform example above, the onboard processor uses an NVIDIA Jetson Orin NanoSuper embedded GPU, which is responsible for encoding the current operation state information, lightweight generative model inference, and real-time calculation of adaptive control laws, meeting the real-time control inference requirements of 50Hz and above; the ground-side processor uses a high-performance processor equipped with an NVIDIA RTX 4080 GPU, which is responsible for offline model training, data preprocessing, and visualization of experimental results, taking into account both the low latency of the embedded end and the high performance of the training end.

[0041] The memory 1200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as hard disks. ROM is used to store the fixed control program, fixed model parameters, and hardware drivers; RAM is used to cache real-time state acquisition data, buffer timing data, and intermediate calculation results; non-volatile memory is used to store offline training datasets, trained model weight files, experimental logs, and trajectory data, supporting data retention after power failure, facilitating subsequent model fine-tuning and experimental reproduction.

[0042] The interface device 1300 includes, for example, a USB interface, a headphone jack, a charging interface, etc.

[0043] Continuing with the mobile robot platform example above, the interface device also includes a CAN bus interface and a Micro XRCE-DDS communication interface, which are used to connect the robotic arm motor, the chassis motion control system, and the onboard computing unit, respectively, to achieve high-speed transmission of control torque and sensing data. The interface transmission rate matches the control frequency, and there are no data packet loss or delay issues.

[0044] The communication device 1400 may be capable of wired or wireless communication, and may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1400 may also include long-range communication devices, such as any device that performs WLAN, GPRS, 2G / 3G / 4G / 5G long-range communication.

[0045] Continuing with the mobile robot platform example above, the preferred communication device is a WiFi wireless communication module with a communication frequency of 50Hz. This enables real-time data interaction between the onboard computing unit and the ground-based end-side and chassis control system, supports wireless remote control, adapts to the long-distance operation requirements of industrial sites, and is also compatible with wired Ethernet communication to ensure communication stability under complex working conditions.

[0046] The display device 1500 is, for example, an LCD screen, a touch screen, etc.

[0047] In this embodiment, the display device is used to display the trajectory tracking status of the end effector, residual prediction error, adaptive gain change and system operating status in real time, so as to facilitate operators to monitor the operation process in real time and troubleshoot abnormalities in a timely manner.

[0048] The input device 1600 may include, for example, a touch screen, a keyboard, etc.

[0049] In this embodiment, the input device is used to input the desired trajectory parameters, load parameters, robotic arm configuration parameters, and control gain parameters to achieve flexible configuration of the task.

[0050] Users can input / output voice information through the speaker 1700 and microphone 1800 for remote voice control and system malfunction alarms, improving ease of operation.

[0051] In this embodiment, the memory 1200 of the control device 120 is used to store instructions for controlling the processor 1100 to operate in order to at least execute the control method according to any embodiment of this application. Those skilled in the art can design instructions based on the schemes disclosed in this application. How the instructions control the processor to operate is well known in the art and will not be described in detail here.

[0052] Despite Figure 1 The control device 120 is shown in the figure, but this application may only involve some of the devices. For example, the control device 120 may only involve the memory 1200 and the processor 1100.

[0053] Figure 2This is a flowchart illustrating a control method according to an embodiment of the present invention. The method can be... Figure 1 The control device 120 is implemented in the device and is suitable for real-time trajectory tracking and residual dynamics compensation of the end effector of the mobile device. According to Figure 2 As shown, the control method of this embodiment may include the following steps S2100~S2600: Step S2100: Obtain the state control torque pair sequence of the end effector of the mobile device at the most recent preset window length.

[0054] In this embodiment, the state control torque pair sequence refers to the time-series data sequence formed by pairing the state information collected at each step of the mobile device's end effector within a continuous control cycle with the corresponding output control torque. This sequence is arranged chronologically according to the control cycle and possesses strict temporal correlation. It is used to fully characterize the state changes and control response relationship during the continuous motion of the end effector, providing complete temporal input for subsequent state feature extraction and residual prediction. The state information is the core data set characterizing the real-time motion attitude and position of the end effector, including position information, velocity information, and acceleration information in generalized coordinates.

[0055] Continuing with the mobile robot platform example above, the status information specifically covers the 6-dimensional spatial motion state of the mobile robot chassis (two-dimensional planar position and heading angle, and three-dimensional spatial position and attitude angle, which can be selected according to the platform type) and the motion state of the two joints of the robotic arm (joint angular displacement, joint angular velocity, and joint angular acceleration), totaling 8-dimensional generalized coordinate status data. All of these are synchronously collected by the joint encoder, inertial measurement unit, and motion capture system, accurately reflecting the real-time motion status of the end effector.

[0056] Control torque can be a generalized torque used in relation to state information. Control torque can be output by a control device and is a direct control signal that drives the end effector to perform motion actions.

[0057] Control torque can be a generalized torque used in relation to state information. Control torque can be output by a control device and is a direct control signal that drives the end effector to perform motion actions.

[0058] Continuing with the mobile robot platform example above, the control torque can include the wheel drive torque or track drive torque of the mobile robot chassis, the steering torque, and the drive torque of each joint of the robotic arm. The control torque corresponds one-to-one with the state information; that is, each set of state information corresponds to a unique set of control torques used for that state information, together forming a single state control torque pair.

[0059] For example, the preset window length is fixed at 10 control cycles, that is, acquiring the continuous state control torque timing data of the most recent 10 control cycles. This length is the optimal value after multiple rounds of experimental optimization, which can fully represent the trend of operation state changes without increasing the amount of computation too much, and is suitable for real-time inference on embedded devices. This step to obtain a valid sequence is only executed when the amount of historical buffer data reaches 10 control cycles; otherwise, the basic control mode is activated.

[0060] In one embodiment of this application, before obtaining the state control torque pair sequence of the end effector of the mobile device at the most recent preset window length, the method further includes: steps S1100 to S1500.

[0061] Step S1100: Initialize the control torque and adaptive gain of the end effector, and clear the history buffer.

[0062] In this embodiment, the initial control torque τ is set to 0, meaning there is no control torque output in the initial state. Adaptive gain initial value. .

[0063] The initial value of the adaptive gain is, for example, 2.0. This value is the optimal initial value for the experiment, ensuring the stability of error compensation in the initial stage.

[0064] At the same time, clear all stored data in the historical buffer to avoid residual historical data interfering with subsequent state coding and residual prediction, ensuring that the system starts from a clean initial state.

[0065] Step S1200: Obtain the current state information of the end effector, and store the current state information and the currently used control torque as a state control torque pair in the history buffer.

[0066] Continuing with the mobile robot platform example above, the current state information includes 8-dimensional generalized coordinate information: 6-dimensional chassis motion state + 2-dimensional robotic arm state. The initial control torque is 0, and subsequent torque commands are calculated and generated from the previous moment. Each time new data of state information and control torque is collected, it is combined into a state-control torque pair and stored in the history buffer in chronological order, achieving real-time data updates and storage.

[0067] Step S1300: Determine whether the number of state control torque pairs stored in the historical buffer reaches the preset window length.

[0068] In this embodiment, the control device counts the number of state control torque pairs in the historical buffer memory in real time, and quickly determines whether the preset window length has been reached through a counting program. This determination step takes very little time and does not affect the real-time performance of the overall control cycle.

[0069] Step S1400: If not, determine the predicted residual value of the end effector as zero, and determine the current control torque of the end effector based on the current state information and the reference state information at the current moment.

[0070] In this embodiment, when the amount of historical buffer data is less than the preset 10 control cycle window length, there is a lack of sufficient time-series state control torque data, making it impossible to generate effective operation state descriptors (i.e., operation state information) through the feature extraction model, and also impossible to complete comprehensive residual prediction based on the generation model. Therefore, the generation model and feature extraction model are directly disabled, and the predicted residual value is forcibly set to 0.

[0071] At this point, the system switches to basic adaptive control mode, relying solely on the current state information and the reference state information at the current moment to calculate the control torque. Specifically, first, the current state information collected in real time by the acquisition device is extracted, namely the actual position, actual velocity, and actual acceleration of the end effector in the generalized coordinate system. Simultaneously, the preset reference state information at the current moment is retrieved, which includes the desired position, desired velocity, and desired acceleration of the end effector. Then, the trajectory tracking error is calculated based on the current state information and the reference state information at the current moment. Combined with the preset nominal inertia matrix and basic error gain parameters, the basic control torque is derived through an adaptive control law. Throughout the process, it relies only on the real-time perceived actual state and the preset desired state, without the need for residual compensation terms.

[0072] Step S1500: If so, perform the step of obtaining the state control torque pair sequence of the end effector of the mobile device at the most recent preset window length.

[0073] In this embodiment, when the buffer accumulates state control torque pairs of a preset window length, the sequence of the most recent state control torque pairs of the preset window length is extracted to prepare for subsequent condition information generation and residual prediction, and then the system switches to high-precision residual compensation control mode.

[0074] Step S2200: Determine condition information based on the state control torque pair sequence; wherein, the condition information includes the current operating state information and current state information of the end effector, as well as the control torque at the previous moment.

[0075] In this embodiment, firstly, the state control moment pair sequence is input into a feature extraction model to obtain the current operation state information of the end effector. The feature extraction model can be a shallow temporal convolutional network. This shallow temporal convolutional network uses two dilated one-dimensional convolutional layers to globally capture features from the state control moment pair sequence, filter out invalid noise, and output 32-dimensional current operation state information.

[0076] The feature extraction model can also be a gated recurrent unit (GRU) / lightweight LSTM, which extracts the current operation state information through sequence modeling and is suitable for scenarios with higher requirements for feature capture.

[0077] The feature extraction model can also adopt a lightweight architecture of attention mechanism + MLP, which only focuses on key state control moment pairs in the historical window, further reducing computational complexity and making it suitable for scenarios with limited airborne computing resources.

[0078] Current operational status information can be used to characterize the operational state of the end effector. This operational state can be one of the following: load grabbing, aerial transport, material release, or smooth hovering.

[0079] Current operation status information can prevent the aliasing of residual features from different operation modes under the same instantaneous state.

[0080] Then, the state information in the latest state control torque pair in the state control torque pair sequence is taken as the current state information, and the control torque in the latest state control torque pair is taken as the control torque at the previous moment.

[0081] Finally, the three types of data—current operation status information, current status information, and control torque at the previous moment—are combined by dimension splicing and feature fusion to form complete conditional information.

[0082] Taking a mobile robotic arm platform as an example, when the end effector is in the 200g load grasping stage, the 32-dimensional current operation state information extracted by the feature extraction model, combined with the real-time position and posture information collected at this time (i.e., current state information) and the control torque at the previous moment, forms conditional information.

[0083] Compared to traditional input methods that rely solely on instantaneous states, this solution uses current operational state information as conditional information to accurately match the dynamic residual characteristics of different operational stages, significantly improving the relevance and accuracy of subsequent residual prediction. This effectively solves the problem that a single instantaneous state cannot distinguish operational modes, leading to confusion in residual prediction.

[0084] Step S2300: Using the conditional information as a guide, the conditional distribution of the comprehensive residual vector of the end effector is modeled by generating a model, and the predicted residual value of the end effector is determined by inverse denoising.

[0085] In this embodiment, firstly, a latent variable with the same dimension as the target residual is sampled from a standard Gaussian distribution as the initial state for the inverse denoising process. Subsequently, conditional information (including current state information, the control torque from the previous moment, and the current operation state information) is synchronously input into the denoising network. Guided by the conditional information, the denoising network progressively performs inverse denoising in K preset diffusion steps. In each denoising step, based on the current noisy residual, the conditional information, and the current diffusion step number, the denoising network predicts the Gaussian noise component injected in that step and removes this noise component according to a preset inverse update formula, thereby restoring a distribution closer to the true distribution.

[0086] By iteratively executing this reverse denoising process, the model eventually transforms the initial Gaussian noise into predicted residual values ​​that can accurately characterize the real physical interaction properties.

[0087] The generative model can be a conditional diffusion model based on a denoising diffusion probability model. This model takes conditional information as its core and models the generation process of the comprehensive residual vector of the end effector as a progressive denoising process from isotropic Gaussian noise to the target residual distribution.

[0088] Furthermore, while meeting the requirements for high-precision modeling, generative models can also be variably replaced based on real-time requirements. In scenarios where sacrificing a small amount of prediction accuracy for higher inference efficiency is acceptable, lightweight variants of diffusion models (such as DDIM or IDDPM) can be used to significantly improve inference speed by reducing the number of iterations for inverse denoising. In applications with extremely high real-time requirements and moderate requirements for distribution representation accuracy, conditional generative adversarial networks (CGANs) can be used as an alternative, leveraging their efficient implicit distribution fitting capabilities to further reduce inference latency and meet the system's high dynamic response requirements.

[0089] Conditional distribution can be the probability distribution of the comprehensive residual vector under the joint constraints of three types of conditions: current operating state information, current state information, and control torque at the previous moment. This differs from traditional unconditional modeling and can significantly improve the relevance and accuracy of residual prediction.

[0090] Through the aforementioned mechanism, this generative model not only avoids the problem of insufficient accuracy of single deterministic predictions under complex operating conditions, but also effectively learns the complete conditional distribution of residual force under conditional information, achieving high-precision modeling and prediction of residual force distribution under multiple operating conditions and states. Compared with traditional point estimation methods, this approach can preserve the multimodal characteristics of the distribution during model inference, providing more statistically robust input for subsequent control compensation and decision-making.

[0091] Step S2400: Determine the current control torque of the end effector based on the reference state information, current state information, predicted residual value, and adaptive gain of the previous moment.

[0092] In this embodiment, the reference state information at the current moment originates from the pre-planned desired trajectory waypoints of the end effector, which are command signals input as the tracking target during online control. The reference state information at the current moment can be the desired state information of the end effector at the current moment.

[0093] The adaptive gain of the previous moment can be the error compensation weight of the previous control cycle.

[0094] In one embodiment of this application, step S2400 determines the current control torque of the end effector based on the reference state information of the end effector at the current moment, the current state information, the predicted residual value and the adaptive gain at the previous moment, including steps S2400.1 to S2400.3.

[0095] Step S2400.1: Determine the tracking error and error variables based on the current state information and the reference state information at the current moment; wherein, the reference state information at the current moment includes the desired position, desired velocity, and desired acceleration.

[0096] In this embodiment, the tracking error is first determined based on the current state information and the reference state information at the current moment. The tracking error characterizes the degree of deviation between the actual state of the end effector at the current moment and the reference state, and is a core metric for trajectory tracking accuracy. The larger the tracking error value, the farther the actual motion deviates from the reference trajectory.

[0097] The reference state vector includes the desired position, desired velocity, and desired acceleration.

[0098] For example, the reference state vector could be the target point (1.0, 1.0, 1.5)m on a preset ∞-shaped trajectory, the desired velocity of 0.5m / s, and the desired joint rotation angle of 25°.

[0099] The desired position is compared with the current actual position to form the positional deviation. The desired velocity is compared with the current actual velocity to form the velocity deviation. The desired acceleration, on the other hand, serves directly as a feedforward term, enabling the control device to predict the trajectory's changing trend and thus apply the appropriate control force in advance, ensuring that the end effector moves precisely along the predetermined trajectory.

[0100] The rate of change of the tracking error is then obtained by differentiating the tracking error, and the error variable is calculated by combining it with a preset positive definite gain matrix. This variable will serve as the core input for subsequent adaptive gain updates and control torque calculations.

[0101] The error variable can be a comprehensive error term that integrates the tracking error and its derivative, taking into account both position deviation and velocity deviation.

[0102] Based on this step, the deviation between the actual motion of the end effector and the target trajectory can be quantified, providing data support for subsequent gain updates and instruction calculations.

[0103] Step S2400.2: Update the adaptive gain of the previous time step according to the error variable to obtain the adaptive gain of the current time step.

[0104] In this embodiment, the adaptive gain is updated using an adaptive iterative law, eliminating the need for manual parameter tuning. The update rule of the adaptive gain closely follows the dynamic changes in error: when the error is large and the trajectory deviates significantly, the magnitude of the error variable is large, and the adaptive gain gradually increases to strengthen error compensation and offset the impact of disturbances. When the error gradually converges and the trajectory returns to the target, the magnitude of the error variable is small, and the adaptive gain smoothly decreases to avoid excessive compensation leading to motion jitter. The adaptive gain at the current moment can be a new parameter updated in real time, used for calculating the control torque in the current control cycle to achieve adaptive error compensation.

[0105] For example, if the end effector is currently in a heavy-load transportation phase, the error variable magnitude may be too large due to uneven ground or external collision disturbances. The adaptive gain was 2.2 at the previous moment, and after updating, the gain at the current moment increases to 2.5, strengthening the compensation. After the disturbance disappears and the error converges, the gain gradually drops back to around 2.0 to ensure stable control.

[0106] Based on this step, it is possible to dynamically correct adaptive correction parameters based on real-time motion deviation, avoiding a decrease in control accuracy caused by residual prediction errors or sudden external disturbances.

[0107] Step S2400.3: Determine the current control torque based on the nominal inertia matrix of the end effector, the desired acceleration, the error variable, the predicted residual value, and the adaptive gain at the current moment.

[0108] This embodiment uses an adaptive control law to calculate the current control torque, which consists of the following four components: The first term is the error correction term, which implements closed-loop feedback adjustment based on the error variable, enabling rapid reduction of tracking error. The positive definite gain matrix is ​​used to adjust the suppression strength of the overall error.

[0109] The second term is the nominal dynamics feedforward term, which is calculated using the user-predefined nominal inertia matrix and the desired acceleration. This nominal inertia matrix is ​​completely identical to the matrix used in the calculation of the integrated residual vector, eliminating the need for precise matching of the actual inertia matrix. Its purpose is to simplify dynamics modeling and reduce the complexity of the solution.

[0110] The third term is the residual compensation term, which introduces the predicted residual value output by the generative model to offset the effects of non-modeled dynamics such as inertial coupling and external disturbances.

[0111] The fourth term is the adaptive robust compensation term, which consists of the adaptive gain at the current moment, the direction vector of the error variable, and the norm of the error variable. The direction vector directs the adaptive compensation force in the direction of error reduction, and the denominator uses the norm of the error variable to ensure that this direction term is a unit direction. In actual calculations, a smooth approximation is usually introduced to avoid numerical problems when the error approaches zero. Furthermore, another positive definite gain matrix is ​​used to accelerate the convergence of the tracking error; this matrix is ​​multiplied by the velocity error to form the velocity error correction term.

[0112] The control torque at the current moment can be obtained by combining the above four factors.

[0113] Step S2500: Control the movement of the end effector according to the current control torque.

[0114] In this embodiment, if the current control torque increases the joint torque and the body thrust, the end effector will respond quickly to complete the weight-grabbing action, while simultaneously fine-tuning its attitude to counteract load inertia and maintain stable movement. If the control torque is a small correction torque, the end effector will fine-tune its position to return to the desired trajectory.

[0115] Taking the aforementioned mobile robot platform as an example, the detailed process of controlling the end effector's movement based on the current control torque is explained below: The control torque can be at least one of the following: the wheel drive torque of the mobile robot chassis, the steering torque, and the joint drive torque of the robotic arm. The end effector can be the actuator of a mobile robot equipped with a 2R tandem robotic arm, encompassing the robotic arm gripper and the robot body, capable of performing tasks such as load grasping, transporting, and releasing. First, the control device sends the current control torque to the chassis motion controller via a communication interface. This transmission delay is less than 5ms, ensuring real-time performance. The chassis motion controller converts the torque command into a motor drive signal, adjusting the wheel speed or track speed to achieve platform position and attitude control. The joint motor driver converts the torque command into motor drive current, controlling the precise rotation of the robotic arm joints.

[0116] Step S2600: Repeat the above steps until the end effector reaches the preset endpoint position and completes the preset operation.

[0117] In this embodiment, the entire process from S2100 to S2500 is executed once within each control cycle. Each cycle synchronously updates the historical buffer data, re-extracts conditional information, generates new prediction residual values, and calculates the real-time control torque, continuously correcting trajectory deviations. Furthermore, during the loop, two core termination conditions are monitored in real time, and the loop stops only when both are simultaneously met: First, the distance between the current position of the end effector and the preset endpoint position is less than a distance threshold (e.g., ±5mm), indicating that the end effector has reached the preset endpoint position. Second, the end effector completes all preset operation actions, with no remaining tasks.

[0118] The preset endpoint position can be a pre-stored target coordinate point, such as the material release point (-10, -1.0, 1.5) m and the landing origin (0.0, 0.0, 1.5) m in the experimental scenario.

[0119] The preset operation actions can be a sequence of actions that covers the entire operation cycle.

[0120] Taking the aforementioned mobile robot platform as an example, the preset operation actions include five stages: startup, load grabbing, ground transportation, fixed-point release, and return to the origin, covering the entire operation cycle.

[0121] This step is an iterative process of online closed-loop control, achieving fully autonomous closed-loop control without manual intervention.

[0122] Based on the above, this application first obtains the state control torque pair sequence of the end effector within the most recent preset window length, and extracts conditional information containing current operating state information, current state information, and control torque from the sequence. This enables the model to perceive historical motion trends and changes in operating conditions, thus avoiding ambiguity caused by relying solely on instantaneous states. Subsequently, using this conditional information as a guide, a generative model is used to model the complete conditional distribution of the comprehensive residual vector, and the predicted residual value is output through inverse denoising. This not only overcomes the limitations of single-point prediction in traditional methods but also preserves the multimodal characteristics of the residual distribution, significantly improving the modeling accuracy for nonlinear and non-stationary disturbances. Finally, the high-precision residual prediction value is combined with adaptive gain to determine the current control torque, achieving real-time compensation for residual prediction errors, and forming a continuously optimized closed-loop control through repeated iterations. This series of mechanisms work together to enable the end effector to maintain high-precision trajectory tracking and stable control performance even under complex operating conditions such as sudden load changes, configuration switching, and speed variations.

[0123] like Figure 3As shown, this application also includes a model offline training method. This method provides a high-precision, high-generalization model weight file for the aforementioned online closed-loop control. The trained model can be directly ported to an airborne embedded control device to achieve low-latency real-time residual inference. The entire process is divided into five core stages: data acquisition and preprocessing, standardized sample construction, conditional information extraction, noise injection training, and model parameter iterative optimization. It balances training efficiency and model accuracy, comprehensively covering various operating conditions of mobile devices, and solving the problem of residual prediction accuracy under non-stationary dynamic disturbances. Figure 3 The flowchart illustrates that this training method includes steps S3100 to S3500, and the detailed explanations of each step and sub-step are as follows: Step S3100: Obtain the training sample set; wherein each training sample includes a sequence of sample state control torque pairs and sample residual values.

[0124] In this embodiment, the sample state control torque pair sequence can refer to the continuous temporal state control torque pair data corresponding to a single sample, which is basically the same as the state control torque pair sequence in step S2100, and will not be elaborated here.

[0125] The sample residual value can refer to the true composite residual value corresponding to a single sample. It serves as label data for model training, guiding the model to learn the true residual distribution pattern.

[0126] In one embodiment of this application, step S3100 of obtaining the training sample set includes: steps S3100.1 to S3100.4.

[0127] Step S3100.1: Obtain the dataset; the dataset includes a sequence of state control torque pairs of the end effector under different operating conditions.

[0128] In this embodiment, the dataset is acquired through two combined methods: first, data collection from real experiments. For example, a wheeled mobile robot platform equipped with a 2R robotic arm uses a baseline PID controller to execute random smooth trajectories, load mutations, and variable configuration movements. Data is collected at a frequency of 100Hz using a data acquisition device to ensure realism. Second, data generation through simulation platforms. For example, a dynamic simulation model matching the real machine is built using IsaacGym and Gazebo to quickly generate a large amount of multi-condition data, overcoming the problems of low efficiency and high cost of real machine data collection. The dataset may include state control torque pairs sequences.

[0129] The dataset can be raw, continuous time-series data that has not undergone any cutting or pairing processing. The dataset contains state information and control torque data of the end effector throughout its motion under multiple operating conditions.

[0130] Operating conditions refer to the external environment and the state conditions of the end effector during operation.

[0131] The working conditions can be load conditions (0g no load, 200g medium load, 400g heavy load), motion conditions (low speed 0.5m / s, high speed 1.0m / s), configuration conditions (different joint extension angles of the 2R robotic arm), disturbance conditions (normal airflow, small sudden disturbance), etc., which are not limited here.

[0132] The meanings of the state control torque pairs in this step and the state control torque pairs in the state control torque pair sequence in step S2100 are basically the same, and will not be elaborated here.

[0133] Step S3100.2: Divide the dataset according to the preset window length to obtain multiple samples; each sample includes a sequence of sample state control torque pairs.

[0134] In this embodiment, a non-overlapping sliding window partitioning method is adopted. A window unit is defined as a preset window length (e.g., 20 steps, also known as 20 control cycles). The data is sequentially cut from the starting position of the original dataset, with no data overlap between adjacent samples, thus avoiding data duplication that could lead to model overfitting. After the partitioning is completed, incomplete samples with a length less than the preset window length are removed, leaving only complete and independent samples. This ensures that the temporal characteristics of each sample are complete and can accurately represent the dynamic characteristics of the corresponding working condition.

[0135] The preset window length can refer to the number of time steps contained in a single sample, which is set in advance. Each step corresponds to a state control torque pair of a control cycle.

[0136] By splitting the continuous raw dataset into independent samples of fixed length, the input format of the model is standardized, ensuring a stable training process without format anomalies.

[0137] Step S3100.3: For any sample, determine the sample residual value based on the sample state control torque pair sequence and the preset comprehensive residual vector calculation formula.

[0138] In this embodiment, the integrated residual vector can be the residual force obtained by integrating the inertial coupling error, Coriolis force, centripetal force, gravity term, external aerodynamic disturbance, and load variation disturbance that are difficult to model accurately in the original mechanical model of the end effector.

[0139] The formula for calculating the comprehensive residual vector is derived through the following steps S1 to S3: Step S1: Establish the original dynamic model of the end effector based on the Euler-Lagrange equations. The original dynamic model includes inertial force terms, Coriolis force and centripetal force terms, gravity terms, and aerodynamic disturbance terms.

[0140] In this embodiment, the Euler-Lagrange equations can be classical equations characterizing the dynamics of multi-degree-of-freedom electromechanical equipment. The original dynamic model can refer to a dynamic model that fully reproduces the actual motion and forces of the end effector, without simplification or omission, and includes all force terms.

[0141] The original dynamic model expression for the end effector is: the sum of the inertial force term, Coriolis and centripetal force term, gravity term, and aerodynamic disturbance term equals the control torque. Among these, the inertial force term reflects the inertial drag of the equipment and its load; the Coriolis and centripetal force terms are nonlinear time-varying terms; the gravity term reflects the gravitational influence of the equipment and its load; the aerodynamic disturbance term represents external airflow and sudden disturbances; and the control torque is the input term. This model comprehensively covers all dynamic influencing factors and forms the basis for calculating the true residual.

[0142] Step S2: Based on the nominal inertia matrix, decompose the inertial force term in the original dynamic model into a nominal inertia term and a comprehensive residual term.

[0143] In this embodiment, the original dynamic model is simplified and reconstructed as follows: the nominal inertia term multiplied by the acceleration, plus the comprehensive residual term, equals the control torque. The nominal inertia term is the simplified inertia matrix, and the comprehensive residual term includes all unmodeled or nonlinear dynamic factors such as Coriolis force, centripetal force, gravity, and aerodynamic disturbances.

[0144] By simplifying the modeling process, the model can focus on predicting residual terms.

[0145] The nominal inertia matrix can refer to a manually preset empirical inertia matrix, which is completely consistent with the matrix used in online control, without the need for precise calibration of the actual inertia parameters.

[0146] The nominal inertia term refers to the fixed part of the inertial force term that can be accurately modeled using the nominal matrix. The comprehensive residual term refers to the integrated term of the deviation part of the inertial force term and other nonlinear and disturbance terms, which is the core content that the model needs to predict.

[0147] Step S3: According to the definition of the comprehensive residual term, the non-nominal inertial force, Coriolis force, centripetal force, gravity and aerodynamic disturbance in the original dynamic model are integrated into the comprehensive residual vector to obtain the preset comprehensive residual vector calculation formula.

[0148] In this embodiment, all time-varying, nonlinear, and disturbance terms in the original dynamic model that cannot be accurately modeled by the nominal inertia matrix are included in the comprehensive residual vector, and the calculation formula for the comprehensive residual vector is finally obtained.

[0149] The formula for calculating the composite residual vector is: the actual inertia matrix minus the nominal inertia matrix equals the position term plus the Coriolis force and centripetal force terms plus the gravity term plus the aerodynamic disturbance term. Here, the actual inertia matrix is ​​the true inertia matrix.

[0150] For any given sample, the process of determining the sample residual value based on the sample state control torque pair sequence and the preset comprehensive residual vector calculation formula includes: First, extracting the position information, velocity information, acceleration information, and control torque at each moment from the sample state control torque pair sequence. Then, based on the position and velocity information at the current moment, calculating the actual inertia matrix, Coriolis force and centripetal force matrix, and gravity term. Finally, substituting the calculated terms (actual inertia matrix, Coriolis force and centripetal force matrix, and gravity term) together with the preset nominal inertia matrix into the comprehensive residual calculation formula, that is, integrating the non-nominal inertial force, Coriolis force, centripetal force, gravity, and external disturbance terms into a comprehensive residual vector, thereby obtaining the sample residual value corresponding to that sample.

[0151] Step S3100.4: Take the sample state control torque pair sequence and sample residual value of a sample as a training sample to obtain the training sample set corresponding to the dataset.

[0152] In this embodiment, the sample state control torque pair sequence for each sample is strictly paired one-to-one with the corresponding calculated sample residual value to ensure no mismatches or omissions. After pairing, the sample order is randomly shuffled and divided into a training sample set and a validation sample set in an 8:2 ratio. The divided training sample set has a uniform format and even distribution of operating conditions, requiring no additional preprocessing and can be directly input into the model for training.

[0153] Step S3200: Determine sample condition information based on the sample state control torque pair sequence; wherein, the condition information includes sample operation state information, sample state information, and sample control torque.

[0154] This step is basically the same as step S2200 above, and will not be elaborated here. For example, the sample state control torque pair sequence can be input into the feature extraction model to obtain the sample operation state information.

[0155] Step S3300: Add noise to the sample residual values ​​of the training samples to obtain a dataset with added noise.

[0156] Step S3400: Input the sample condition information corresponding to the training sample and the dataset with added noise into the generation model, so that the generation model predicts the noise added to the dataset based on the sample condition information and obtains the predicted noise value.

[0157] In this embodiment, the sample condition information, the residual data after adding noise, and the iteration step embedding vector are concatenated and used as input to the generative model. The generative model, through convolution, pooling, and upsampling operations, accurately predicts the corresponding injected noise based on the constraints of this sample condition information, outputting a predicted noise value with the same dimension as the ground truth noise. The core task of the generative model is to predict the injected noise, rather than directly predicting the residual, which corresponds to the reverse denoising logic described earlier, ensuring compatibility between training and online inference processes.

[0158] Step S3500: Construct a loss function using the loss of the predicted noise value relative to the true noise value, and update the model parameters of the generated model.

[0159] In this embodiment, the model parameters of the generated model are updated based on the loss function until the training stops under certain conditions, such as the loss function converging or reaching a set number of training steps, which are not limited here.

[0160] The loss function is used to characterize the deviation between the predicted noise value and the true noise value. The smaller the deviation, the higher the model accuracy. Model parameters refer to the trainable parameters such as weights and biases inside the generated model. Accurate prediction is achieved through iterative updates and optimization.

[0161] For example, the parameters of the generative model can be updated using the Adam adaptive optimizer, with a learning rate of 0.0002, a batch size of 256, and a total training step count of 50k steps.

[0162] In an embodiment where the system includes a feature extraction model, step S3500 involves constructing a loss function based on the loss of the predicted noise value relative to the true noise value, and updating the model parameters of the generating model, including: A loss function is constructed using the loss of the predicted noise value relative to the true noise value, and the model parameters of the generative model and the parameters of the feature extraction model are updated.

[0163] In this embodiment, if the system includes a feature extraction model, the gradient will synchronously update the generation model parameters and the feature extraction model parameters to achieve dual-model collaborative optimization, ensuring the consistency between state feature extraction and residual noise prediction, without the need to train the encoder separately.

[0164] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above-described method embodiments.

[0165] In addition, experiments were designed to verify the residual prediction accuracy of this application under complex working conditions, as detailed below: 1. Hardware and Experimental Environment Configuration 1.1 Hardware Platform Configuration The experimental verification of this invention is based on a self-built end effector hardware platform. To verify the spatial effectiveness of this method, a wheeled mobile robot platform is used, and the working conditions fully cover the planar working conditions of a quadruped robot dog. The robotic arms adopt a 2R tandem robotic arm configuration, with the following specific configuration: Mobile chassis: Customized aluminum alloy chassis, equipped with four independent drive motors (each wheel is equipped with a brushless DC motor), and the tires are omnidirectional wheels or Mecanum wheels; Robotic arm: 2R tandem robotic arm, single arm length ≈18cm, equipped with Dynamixel XM430-W210-T torque control motor, and end effector with gripper; Overall parameters: Total weight ≈ 5.0kg, robotic arm degrees of freedom n=2, generalized coordinate dimensions 6+n=8; Control system: Embedded motion controller (such as STM32H7 series), running a real-time operating system, supporting MicroXRCE-DDS communication; Airborne computing: Jetson Orin Nano Super airborne computer enables sensor data acquisition and low-latency communication; Edge computing: A host equipped with an NVIDIA RTX 4080 GPU to perform model inference and control law calculation; Perception system: OptiTrack motion capture system (120fps) + onboard IMU (integrating robot chassis status), Dynamixel encoder (collecting robotic arm joint position / velocity); Communication link: The airborne computer communicates with the terminal host via WiFi. The uplink frequency of sensor data and the downlink frequency of control commands are 50Hz to meet the real-time control requirements. Power supply system: A high-capacity lithium battery (24V, 20Ah) provides independent power to the chassis motor and robotic arm, ensuring long-term operation.

[0166] 1.2 Software and Simulation Environment Configuration Operating system: Ubuntu 20.04, ROS2 Humble (robotic arm and chassis control); Deep learning framework: JAX (with GPU acceleration), used for training and inference of feature extraction and diffusion models; Control framework: A custom C++ / Python control library is used to realize real-time calculation of adaptive control laws, which is compatible with PX4's uORB topic and ROS2's Joint Trajectory Controller; Data processing: Pandas / Numpy is used for data acquisition and preprocessing, and Matplotlib / Seaborn is used for result analysis and visualization; Simulation platform (pre-training): Gazebo / Isaac Gym, used to generate simulation data and pre-train models.

[0167] 2. Model Network Architecture and Training Settings 2.1. Feature Extraction Model (TCN) Architecture Shallow temporal convolutional networks are used as the feature extraction model. The input is the state-input sequence of the most recent L=10 steps (dimension 10×(8+8)=160, 8-dimensional state + 8-dimensional control input), and the output is a 32-dimensional state descriptor. Specific architecture: Table 1 TCN Architecture

[0168] 2.2. Denoising Network Architecture for Diffusion Model (Temporal U-Net) Based on the DroneDiffusion architecture, a temporal U-Net is used as the denoising network, fusing residual temporal convolutional blocks, diffusion step embedding, and conditional embedding (state / input / state). Specific design details are as follows: Input: Noisy residual (8-dimensional) + state (8-dimensional) + control input (8-dimensional) + state descriptor (32-dimensional) + diffusion step embedding (16-dimensional), concatenated to 72 dimensions; Core module: 4-layer residual temporal convolutional blocks, each layer containing 1D convolution, normalization, and SiLU activation, to achieve multi-scale feature extraction; Diffusion step embedding: The diffusion step k is converted into a 16-dimensional vector through sinusoidal position encoding, incorporating temporal features; Conditional embedding: The state, input, and state descriptor are converted into a 64-dimensional embedding vector through an MLP and then fused with convolutional features; Output: Predicted Gaussian noise (8-dimensional), consistent with the dimensions of the residual vector.

[0169] 2.3. Core Training Hyperparameters The model's training hyperparameters were optimized through multiple rounds of iterations, balancing convergence speed, prediction accuracy, and generalization ability. The specific settings are shown in Table 2. Table 2 Core Hyperparameter Settings for Model Training

[0170] 3. Adaptive Controller Parameter Settings The gain parameters of the adaptive controller are optimized for the dynamic characteristics of the end effector, balancing tracking accuracy and control stability. The specific settings are shown in Table 3 (all matrices are diagonal matrices): Table 3 Adaptive Controller Core Parameter Settings

[0171] 4. Experimental Scenario and Evaluation Indicators 4.1 Experimental Scenario Design The core experimental scenario uses a payload grabbing-transporting-releasing task with an ∞-shaped trajectory to simulate actual ground movement operations. Specific scenario parameters are as follows: Movement trajectory: XY plane ∞ shape trajectory, movement height is the ground height (z=0), including five stages: start-up, load grabbing, transportation, release, and return to origin; Key waypoints: origin (0.0, 0.0), grab point (1.0, 1.0), release point (-1.0, -1.0); Robotic arm configuration: start / return phase (retraction), grasping phase, transport / release phase; Test conditions: 2 load masses (300g / 500g, which are 6% / 10% of the total weight of the machine, respectively) and 2 moving speeds (0.5m / s / 1.0m / s), covering low-speed stable and high-speed dynamic conditions; Comparison methods: Classical methods (SysID / ASMC), data-driven baseline methods (DNN / GP / GPT-2 / traditional diffusion models), all comparison methods use the same training dataset and controller parameters.

[0172] 4.2 Core Evaluation Indicators Two metrics, open-loop prediction accuracy and closed-loop tracking accuracy, were used to quantitatively evaluate the performance of the model and control framework. All results are the average of 10 independent experiments. Open-loop prediction metric: Root mean square error (RMSE) of the residual vector. The RMSE of the three dimensions of position (H1-H3), attitude (H4-H6), and robotic arm joints (H7-H8) are calculated separately. The lower the RMSE, the higher the prediction accuracy. Closed-loop tracking metric: Track tracking RMSE (unit: meters) in generalized coordinates. The tracking RMSE under different load-velocity combinations is calculated separately. The lower the RMSE, the higher the control accuracy. Stability metrics: size of the uniform final bounded set of the closed-loop system, convergence time of the tracking error, used to evaluate the stability and dynamic response characteristics of the control.

[0173] 4. Experimental Results and Analysis 5.1 Open-loop residual prediction results Table 4 shows the residual prediction RMSE of different methods. The state condition diffusion model of the present invention (Proposed) achieves the lowest RMSE in all dimensions, significantly outperforming the comparison methods. Table 4 Comparison of Residual Prediction RMSE

[0174] 5. Results Analysis: The classic SysID method relies on a fixed parameterized model, which results in the largest prediction error and cannot adapt to non-stationary changes in dynamics. While traditional data-driven methods such as DNN / GP / GPT-2 reduce errors, their prediction accuracy is limited by deterministic mapping / smoothing assumptions / error accumulation. Traditional diffusion models improve accuracy by learning the distribution, but due to the lack of state conditions, they cannot distinguish the residual characteristics of different states, and the error is still higher than that of this invention. This invention achieves state-aware distribution prediction through state descriptor encoding, effectively avoiding prediction aliasing and improving prediction accuracy by approximately 30% / 15% / 25% in position / attitude / robotic arm dimensions, respectively.

[0175] 5.2 Closed-loop trajectory tracking results Table 5 shows the trajectory tracking RMSE of different methods. The framework of this invention achieves the lowest tracking RMSE under all test conditions and maintains excellent performance even under harsh conditions of high load / high speed. Table 5. Comparison of RMSE for trajectory tracking (unit: meters)

[0176] Results analysis: As a classic adaptive control method, ASMC lacks data-driven residual compensation, resulting in the largest tracking error, which increases sharply with increasing load / velocity. Traditional data-driven methods combined with adaptive control improve tracking accuracy, but performance degrades significantly under high load / high speed due to limitations in residual prediction accuracy. The residual compensation of the traditional diffusion model improves the tracking accuracy, but the state aliasing causes it to be untimely in the event of sudden load changes / configuration switching, resulting in a higher error than that of the present invention. The state-condition diffusion model of this invention provides high-precision residual prediction. Combined with the error compensation of the adaptive controller, under the out-of-range conditions of 500g / 1.0m / s, the tracking accuracy is still about 18% higher than the traditional diffusion model and about 52% higher than ASMC, demonstrating excellent robustness.

[0177] 5.3 Generalization ability and stability results Out-of-distribution generalization: Under a 500g load condition outside the training set, the predicted RMSE of this invention only increased by 0.005 to 0.01, and the tracking RMSE increased by 0.03 to 0.04, which is much lower than the comparison method (the traditional diffusion model's predicted RMSE increased by 0.01 to 0.02, and the tracking RMSE increased by 0.05 to 0.06), demonstrating strong out-of-distribution adaptation capability; Closed-loop stability: The closed-loop system of this invention has a consistent final bounded set size of 0.05m and a convergence time of 0.2s for tracking error. Compared with the traditional diffusion model (bounded set 0.08m, convergence time 0.3s), it has better stability and dynamic response characteristics, and there was no instability in 10 independent experiments, which significantly improves reliability.

[0178] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0179] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement any of the methods in the foregoing embodiments of this disclosure.

[0180] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media may include, for example, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), compact disc-read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any combination thereof. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0181] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include one or more of copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media in the respective computing / processing device.

[0182] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object programs written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as the "C" language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network (e.g., a local area network or a wide area network), or it may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays, or programmable logic arrays, can execute computer-readable program instructions to implement various aspects of the embodiments of this disclosure by utilizing state information from the computer-readable program instructions.

[0183] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0184] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0185] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0186] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It should be noted that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are all equivalent.

[0187] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.

Claims

1. A control method, characterized in that, include: Obtain the sequence of state control torque pairs of the end effector of the mobile device within the most recent preset window length; Based on the state control torque pair sequence, condition information is determined; wherein, the condition information includes the current operating state information and current state information of the end effector, as well as the control torque at the previous moment; By generating a model guided by the conditional information, the conditional distribution of the comprehensive residual vector of the end effector is modeled, and the predicted residual value of the end effector is determined by inverse denoising. The current control torque of the end effector is determined based on the current reference state information, current state information, predicted residual value, and adaptive gain of the previous moment. The end effector is moved according to the current control torque; Repeat the above steps until the end effector reaches the preset endpoint position and completes the preset operation.

2. The method according to claim 1, characterized in that, The step of determining the current control torque of the end effector based on the current reference state information, current state information, predicted residual value, and adaptive gain from the previous moment includes: Based on the current state information and the reference state information at the current moment, the tracking error and error variables are determined; wherein, the reference state information at the current moment includes the desired position, desired velocity, and desired acceleration; The adaptive gain at the previous time step is updated based on the error variable to obtain the adaptive gain at the current time step. The current control torque is determined based on the nominal inertia matrix of the end effector, the desired acceleration, the error variable, the predicted residual value, and the adaptive gain at the current moment.

3. The method according to claim 2, characterized in that, Before obtaining the sequence of state control torque pairs of the end effector of the mobile device at the most recent preset window length, the method further includes: Initialize the control torque and adaptive gain of the end effector, and clear the history buffer. Obtain the current state information of the end effector, and store the current state information and the currently used control torque as a state control torque pair in the history buffer; Determine whether the number of state control torque pairs stored in the historical buffer reaches the preset window length; If not, the predicted residual value of the end effector is set to zero, and the current control torque of the end effector is determined based on the current state information and the reference state information at the current moment. In this case, the step of obtaining the state control torque pair sequence of the end effector of the mobile device at the most recent preset window length is performed.

4. A training method, characterized in that, include: Obtain a training sample set; where each training sample includes a sequence of sample state control torque pairs and sample residual values; Based on the sample state control torque pair sequence, sample condition information is determined; wherein, the condition information includes sample operation state information, sample state information, and sample control torque; Noise is added to the sample residuals of the training samples to obtain a dataset with added noise; The sample condition information corresponding to the training samples and the dataset with added noise are input into the generative model, so that the generative model predicts the noise added to the dataset based on the sample condition information, and obtains the predicted noise value. A loss function is constructed using the loss of the predicted noise value relative to the true noise value, and the model parameters of the generated model are updated.

5. The method according to claim 4, characterized in that, The acquisition of the training sample set includes: Obtain the dataset; the dataset includes a sequence of state control torque pairs of the end effector under different operating conditions; The dataset is divided into multiple samples according to a preset window length; each sample includes a sequence of sample state control torque pairs. For any sample, the sample residual value is determined based on the sample state control torque pair sequence and the preset comprehensive residual vector calculation formula. The sample state control torque pair sequence and sample residual value of a sample are used as a training sample to obtain the training sample set corresponding to the dataset.

6. The method according to claim 4, characterized in that, The step of determining sample condition information based on the sample state control torque pair sequence includes: The sample state control torque pair sequence of the sample is input into the feature extraction model to obtain the sample operation state information of the sample; The step of constructing a loss function based on the loss between the predicted noise value and the true noise value, and updating the model parameters of the generated model, includes: A loss function is constructed using the loss of the predicted noise value relative to the true noise value, and the model parameters of the generative model and the parameters of the feature extraction model are updated.

7. The method according to claim 5, characterized in that, The preset formula for calculating the comprehensive residual vector is determined through the following steps: The original dynamic model of the end effector is established based on the Euler-Lagrange equations. The original dynamic model includes inertial force terms, Coriolis force and centripetal force terms, gravity terms, and aerodynamic disturbance terms. Based on the nominal inertia matrix, the inertial force term in the original dynamic model is decomposed into a nominal inertia term and a comprehensive residual term; Based on the definition of the comprehensive residual term, the non-nominal inertial force, Coriolis force, centripetal force, gravity, and aerodynamic disturbance in the original dynamic model are integrated into the comprehensive residual vector, and the preset comprehensive residual vector calculation formula is obtained.

8. A control device, characterized in that, It includes a memory and a processor, the memory being used to store executable instructions; the processor being used to operate under the control of the instructions to perform the method as described in any one of claims 1 to 7.

9. A mobile device, characterized in that, It includes a data acquisition device and a control device according to claim 8.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.