A method, system, and apparatus for coordinated execution of robots
Patent Information
- Application Number
- CN202610844130.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]现有技术直接将VLA模型输出的连续动作块直接用于控制执行器,其动作频率、动作幅值变化以及动作帧分布未必与真实机器人硬件的控制接口和执行器动态特性相匹配
[0019]这样通过保持第一动作子序列不变,将第二动作子序列整体后移偏移量的帧位置,直接实现两种执行器在物理时间上的同步到达,根本性解决夹爪超前问题,显著提升操作稳定性与任务成功率。
Smart Images

Figure CN122829818A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of machine learning and robot control, and in particular to a method, system, and device for cooperative execution of robots. Background Technology
[0002] When robots perform complex tasks, they typically require coordinated movement between first-type actuators (such as the joints of a robotic arm) and second-type actuators (such as an end effector) to complete actions such as grasping, assembly, and manipulation. With the development of Vision-Language-Action (VLA) models, these models can directly output action blocks containing multiple frames of continuous motion. Each frame includes both joint motion commands and gripper motion commands, with the two types of commands perfectly aligned in time. However, in real-world robot systems, first-type actuators are characterized by high inertia, low control bandwidth, and high precision, resulting in a slower dynamic response speed; while second-type actuators are characterized by low inertia, high control bandwidth, and low precision, resulting in a faster response speed. There is an order-of-magnitude difference in their physical response characteristics. Therefore, enabling the ideal synchronized action blocks output by the VLA model to achieve task-level coordinated motion on real robots with heterogeneous actuators has become a key issue in improving the stability and success rate of robot operations.
[0003] Existing technologies directly use the continuous action blocks output by the VLA model to control the actuator. However, the action frequency, amplitude variations, and action frame distribution may not match the control interface and dynamic characteristics of the actual robot hardware. Furthermore, if the original action blocks output by the model are directly sent to the robot controller frame by frame, it may increase the control communication load and execution latency due to the large number of action frames and frequent sending of small actions. On the other hand, key action frames in the action blocks are usually implicitly generated by the model. If these key action frames are not identified and retained, critical actions such as grasping and releasing may be lost, or even abnormal grasping timing may occur. Moreover, the dynamic characteristics of the gripper and the first type of actuator also differ to some extent, amplifying the risk of timing inconsistencies. Summary of the Invention
[0004] This invention provides a method, system, and device for collaborative execution of robots, which can improve the timing matching between the robot motion model output and the collaborative execution between heterogeneous actuators, thereby improving the stability of robot operation.
[0005] The first aspect of this invention provides a method for cooperative execution of robots, comprising: Receive the original motion block output by the robot motion model, wherein the original motion block contains motion components corresponding to at least two types of actuators; The original action block is decoupled into at least two types of action sub-sequences, wherein each type of action sub-sequence corresponds to one type of executor; The correction compensation amount is calculated in real time based on the robot's current task operating parameters, and the dynamic timing offset between various actuators is determined by using the preset dynamic response benchmark deviation and the correction compensation amount. According to the dynamic temporal offset, the at least two types of action subsequences are relatively shifted on the time axis to generate asynchronously aligned composite action blocks; The composite action block is sent to the robot's controller to drive various actuators to perform coordinated movements.
[0006] This invention decouples the original action block into action sub-sequences corresponding to each actuator, laying the foundation for subsequent differentiated adjustments. Instead of using fixed delays or simple filtering, it collects current task parameters in real time and dynamically calculates a timing offset adapted to the current load, speed, and operating object, based on a pre-calibrated dynamic response benchmark deviation. This allows the offset to change with the operating conditions, avoiding the mismatch risk of static compensation. The action sub-sequences are then relatively shifted along the time axis according to this offset, generating asynchronously aligned composite action blocks. This "asynchronous alignment" breaks the original synchronous assumption of the model, ensuring that after being sent to the controller, the first and second types of actuators can reach their respective key action states synchronously in physical time, fundamentally solving the timing mismatch problem between the model output and heterogeneous actuators. Finally, by sending the composite action blocks as a whole, communication overhead is reduced, and the smooth, conflict-free movement of various actuators during task execution is ensured, significantly improving the stability of robot operation and the success rate of the task.
[0007] Furthermore, the at least two types of actuators include a first type of actuator and a second type of actuator. The step of calculating the correction compensation amount in real time based on the robot's current task operating parameters, and determining the dynamic timing offset between the various types of actuators using a preset dynamic response benchmark deviation and the correction compensation amount, includes: Obtain the pre-calibrated dynamic response benchmark deviation between the first type of actuator and the second type of actuator; The robot's current task operating condition parameters are collected in real time, wherein the current task operating condition parameters include at least the load inertia and end effector speed of the first type of actuator and the second type of actuator; The correction compensation amount is calculated based on the load inertia and the end velocity. The dynamic timing offset between the first type of actuator and the second type of actuator is obtained based on the dynamic response reference deviation and the correction compensation amount.
[0008] By pre-calibrating the baseline deviation and dynamically correcting the offset by collecting working parameters in real time, the timing compensation can accurately adapt to changes in the task, effectively offsetting the differences in joint and gripper response, and improving the coordination and operational stability.
[0009] Furthermore, the at least two types of actuators include a first type of actuator and a second type of actuator, and the decoupling of the original action block into at least two types of action sub-sequences includes: The motion components corresponding to at least one type of actuator in the original motion block are preprocessed to obtain the target motion block; According to the preset dimension division rules of each action frame in the target action block, the action components belonging to the same executor type are extracted to form the first action subsequence corresponding to the first type of executor and the second action subsequence corresponding to the second type of executor.
[0010] By removing motion noise and redundancy and retaining keyframes, this provides high-quality subsequences for timing alignment, reduces communication latency, and improves execution smoothness and collaborative stability.
[0011] Further, the preprocessing of the motion components corresponding to at least one type of actuator in the original motion block to obtain the target motion block includes: The motion components corresponding to at least one type of actuator in the original motion block are nonlinearly mapped to obtain the first motion block; The first action block is subjected to a moving average filter to obtain the second action block; Redundant action frames of the second action block are removed to obtain the third action block; Adaptive keyframe downsampling is performed on the third action block to obtain the target action block.
[0012] Further, the step of performing a nonlinear mapping on the action components corresponding to at least one type of actuator in the original action block to obtain the first action block includes: Obtain the first action component corresponding to the second type of executor in the original action block; Based on the attribute type of the current operation object, the current task stage, and the interaction state between the second type of executor and the current operation object, the steepness parameter and closure threshold parameter of the nonlinear mapping curve are dynamically selected. The nonlinear mapping curve is constructed using the steepness parameter and the closure threshold parameter. The first motion component corresponding to the second type of actuator is numerically transformed through the nonlinear mapping curve to obtain the transformation result. The transformation result is then subjected to amplitude limiting processing to obtain the adjusted gripper motion component. The gripper motion component is merged with the second motion component corresponding to other unadjusted actuators in the original motion block to form the first motion block.
[0013] This dynamically adjusts the gripper mapping curve based on the object's hardness, fragility, task stage, and gripping interaction state, avoiding sudden changes in the gripper or unstable gripping, making the timing of gripper movements and joint movements more coordinated, and improving the success rate of grasping and operational stability.
[0014] Further, the step of removing redundant action frames from the second action block to obtain the third action block includes: Obtain the third motion component in each dimension for each motion frame in the second action block; For any two adjacent action frames in the second action block, calculate the change of the third action component in each dimension of the latter action frame relative to the former action frame. If the absolute value of the change in each dimension is less than the change threshold corresponding to that dimension, then the next action frame is determined to be a redundant action frame and the redundant action frame is removed from the second action block until all non-redundant action frames are retained, and all the non-redundant action frames are arranged into the third action block in the original time order.
[0015] By judging and eliminating action frames with minor changes based on thresholds of various dimensions, redundant frames in action blocks are reduced, controller communication load and delivery latency are lowered, and composite action blocks after timing alignment are made more concise, improving execution efficiency and collaborative response speed.
[0016] Further, the adaptive keyframe downsampling of the third action block to obtain the target action block includes: Obtain the joint motion components of each motion frame in the third motion block, and calculate the rate of change of joint angular velocity between adjacent motion frames based on the joint motion components. The joint angular velocity change rate of each action frame is compared with a preset motion change threshold. If the joint angular velocity change rate is greater than the motion change threshold, the action frame is determined to be a key action frame and is retained, resulting in several first retained frames. The last action frame in the third action block is used as the second reserved frame; The first reserved frame and the second reserved frame are combined in their original time order to form the target action block.
[0017] This approach, which adaptively preserves key motion frames based on the rate of change of joint angular velocity, avoids losing important poses, and reduces non-key frames, ensures that the joints and grippers are precisely synchronized at key points after the time shift, and balances execution efficiency and fine operation capabilities.
[0018] Further, the at least two types of actuators include a first type of actuator and a second type of actuator, and the at least two types of action sub-sequences include a first action sub-sequence and a second action sub-sequence. The step of relative shifting the at least two types of action sub-sequences on the time axis according to the dynamic timing offset to generate asynchronously aligned composite action blocks includes: Keep the position of the first action subsequence corresponding to the first type of actuator unchanged on the time axis, and shift the frame position of the second action subsequence corresponding to the second type of actuator in the positive direction of the time axis by the frame position of the dynamic timing offset to obtain the third action subsequence after the second type of actuator is shifted; The third action subsequence is merged with the unshifted first action subsequence to generate the composite action block.
[0019] By keeping the first action subsequence unchanged and shifting the second action subsequence backward by an offset frame position, the two actuators can be synchronized in physical time, fundamentally solving the gripper advance problem and significantly improving operational stability and task success rate.
[0020] Another embodiment of the present invention provides a cooperative execution system for robots, comprising: the cooperative execution system being used to perform the steps of the cooperative execution method for robots as described in the present invention.
[0021] Another embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the steps of the robot cooperative execution method as described in the present invention. Attached Figure Description
[0022] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is an example diagram of a dual-arm robot service scenario provided in this application; Figure 2 This is a flowchart illustrating one embodiment of the robot collaborative execution method provided in this application; Figure 3 This is a numerical example diagram of the action block provided in this application; Figure 4 This is a flowchart illustrating steps S401 to S402 provided in this application; Figure 5This is a flowchart illustrating steps S501 to S504 provided in this application; Figure 6 Here are example diagrams of the gripper mapping curves for the three types of objects provided in this application; Figure 7 This is an example diagram of the gripper mapping curve for the enhanced closure mode provided in this application; Figure 8 This is an example diagram comparing the effects of the moving average filtering process before and after the processing provided in this application; Figure 9 This is an example diagram of the redundant action frame removal results provided in this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0026] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0029] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0030] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0031] This invention is applicable to scenarios where dual-arm service robots need to perform complex operational tasks in unstructured environments, such as... Figure 1 The diagram illustrates a dual-arm robot service scenario, showcasing typical applications of this solution, such as dual-arm robot operation in home service, catering service, or medical assistance. The left and right arms perform different tasks (e.g., holding objects and grasping / manipulating), highlighting the need for coordinated joint and gripper work. Examples include home service, catering service, medical assistance, and commercial cleaning. In these scenarios, the robot needs to generate and execute continuous action blocks, including both arm joints and grippers, in real time based on visual and verbal instructions (e.g., "pick up the fragile glass on the table" or "use the left arm to hold the box and the right arm to open the lid"). Because these scenarios often face challenges such as variable object properties (hard, soft, fragile), frequent transitions between operation stages (approach, gripping, moving, releasing), and high requirements for dual-arm coordination, directly using the raw actions output by the VLA model can easily lead to problems such as gripper abrupt changes, joint tremors, loss of key actions, or timing discrepancies between the two arms. By adopting the post-processing and efficient execution methods of this solution, the joint and gripper movements can be decoupled, smoothed, downsampled, and dynamically time-compensated, significantly improving the smoothness, stability, and success rate of the action execution. It is especially suitable for service robot applications that have high requirements for grasping stability, object safety, and execution efficiency.
[0032] Beyond the aforementioned dual-arm service robot scenarios, this solution can also be extended to a wider range of robot platforms, including single-arm or omnidirectional mobile manipulation robots, industrial assembly robots, and humanoid robots. For example, in precision electronic assembly or parts insertion scenarios, the adaptive keyframe downsampling mechanism of this solution can retain key assembly actions and eliminate minor redundant movements, thereby improving assembly accuracy and cycle time efficiency. In logistics sorting or warehousing and handling scenarios, when faced with packages of different hardness and shape, the gripper dynamic adjustment mechanism can optimize the gripping force in real time according to the object type, reducing the risk of damage to goods. In elderly care or rehabilitation assistance scenarios, smooth motion trajectories and timing compensation mechanisms can ensure that the robot interacts with the human body in a gentler and safer manner, avoiding discomfort or injury caused by sudden gripper movements. In addition, this solution can also be combined with teleoperation or teaching learning systems to post-process offline collected expert trajectories, achieving efficient motion replay and generalized deployment. Essentially, any robotic system that needs to convert discrete motion blocks output by VLA models or similar motion generation models into smooth, stable, and efficient execution by real physical actuators (joints, grippers, dexterous hands, etc.) can benefit from this solution.
[0033] See Figure 2 To improve the timing matching between the robot motion model output and the collaborative execution of heterogeneous actuators, an embodiment of the present invention provides a collaborative execution method for robots, including steps S201 to S205: Step S201: Receive the original motion block output by the robot motion model, wherein the original motion block contains motion components corresponding to at least two types of actuators; In some embodiments, motion blocks generated iteratively by the diffusion module in the vision-language-action model (VLA model) are received, which are the original motion blocks output by the robot motion model.
[0034] It should be noted that the original motion block contains motion components of at least two types of actuators, including but not limited to: motion components of joint actuators (7-dimensional left arm and 7-dimensional right arm, totaling 14-dimensional arc values), and motion components of gripper actuators (1-dimensional left arm gripper and 1-dimensional right arm gripper, totaling 2-dimensional opening and closing values).
[0035] Those skilled in the art will readily understand that the structure of this action block can be extended to other types of actuators depending on different robot configurations, such as the joints and grippers of a single-arm robot, or the action dimensions of other types of actuators (such as the waist, head, etc.). Therefore, the received original action block is essentially a multi-moment action sequence that organizes the action instructions of different actuators at the same moment in a preset dimensional order, providing basic data for subsequent decoupling and actuator-specific processing.
[0036] It should be noted that the steps of the diffusion module in the vision-language-action model (VLA model) to iteratively generate action blocks are not the focus of this application, and therefore will not be elaborated here.
[0037] It should be noted that the numerical diagram of the action block is as follows: Figure 3 As shown in the figure, a 50-row × 16-column matrix example is illustrated. Each row represents the action of a time frame, where 50 represents the number of actions contained in the action block (i.e., 50 time-step action frames), 16 represents the action dimensions of each action frame, and the 16 columns correspond to the values of the 7 joints of the left arm, the left arm gripper, the 7 joints of the right arm, and the right arm gripper. The specific composition of these 16 action dimensions is as follows: the 7 joints of the left arm (each corresponding to a motor rotation radian value), the left arm gripper (corresponding to a gripper opening / closing degree value), the 7 joints of the right arm (each corresponding to a motor rotation radian value), and the right arm gripper (corresponding to a gripper opening / closing degree value). This example provides a visual understanding of the data structure of the original action block output by the model.
[0038] Step S202: Decouple the original action block into at least two types of action sub-sequences, wherein each type of action sub-sequence corresponds to one type of executor; Please refer to Figure 4 In some embodiments, the at least two types of actuators include a first type of actuator and a second type of actuator, and step S202 includes steps S401 to S402: Step S401: Preprocess the motion components corresponding to at least one type of actuator in the original motion block to obtain the target motion block; Step S402: According to the preset dimension division rules of each action frame in the target action block, the action components belonging to the same actuator type are extracted to form a first action subsequence corresponding to the first type of actuator and a second action subsequence corresponding to the second type of actuator.
[0039] By removing motion noise and redundancy and retaining keyframes, this provides high-quality subsequences for timing alignment, reduces communication latency, and improves execution smoothness and collaborative stability.
[0040] Please refer to Figure 5 In some embodiments, step S401 includes steps S501 to S504: Step S501: Perform nonlinear mapping on the motion components corresponding to at least one type of actuator in the original action block to obtain the first action block; Step S502: Perform a moving average filter on the first action block to obtain the second action block; Step S503: Remove redundant action frames from the second action block to obtain the third action block; Step S504: Adaptive keyframe downsampling is performed on the third action block to obtain the target action block.
[0041] In some embodiments, step S501 includes: obtaining the first motion component corresponding to the second type of actuator in the original action block; dynamically selecting the steepness parameter and closure threshold parameter of the nonlinear mapping curve according to the attribute type of the current operation object, the current task stage, and the interaction state between the second type of actuator and the current operation object; constructing the nonlinear mapping curve using the steepness parameter and the closure threshold parameter; numerically transforming the first motion component corresponding to the second type of actuator through the nonlinear mapping curve to obtain the transformation result; and performing amplitude limiting processing on the transformation result to obtain the adjusted gripper motion component; merging the gripper motion component with the second motion components corresponding to other unadjusted actuators in the original action block to form the first action block. Specifically, firstly, the motion component corresponding to the gripper in the VLA model action block is obtained, and the actual maximum opening parameter MAXG of the gripper is set (e.g., a value of 70). Then, according to the attribute type of the current operation object, the task stage, and the interaction state between the gripper and the object, the steepness parameter α and the closure threshold parameter k of a Sigmoid mapping curve are dynamically adjusted. In terms of the task phase, before gripping the object, k can be set to 0.1 and α=10 to suppress noise; switch to gripping mode after contact with the object. After selecting the parameters, input the gripper values using the following formula. The transformation and calculation formula is as follows: ,in, This represents the actual maximum opening of the gripper (70 in this scheme). For nonlinear smoothing functions, For steepness parameters, The closing threshold, This is the input of the original gripper motion value. This allows us to obtain the transformed gripper components and the components of other actuators (such as joints) in the original motion block. These components are then combined to obtain the first motion block.
[0042] It should be noted that the object type can be obtained through the visual question-answering capability of the VLA model. This solution considers three types of objects: hard, soft, and fragile. The object attributes can be obtained by asking the VLA model with a prompt: "Based on the task {task} and the current visual observation {image}, determine whether the object to be operated on is hard, soft, or fragile." After obtaining the model's response, the type string can be extracted using a simple regular expression matching formula "(hard|soft|fragile)".
[0043] It should be noted that for hard objects, the grippers should close as early as possible to prevent the object from falling; therefore, a slightly larger gripper can be used. Values with larger ones Values (e.g.) For soft objects, the gripper closure degree should be moderate to avoid excessive deformation of the object; a moderate gripper can be selected. Value and moderate Values (e.g.) For fragile objects, the gripper dynamics should be as conservative as possible; a moderate gripper speed can be selected. Values with smaller Values (e.g.) If visual detection detects object slippage, the k value is gradually increased in increments of 0.05, up to a maximum of approximately 0.7, to provide greater clamping force. Example diagrams of the gripper mapping curves for three types of objects (hard, soft, and fragile) are shown below. Figure 6 As shown in the figure, the Sigmoid mapping curves are compared under three different α and k parameters. The curve is steepest for hard objects and flattest for fragile objects. This illustration demonstrates how the gripper opening and closing values are non-linearly adjusted according to the object's properties.
[0044] It should be noted that, during the task phase, the grippers should remain open and non-shaking before grasping the object; therefore, the closure threshold can be set relatively low to avoid noise (e.g., Once the force sensor or vision detector determines that the object is in contact, the curve is adjusted to clamping mode.
[0045] This dynamically adjusts the gripper mapping curve based on the object's hardness, fragility, task stage, and gripping interaction state, avoiding sudden changes in the gripper or unstable gripping, making the timing of gripper movements and joint movements more coordinated, and improving the success rate of grasping and operational stability.
[0046] Figure 7 This application provides an enhanced closure mode ( The example diagram shows the gripper mapping curve. The curve is close to a step shape, indicating that the gripper quickly saturates to its maximum opening once the input value exceeds the closing threshold. This mode provides greater gripping force when the object slips. Regarding gripper-object interaction, if the visual detector determines that the object is not firmly gripped and has slipped, the gripper force is gradually increased in increments of 0.05. value. The closer the curve is to the reference value of 0.7, the closer it is to a step curve, and the more the gripper tends to be in a fully closed state, which can provide greater clamping force to prevent slippage.
[0047] This dynamic adjustment process effectively avoids problems such as sudden changes in the gripper or unstable gripping caused by the mismatch between training data and the inference object.
[0048] In some embodiments, step S502 includes: after obtaining the first action block, to suppress possible motion jitter and abrupt changes in the VLA model prediction, especially numerical jumps in the gripper component, a moving average filtering process needs to be performed. Specifically, a queue storing historical action frames can be maintained, with a window size denoted as W, generally based on 50. The window size can be fine-tuned according to the specific scenario and task (increasing the window makes the action smoother but increases execution latency, while decreasing the window has the opposite effect). For each time step t in the action block, the action values of the current time and the previous W-1 times are taken (if less than W frames, the actual number of frames is used), and the arithmetic mean is calculated for each dimension to serve as the smoothed output action. After processing all time steps, the second action block is obtained. Enlarging the window can make the action smoother, but it also risks increasing execution latency. Therefore, the window size needs to be fine-tuned according to the scenario and task.
[0049] It should be noted that, Figure 8 This is an example of a comparison between the effects of the moving average filtering process provided in this application. Taking the gripper motion value as an example, the figure shows that the original curve has sharp peaks and jumps, while the curve becomes continuous and smooth after the moving average filtering. This figure illustrates the role of smoothing in suppressing motion jitter and enhancing execution smoothness.
[0050] In some embodiments, step S503 includes: obtaining the third action component of each action frame in the second action block in each dimension; for any two adjacent action frames in the second action block, calculating the change of the third action component of the latter action frame relative to the former action frame in each dimension; if the absolute value of the change is less than the change threshold corresponding to that dimension in each dimension, then the latter action frame is determined to be a redundant action frame and the redundant action frame is removed from the second action block until all non-redundant action frames are retained, and all the non-redundant action frames of the second action block are combined into the third action block according to their original time order. Specifically, after obtaining the second action block... Next, the motion components of each motion frame in the second motion block are obtained in each dimension. In this scheme, a motion frame has 16 dimensions, corresponding to the 7 joints of the left arm, the left arm gripper, the 7 joints of the right arm, and the right arm gripper. For any two adjacent motion frames, the motion components of the next frame are calculated. Compared to the previous frame The absolute value of the change in each dimension j. If, in each dimension, the absolute value of this change is less than a pre-defined threshold for that dimension. , that is If the condition is met, the i-th frame is discarded. Repeat the process of traversing the entire action block until all non-redundant action frames that do not meet the condition are retained. These frames are then arranged in the original time order to form the third action block.
[0051] It should be noted that the schematic diagram of the adaptive frame filtering result provided by the present invention is as follows: Figure 9 As shown in the figure, the shaded areas indicate small motion segments where the change between adjacent frames is less than a threshold and can be removed. This figure visually illustrates how to identify and filter redundant frames by using a dimensionality change threshold, thereby reducing the number of actions sent.
[0052] It should be noted that the thresholds for each dimension... The empirical threshold vector (arranged in 16-dimensional order) can be obtained by multiplying the standard deviation of the statistical training samples by a certain factor. Among them, the 8th and 16th dimensions correspond to the left and right grippers, with relatively large thresholds (0.835 and 0.903), because the gripper opening and closing value range is different from the joint curvature.
[0053] By judging and eliminating action frames with minor changes based on thresholds of various dimensions, redundant frames in action blocks are reduced, controller communication load and delivery latency are lowered, and composite action blocks after timing alignment are made more concise, improving execution efficiency and collaborative response speed.
[0054] In some embodiments, step S504 includes: acquiring the joint motion components of each motion frame in the third motion block, and calculating the joint angular velocity change rate between adjacent motion frames based on the joint motion components; comparing the joint angular velocity change rate of each motion frame with a preset motion change threshold; if the joint angular velocity change rate is greater than the motion change threshold, determining that the motion frame is a key motion frame and retaining it, thus obtaining several first retained frames; using the last motion frame in the third motion block as a second retained frame; and combining each of the first retained frames and the second retained frames in their original time order to form the target motion block. Specifically, firstly, the joint motion components of each motion frame in the third motion block are acquired, and the joint angular velocity change rate between adjacent motion frames is calculated based on these components. Let the joint angular velocity of the t-th frame be denoted as . The joint angular velocity in frame t-1 is Then the rate of change of motion The calculation formula is: ,in It is a very small positive number. Then, the rate of change of motion state corresponding to each action frame is... Compared with the preset motion change threshold (For example, 0.10) for comparison. If Greater than the threshold If a significant change in joint velocity is observed, the resulting action frame is identified as a critical action frame and retained, resulting in several first-retained frames. Simultaneously, to ensure trajectory integrity, the last action frame in the third action block is designated as the second-retained frame and must be retained. The target action block after redundancy removal and downsampling is... .
[0055] It should be noted that the required level of precision varies at different stages of the task. During the end-effector movement, the precision required can be increased. This reduces the number of actions required and improves movement efficiency; during object manipulation, it can reduce... To ensure the operation is completed, the number of keyframes is increased. It is worth noting that, to ensure the integrity of the trajectory, the action block will retain at least one frame of the action, and the last frame must be retained.
[0056] It should be noted that the joint angular velocity is It can be calculated by dividing the joint position values of adjacent frames by the time interval.
[0057] It should be noted that during the execution of an action block, only a small number of actions are often critical. The process of the robot moving to a certain action can be achieved through low-level motion control, without needing to execute each action block completely. Based on this, this solution employs an adaptive downsampling strategy to reduce the number of actions issued, greatly alleviating the slow execution issue. For scenarios where the precision of the actions is not critical, this approach can be used... The method of uniform downsampling, that is, each Only one of each action is retained, reducing the number of actions to the original value. The advantage of uniform downsampling is that it is simple and efficient to process, but it also has the potential to miss critical actions in the action block.
[0058] This approach, which adaptively preserves key motion frames based on the rate of change of joint angular velocity, avoids losing important poses, and reduces non-key frames, ensures that the joints and grippers are precisely synchronized at key points after the time shift, and balances execution efficiency and fine operation capabilities.
[0059] By preprocessing the motion components, redundant frames are obtained, improving the accuracy and efficiency of subsequent processing.
[0060] In some embodiments, step S402 includes: for the target action block Extract the 14 dimensions (indices 0~6 and 8~14) belonging to the joints from all action frames and merge them to form a joint action subsequence (corresponding to the first type of actuator); extract the 2 dimensions (indices 7 and 15) belonging to the gripper from all action frames and merge them to form a gripper action subsequence (corresponding to the second type of actuator).
[0061] It should be noted that if the robot is configured with a single arm or only a single gripper, the number of dimensions should be adjusted accordingly, but the partitioning logic remains the same. The two extracted subsequences retain their original temporal order, providing an independent data foundation for subsequent dynamic temporal compensation.
[0062] This approach, which adaptively preserves key motion frames based on the rate of change of joint angular velocity, avoids losing important poses, and reduces non-key frames, ensures that the joints and grippers are precisely synchronized at key points after the time shift, and balances execution efficiency and fine operation capabilities.
[0063] Step S203: Calculate the correction compensation amount in real time based on the robot's current task working parameters, and determine the dynamic timing offset between various actuators using the preset dynamic response benchmark deviation and the correction compensation amount. In some embodiments, the at least two types of actuators include a first type of actuator and a second type of actuator. Step S203 includes: acquiring a pre-calibrated dynamic response reference deviation between the first type of actuator and the second type of actuator; real-time acquisition of the robot's current task condition parameters, wherein the current task condition parameters include at least the load inertia and end-effector velocity of the first type of actuator and the second type of actuator; calculating the correction compensation amount based on the load inertia and the end-effector velocity; and obtaining the dynamic timing offset between the first type of actuator and the second type of actuator based on the dynamic response reference deviation and the correction compensation amount. Specifically, firstly, motion commands are sent to the first type of actuator and the second type of actuator at the same time, the time deviation of the two actuators reaching the specified posture is recorded, and then the motion frame number deviation is calculated based on the control frequency, which is the basic compensation amount. Then, during the actual execution of the robot, the current task status parameters, namely load inertia, are collected in real time. and terminal velocity Then, according to the formula The corrected compensation amount is calculated, where and The empirical weights obtained through operating condition calibration, The influence coefficient corresponding to the load inertia. The influence coefficient corresponding to the end velocity. Then, when the correction compensation amount is calculated... Then, the deviation is compared with the pre-calibrated dynamic response benchmark. Adding them together gives the final dynamic timing offset K= + .
[0064] It should be noted that load inertia This reflects the mass and moment of inertia of the object currently being grasped or moved by the robot. The greater the load and the higher the moment of inertia, the slower the response of the joint actuators; in this case, the correction compensation should be increased. End-effector velocity... This directly reflects the speed of the actuator's movement. The lower the end effector speed, the more obvious the timing deviation between the joint and the gripper, and the more compensation is needed.
[0065] As another embodiment, for scenarios with a large range of working conditions or high requirements for coordination accuracy, a complete formula for calculating the correction compensation amount can be used, that is, taking into account both the joint-gripper speed ratio and the overall speed of the gripper. Load inertia Dimensions of the object being operated on and terminal velocity Four parameters, correction compensation amount The calculation formula can be: ,in, This is the ratio of the joint actuator speed to the gripper actuator speed. When the joint speed is relatively slow, the ratio becomes smaller, and the compensation amount needs to be reduced appropriately. The size of the object being manipulated is considered. Larger objects may require longer time for the grippers to go from initial closing to full contact; therefore, this factor also contributes positively to the compensation amount. The coefficients β, γ, and δ represent these dimensions. All of these are empirical weights obtained through working condition calibration, which can be pre-determined based on the actual robot platform and task type.
[0066] It should be noted that due to differences in inertia, precision, and control strategies, there are often dynamic response differences between first-type actuators, such as gripper actuators (high speed, low precision), and second-type actuators, such as articulated actuators (low speed, high precision), leading to a mismatch in coordination between the two types of actuators. Therefore, it is necessary to delay the action of the first-type actuator, such as the gripper actuator, to compensate for this deviation, and to update the compensation amount using real-time feedback information. .
[0067] By pre-calibrating the baseline deviation and dynamically correcting the offset by collecting working parameters in real time, the timing compensation can accurately adapt to changes in the task, effectively offsetting the differences in joint and gripper response, and improving the coordination and operational stability.
[0068] Step S204: According to the dynamic time offset, the at least two types of action sub-sequences are relatively shifted on the time axis to generate asynchronously aligned composite action blocks; Further, the at least two types of actuators include a first type of actuator and a second type of actuator, and the at least two types of action sub-sequences include a first action sub-sequence and a second action sub-sequence. Step S204 includes: The first action subsequence corresponding to the first type of actuator is kept in its position on the time axis, and the second action subsequence corresponding to the second type of actuator is shifted as a whole in the positive direction of the time axis by the dynamic timing offset frame position, resulting in the third action subsequence after the second type of actuator is shifted; the third action subsequence is then merged with the unshifted first action subsequence to generate the composite action block. Specifically, after obtaining the dynamic timing offset K, the two types of action subsequences after decoupling are asynchronously aligned. Specifically, the original position of the action subsequence corresponding to the first type of actuator (joint actuator) on the time axis is kept unchanged, while the action subsequence corresponding to the second type of actuator (gripper actuator) is shifted as a whole in the positive direction of the time axis (i.e., the future direction) by K frames, thereby obtaining the shifted gripper action subsequence to offset the timing lead caused by the gripper actuator's faster response speed than the joint actuator. After the shift is completed, the unshifted joint action subsequence is merged with the shifted gripper action subsequence. During merging, the gripper sequence shifts backward by K frames, increasing the total number of frames for the entire composite motion block on the timeline. During merging, for the first K frames (before the gripper begins its movement), the gripper components can be filled with their initial state (e.g., fully open or maintaining the previous valid value); for the remaining frames after the joint sequence ends (where only gripper movement exists), the joint components can maintain the posture of the last frame. The resulting composite motion block can then be directly sent to the robot control service for execution, ensuring synchronization between the joints and gripper in actual physical response.
[0069] By keeping the first action subsequence unchanged and shifting the second action subsequence backward by an offset frame position, the two actuators can be synchronized in physical time, fundamentally solving the gripper advance problem and significantly improving operational stability and task success rate.
[0070] Step S205: The composite action block is sent to the robot's controller to drive various actuators to perform coordinated movements.
[0071] In some embodiments, after generating the composite action block B, it is encapsulated as a sequence control request and sent to the robot's control service for execution. The underlying controller can further smooth the trajectory based on the relationships between data points, reducing communication latency caused by multiple transmissions and thus improving execution efficiency. Finally, the control service parses the joint motion sub-sequences and gripper motion sub-sequences in the composite action block and drives the joint actuators and gripper actuators to move collaboratively according to the asynchronously aligned time axis.
[0072] This invention decouples the original action block into action sub-sequences corresponding to each actuator, laying the foundation for subsequent differentiated adjustments. Instead of using fixed delays or simple filtering, it collects current task parameters in real time and dynamically calculates a timing offset adapted to the current load, speed, and operating object, based on a pre-calibrated dynamic response benchmark deviation. This allows the offset to change with the operating conditions, avoiding the mismatch risk of static compensation. The action sub-sequences are then relatively shifted along the time axis according to this offset, generating asynchronously aligned composite action blocks. This "asynchronous alignment" breaks the original synchronous assumption of the model, ensuring that after being sent to the controller, the first and second types of actuators can reach their respective key action states synchronously in physical time, fundamentally solving the timing mismatch problem between the model output and heterogeneous actuators. Finally, by sending the composite action blocks as a whole, communication overhead is reduced, and the smooth, conflict-free movement of various actuators during task execution is ensured, significantly improving the stability of robot operation and the success rate of the task.
[0073] Compared to directly issuing the original actions in each action block one by one, this invention optimizes the processing and execution logic of gripper-related actions, improving the stability of gripping task execution. By reasonably reducing the number of actions issued, the latency caused by control command transmission is reduced, improving overall execution efficiency. Multiple dynamic and adaptive adjustment mechanisms enable the robot to cope with different task conditions and optimize task completion. In addition, this solution optimizes the trajectory smoothness in action blocks, ensuring more stable and smooth robot movement. This solution can form a plug-and-play modular access method for building a robot VLA dual-arm service system, ensuring efficient and stable task execution.
[0074] An embodiment of the present invention provides a cooperative execution system for robots, comprising: the cooperative execution system being used to perform the steps of the cooperative execution method for robots as described in the present invention.
[0075] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the robot cooperative execution method provided by any of the above-described method embodiments of the present invention.
[0076] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0077] Based on the above embodiments of the robot collaborative execution method, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the robot collaborative execution method of any embodiment of the present invention.
[0078] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0079] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0080] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0081] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the robot cooperative execution method described in any of the above-described method embodiments of the present invention.
[0082] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0083] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for cooperative execution of robots, characterized in that, include: Receive the original motion block output by the robot motion model, wherein the original motion block contains motion components corresponding to at least two types of actuators; The original action block is decoupled into at least two types of action sub-sequences, wherein each type of action sub-sequence corresponds to one type of executor; The correction compensation amount is calculated in real time based on the robot's current task operating parameters, and the dynamic timing offset between various actuators is determined by using the preset dynamic response benchmark deviation and the correction compensation amount. According to the dynamic temporal offset, the at least two types of action subsequences are relatively shifted on the time axis to generate asynchronously aligned composite action blocks; The composite action block is sent to the robot's controller to drive various actuators to perform coordinated movements.
2. The cooperative execution method for robots according to claim 1, characterized in that, The at least two types of actuators include a first type of actuator and a second type of actuator. The step of calculating a correction compensation amount in real time based on the robot's current task operating parameters, and determining the dynamic timing offset between different types of actuators using a preset dynamic response benchmark deviation and the correction compensation amount, includes: Obtain the pre-calibrated dynamic response benchmark deviation between the first type of actuator and the second type of actuator; The robot's current task condition parameters are collected in real time, wherein the current task condition parameters include at least the load inertia and / or end effector speed of the first type of actuator and the second type of actuator; The correction compensation amount is calculated based on the load inertia and the end velocity. The dynamic timing offset between the first type of actuator and the second type of actuator is obtained based on the dynamic response reference deviation and the correction compensation amount.
3. The cooperative execution method for robots according to claim 1, characterized in that, The at least two types of actuators include a first type of actuator and a second type of actuator. The decoupling of the original action block into at least two types of action sub-sequences includes: The motion components corresponding to at least one type of actuator in the original motion block are preprocessed to obtain the target motion block; According to the preset dimension division rules of each action frame in the target action block, the action components belonging to the same executor type are extracted to form the first action subsequence corresponding to the first type of executor and the second action subsequence corresponding to the second type of executor.
4. The robot cooperative execution method according to claim 3, characterized in that, The preprocessing of the motion components corresponding to at least one type of actuator in the original motion block to obtain the target motion block includes: The motion components corresponding to at least one type of actuator in the original motion block are nonlinearly mapped to obtain the first motion block; The first action block is subjected to a moving average filter to obtain the second action block; Redundant action frames of the second action block are removed to obtain the third action block; Adaptive keyframe downsampling is performed on the third action block to obtain the target action block.
5. The robot cooperative execution method according to claim 4, characterized in that, The step of performing a nonlinear mapping on the action components corresponding to at least one type of actuator in the original action block to obtain a first action block includes: Obtain the first action component corresponding to the second type of executor in the original action block; Based on the attribute type of the current operation object, the current task stage, and the interaction state between the second type of executor and the current operation object, the steepness parameter and closure threshold parameter of the nonlinear mapping curve are dynamically selected. The nonlinear mapping curve is constructed using the steepness parameter and the closure threshold parameter. The first motion component corresponding to the second type of actuator is numerically transformed through the nonlinear mapping curve to obtain the transformation result. The transformation result is then subjected to amplitude limiting processing to obtain the adjusted gripper motion component. The gripper motion component is merged with the second motion component corresponding to other unadjusted actuators in the original motion block to form the first motion block.
6. The cooperative execution method of robots according to claim 4, characterized in that, The process of removing redundant action frames from the second action block to obtain the third action block includes: Obtain the third motion component in each dimension for each motion frame in the second action block; For any two adjacent action frames in the second action block, calculate the change of the third action component in each dimension of the latter action frame relative to the former action frame. If the absolute value of the change in each dimension is less than the change threshold corresponding to that dimension, then the next action frame is determined to be a redundant action frame and the redundant action frame is removed from the second action block until all non-redundant action frames are retained, and all the non-redundant action frames are arranged into the third action block in the original time order.
7. The robot cooperative execution method according to claim 4, characterized in that, The adaptive keyframe downsampling of the third action block to obtain the target action block includes: Obtain the joint motion components of each motion frame in the third motion block, and calculate the rate of change of joint angular velocity between adjacent motion frames based on the joint motion components. The joint angular velocity change rate of each action frame is compared with a preset motion change threshold. If the joint angular velocity change rate is greater than the motion change threshold, the action frame is determined to be a key action frame and is retained, resulting in several first retained frames. The last action frame in the third action block is used as the second reserved frame; The first reserved frame and the second reserved frame are combined in their original time order to form the target action block.
8. The cooperative execution method of robots according to claim 1, characterized in that, The at least two types of actuators include a first type of actuator and a second type of actuator; the at least two types of action sub-sequences include a first action sub-sequence and a second action sub-sequence; the step of relative shifting the at least two types of action sub-sequences on the time axis according to the dynamic timing offset to generate asynchronously aligned composite action blocks includes: Keep the position of the first action subsequence corresponding to the first type of actuator unchanged on the time axis, and shift the frame position of the second action subsequence corresponding to the second type of actuator in the positive direction of the time axis by the frame position of the dynamic timing offset to obtain the third action subsequence after the second type of actuator is shifted; The third action subsequence is merged with the unshifted first action subsequence to generate the composite action block.
9. A cooperative execution system for robots, characterized in that, include: The cooperative execution system is used to perform the steps of the cooperative execution method of the robot as described in any one of claims 1-7.
10. A terminal device, characterized in that, include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the cooperative execution method of the robot as described in any one of claims 1-8.