Methods, apparatus, devices, media, and program product for training a robotic system

By analyzing the deviation of the failed task trajectories of the robot system from the task constraints, training priorities are determined and high-value samples are constructed, which solves the problem of uneven quality of failed samples and improves training efficiency and effectiveness.

CN121572340BActive Publication Date: 2026-04-28VASTAI TECH (SHANGHAI) INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VASTAI TECH (SHANGHAI) INC
Filing Date
2026-01-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the quality of failure samples used to train embodied robots varies, resulting in low training value, severe noise impact, and failure to effectively distinguish high-value samples, thus affecting the reliability and efficiency of training results.

Method used

By acquiring historical task information of the robot system, analyzing the state and action sequences of failed tasks, determining the degree of deviation between the trajectory and task constraints, determining training priorities based on the degree of deviation, and using high-value failure samples to train the control model.

Benefits of technology

This improved the correlation between training strategies and failure samples, enhanced the model training effect and the training efficiency of the robot system, and ensured the relevance and accuracy of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121572340B_ABST
    Figure CN121572340B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, apparatus, device, medium and program product for training a robot system. The method comprises: obtaining a plurality of historical trajectories associated with a plurality of failed tasks based on historical task information of the robot system; determining a degree of deviation of the plurality of historical trajectories from at least one constraint, the at least one constraint describing a condition to be followed for the robot system to successfully complete a task; determining a training priority of the plurality of historical trajectories based on the degree of deviation; determining a training strategy for the robot system based on the training priority, the training strategy indicating how to use at least one historical trajectory of the plurality of historical trajectories to perform a training process; and training a control model associated with the robot system based on the training strategy. In this way, embodiments of the present disclosure can improve the model training effect based on failed samples and improve the training efficiency of the robot system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, electronic devices, computer-readable storage media, and computer program products for training robot systems. Background Technology

[0002] With the rapid development of computer technology, the application of intelligent robots is becoming increasingly widespread. In existing embodied robot training, the use of failure samples for reinforcement learning or contrastive learning has been widely verified to improve policy robustness, boundary awareness, and adaptive correction capabilities. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for training a robot system is provided. The method includes: acquiring multiple historical trajectories associated with multiple failed tasks based on historical task information of the robot system, each historical trajectory indicating a state sequence and an action sequence of the robot system during the execution of the corresponding failed task, the state sequence indicating the control state of the robot system at multiple time points, and the action sequence indicating the control actions of the robot system at multiple time points; determining the degree of deviation of the multiple historical trajectories from at least one constraint, the at least one constraint describing the conditions required for the robot system to successfully complete the task; determining a training priority for the multiple historical trajectories based on the degree of deviation; determining a training strategy for the robot system based on the training priority, the training strategy indicating how to use at least one of the multiple historical trajectories to perform the training process; and training a control model associated with the robot system based on the training strategy.

[0004] In a second aspect of this disclosure, an apparatus for training a robot system is provided. The apparatus includes: a trajectory acquisition module, a deviation determination module, a priority determination module, a policy determination module, and a model training module. The trajectory acquisition module is configured to acquire multiple historical trajectories associated with multiple failed tasks based on historical task information of the robot system. Each historical trajectory indicates a state sequence and an action sequence of the robot system during the execution of the corresponding failed task. The state sequence indicates the control state of the robot system at multiple time points, and the action sequence indicates the control actions of the robot system at multiple time points. The deviation determination module is configured to determine the degree of deviation of the multiple historical trajectories from at least one constraint, the at least one constraint describing the conditions required for the robot system to successfully complete the task. The priority determination module is configured to determine the training priority of the multiple historical trajectories based on the degree of deviation. The policy determination module is configured to determine a training policy for the robot system based on the training priority, the training policy indicating how to use at least one of the multiple historical trajectories to perform the training process. The model training module is configured to train a control model associated with the robot system based on the training policy.

[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a computing unit to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] The solution provided in this disclosure addresses multiple historical trajectories associated with a robot system and multiple failed tasks. It determines the training priority of these historical trajectories by assessing their deviation from at least one constraint. Then, based on the training priority, it determines a training strategy for the robot system using failure samples corresponding to these historical trajectories. This strategy is then used to train a control model associated with the robot system. This effectively ensures the correlation between the training strategy and the deviation of the historical trajectories in the failure samples from at least one constraint, guaranteeing the sample value of the failure trajectories used for training. Consequently, it improves the model training effect based on failure samples and enhances the training efficiency of the robot system.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 A schematic diagram illustrating the implementation process of a training robot system according to some embodiments of the present disclosure is shown;

[0013] Figure 3A flowchart illustrating an example process for training a robot system according to some embodiments of this disclosure is shown;

[0014] Figure 4 A schematic structural block diagram of an example device for a training robot system according to some embodiments of the present disclosure is shown;

[0015] Figure 5 A schematic structural block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0020] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0021] As mentioned above, in existing embodied robot training, the use of failure samples for reinforcement learning or contrastive learning has been widely verified to improve policy robustness, boundary awareness, and adaptive correction capabilities.

[0022] However, in related technologies, the quality of failed samples used for training varies, resulting in low training value for some samples and potentially introducing noise, which seriously affects the reliability of training results.

[0023] In existing technologies, all failure samples are treated equally when performing training tasks, failing to effectively distinguish or label high-value samples that are sensitive to task success. As a result, it is difficult to distinguish the contribution of different failure samples to the training results, and thus there is no way to screen which samples are high-quality.

[0024] Furthermore, performing the training process using all failure samples not only increases training time but also reduces the convergence speed and generalization ability of the policy, and limits the policy's learning effect and cognitive ability on key boundaries.

[0025] Embodiments of this disclosure propose a scheme for data processing. The scheme includes: acquiring multiple historical trajectories associated with multiple failed tasks based on historical task information of a robot system, each historical trajectory indicating a state sequence and action sequence of the robot system during the execution of the corresponding failed task, the state sequence indicating the state of the robot system at multiple time points, and the action sequence indicating the control actions of the robot system at multiple time points; determining the degree of deviation of the multiple historical trajectories from at least one constraint, the at least one constraint describing the conditions required for the robot system to successfully complete the task; determining the training priority of the multiple historical trajectories based on the degree of deviation; determining a training strategy for the robot system based on the training priority, the training strategy indicating how to use at least one of the multiple historical trajectories to perform the training process; and training a control model associated with the robot system based on the training strategy.

[0026] In this manner, embodiments of the present disclosure target a robot system with multiple historical trajectories associated with multiple failed tasks. The training priority of the multiple historical trajectories is determined by the degree of deviation of the multiple historical trajectories from at least one constraint. Then, based on the training priority, a training strategy for training the robot system is determined using the failure samples corresponding to the multiple historical trajectories. The control model associated with the robot system is then trained accordingly. This effectively ensures the correlation between the training strategy and the degree of deviation of the historical trajectories in the failure samples from at least one constraint, guarantees the sample value of the failure trajectories used for training, and thus improves the model training effect based on failure samples, thereby improving the training efficiency of the robot system.

[0027] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0028] Example environment:

[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110 and robot system 120.

[0030] In some embodiments, the robot system 120 may include a robot. The robot system 120 may, for example, be capable of performing tasks such as work or movement through programming and automatic control. As an example, the robot can be implemented in various forms, such as a humanoid robot, a robotic arm, a wheeled robot, etc.

[0031] In some embodiments, the robot system 120 may include an intelligent system that performs tasks semi-autonomously or fully autonomously through perception, decision-making, and execution. In some embodiments, the robot system 120 may include an embodied intelligent robot. As an example, an embodied intelligent robot may be a system that possesses a physical body (e.g., motors, robotic arms, sensors, end effectors, etc.) and is capable of perceiving, learning, and intelligently interacting with the real world. For example, the end effector may include structures such as grippers, suction cups, and robotic hands; sensors may include force / torque sensors, vision sensors, position encoders, etc.

[0032] In some embodiments, the robot system 120 may acquire perception information based on configured sensors. The robot system 120 may be driven, for example, by a control model 160. The control model 160 may generate decisions based on the perception information acquired by the robot system 120 to control the robot system 120 to perform corresponding actions or tasks.

[0033] In some embodiments, the electronic device 110 may be a standalone hardware device or a software or hardware unit embedded within a hardware device. As an example, such an electronic device may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device may also support any type of user-facing interface (such as "wearable" circuitry).

[0034] In some embodiments, the electronic device 110 can acquire historical task information 130 of the robot system 120. In some embodiments, such historical task information 130 can be associated with at least one historical task performed by the robot system 120, and may include, for example, records of successful task execution and / or unsuccessful task execution by the robot system 120. Such historical task information 130 can indicate associated data of the robot system 120 during the execution of historical tasks. As an example, the electronic device 110 can extract multiple failure trajectories 135 associated with multiple failed tasks based on the historical task information, and determine the operating data associated with each failure trajectory 135. For example, such operating data may include, but is not limited to: the operating status, operating parameters, control actions, etc. of the robot system 120 at multiple time points corresponding to the failure trajectory 135.

[0035] In some embodiments, the electronic device 110 may construct a constraint model 140 associated with the task performed by the robot system 120 based on historical task information 130. The constraint model 140 may characterize at least one constraint (or condition) that needs to be satisfied to successfully complete the task. Such at least one constraint may be determined based on historical task information 130 or may be preset based on the task's execution requirements.

[0036] In some embodiments, the electronic device 110 may also utilize the constructed constraint model 140 to evaluate the extracted failure trajectory 135 and its associated data to determine the degree of deviation of the failure trajectory 135 from at least one constraint, in order to construct a training strategy 150 based on failure samples. The electronic device 110 can then use the failure samples constructed based on the failure trajectory 135 to train a control model 160 associated with the robot system 120 based on such training strategy 150.

[0037] In some embodiments, the control model 160 may include a controller built based on a neural network, a reinforcement learning strategy, an imitation learning model, or other machine learning models.

[0038] It should be understood that the structure and function of the various elements in the example environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0039] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0040] Example process:

[0041] Figure 2 A flowchart illustrating an implementation process 200 of a training robot system according to some embodiments of the present disclosure is shown. Implementation process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe the implementation process 200.

[0042] like Figure 2 As shown, in step 210, the electronic device 110 controls the robot system 120 to perform the task.

[0043] As an example, electronic device 110 can use control model 160 to control robot system 120 to perform at least one task. Such tasks may include, but are not limited to: grasping tasks, handling tasks, packaging tasks, processing tasks, etc.

[0044] For the data generated during task execution, such as execution trajectory, control actions, action sequences, operating parameters of robot system 120, and state sequences, robot system 120 can generate corresponding task logs as part of historical task information 130.

[0045] In step 220, the electronic device 110 determines whether the robot system 120 has successfully executed the task. If the task execution fails, the electronic device 110 can proceed to step 221; if the task execution is successful, the electronic device 110 can proceed to step 222.

[0046] In step 221, in response to determining that the robot system 120 has failed to perform a task, the electronic device 110 may collect the failure trajectory associated with the failed task.

[0047] As an example, such a failure trajectory includes the sequence of movement positions of the robot system 120 during the execution of the failed task, as well as the associated data such as control parameters, control states or actions, and operating states corresponding to such movement position sequences.

[0048] In some embodiments, such a failure trajectory may include at least a sequence of states and a sequence of actions of the robot system 120 during the execution of the failed task. The sequence of actions indicates the control states of the robot system 120 at multiple points in time, and the sequence of actions indicates the control actions of the robot system 120 at multiple points in time.

[0049] Such multiple time points can include multiple discrete or continuous time points. As an example, such multiple time points can be multiple time points determined based on a preset time interval, or multiple time points associated with the control sequence of the failed task.

[0050] In some scenarios, the state sequence may include, but is not limited to, the following information corresponding to multiple points in time: the angles and angular velocities of each joint of the embodied robot, the position (coordinates) and orientation (e.g., angles or Euler angles) of the end effector in three-dimensional space, the distance between the end effector and the target object, and the readings of each sensor (e.g., contact force, torque, image features, etc.).

[0051] The action sequence may include action instructions sent to the robot system 120 at multiple points in time, control instructions executed by the robot system 120, etc. As an example, such action instructions or control instructions may include, but are not limited to: position, velocity, torque, etc., associated with at least one joint of the robot, and pose, gripper opening and closing state, contact force with the target object, etc., associated with the end effector.

[0052] In some embodiments, the electronic device 110 may perform steps 230 to 260 after each failure trajectory is collected, or may perform steps 230 to 260 once after a preset time interval, or after a preset number of failure trajectories are collected, or after at least one failure trajectory associated with one or more tasks, without limitation.

[0053] In step 230, the electronic device 110 calculates the degree of deviation of each failed trajectory from the task constraint corresponding to the failed task.

[0054] Therefore, the electronic device 110 can accurately locate and quantitatively analyze the reasons for task execution failure, so as to determine the location or time point of deviation in the execution process corresponding to the failure trajectory and the corresponding degree of deviation.

[0055] In some embodiments, the task constraint corresponding to the failed task describes at least one condition that must be followed for the successful execution process of the failed task. As an example, such a task constraint may include at least one constraint indicated by the constraint model 140. For example, such at least one constraint may include, but is not limited to, at least one of the following: spatial constraints, temporal constraints, contact constraints, state constraints, etc.

[0056] As an example, given constraint k corresponding to time point t, electronic device 110 can calculate the failure trajectory based on the following formula. Degree of deviation from constraint k:

[0057] ;

[0058] In the formula Indicates the failure trajectory The state at time point t, Indicates the failure trajectory The action at time point t, Indicates the failure trajectory The degree of deviation from constraint k at time point t.

[0059] In some embodiments, spatial constraints can indicate the pose conditions that the end effector of the robot system 120 must follow to successfully complete a task. For example, spatial constraints can be used to describe specific positional and orientation conditions that the end effector of the robot system 120 needs to achieve at critical task nodes or throughout the task. For instance, spatial constraints can be used to describe whether the position of the end effector or target object in three-dimensional space enters the target recognition area, and whether the orientation of the end effector meets the allowable angular deviation range, etc., during task execution. For example, in an assembly task, based on spatial constraints, it can be required that the end effector control pin be moved above the hole and aligned with the centerline of the hole before insertion.

[0060] In this way, the present solution can construct the spatial constraints required to successfully complete the task based on the pose information of the end effector of the robot system 120, thereby constraining the task execution trajectory with the position and posture of the robot system 120 during the task completion process, in order to control the accuracy of the task execution process.

[0061] Relative to spatial constraints, electronic device 110 can determine the pose offset of the failed trajectory relative to the corresponding spatial constraints at each time point. As an example, such pose offset may include, but is not limited to, at least one of the following: the distance offset (e.g., Euclidean distance) between the end effector of robot system 120 and the target position, the angular offset (e.g., angular error) between the end effector and the target pose, etc. Such target position and target pose are used to define the spatial conditions required for mission success.

[0062] In some scenarios, the electronic device 110 can obtain the target position and target pose of the task at multiple time points from the corresponding task description, or it can extract the target position and target pose at multiple time points from the reference trajectory of successfully executing the task (e.g., extracting from a reference trajectory or extracting after cluster analysis of multiple reference trajectories). Then, the electronic device 110 uses the obtained target position and target pose as prior parameters and compares them with the actual position and actual pose of the failed trajectory at multiple time points to determine the pose offset of the failed trajectory relative to the spatial constraints at each time point.

[0063] In some embodiments, the total deviation of a failed trajectory from spatial constraints can be determined based on the weighted average, maximum, or sum of squares of the position and angle offsets of the failed trajectory at multiple time points.

[0064] In this way, the embodiments of this disclosure can use the distance offset between the end effector and the target position and the angular offset between the end effector and the target attitude to determine the degree of deviation between the failed trajectory and the spatial constraints in terms of distance and angle, effectively ensuring the accuracy of the degree of deviation corresponding to the spatial constraints.

[0065] In some embodiments, timing constraints can instruct the robot system 120 on the sequence of actions required to successfully complete a task. For example, timing constraints can describe the sequential relationship that the robot system 120 needs to satisfy during the execution of a task. For instance, in a "grasp-move-place" task, the gripper must first open and grasp the target object before the actions of lifting, moving, placing, and releasing the target object can be performed.

[0066] In this way, embodiments of the present disclosure can construct timing constraints for the sequence of actions or states that the robot system 120 needs to follow, thereby constraining the orderliness of the actions and states performed by the robot system 120 in the process of completing the task.

[0067] In some embodiments, the electronic device 110 can determine the degree of deviation between the failure trajectory and the timing constraints based on the prior parameters corresponding to the action sequence.

[0068] As an example, whether the execution process corresponding to a single point in time or time step meets the timing constraints can also depend on whether the preconditions required for the current step are met. Such preconditions are part of the task definition and can be obtained from information such as the task description, task planning, or at least one reference trajectory for successful task execution.

[0069] In some scenarios, electronic device 110 can determine the prior model (e.g., constraint model 140 or a sub-model within constraint model 140) corresponding to the temporal constraints using prior information associated with the task or at least a reference trajectory. Such a prior model could be a Hidden Markov Model, a temporal logic rule model, etc. Electronic device 110 can then calculate the probability or degree of fit of the action sequence corresponding to the failed trajectory to the prior model. The deviation of the failed trajectory from the temporal constraints can be defined as (1 - degree of fit), or as the number / severity of the number of times the action sequence corresponding to the failed trajectory violates the temporal constraints.

[0070] In this way, embodiments of the present disclosure can use prior parameters of the action sequence to determine the degree of deviation of the failure trajectory from the timing constraints, thereby ensuring the accuracy of the result of the degree of deviation of the failure trajectory from the timing constraints.

[0071] In some embodiments, contact constraints may indicate the contact conditions that the end effector of the robot system 120 must follow to successfully complete a task. These contact conditions indicate the contact state between the end effector and the target object. For example, such contact states may include, but are not limited to: whether physical contact occurs between the end effector and the target object, whether the contact is stable (e.g., whether the force / torque reaches a preset value), and whether undesired contact (e.g., a collision) occurs. For example, in a screw-tightening task, contact constraints may require the screwdriver bit (e.g., the end of the end effector) to maintain stable contact with the screw slot and apply positive pressure.

[0072] In some implementation scenarios, contact constraints can be determined based on sensor data such as tactile sensors and torque sensors.

[0073] In this way, embodiments of the present disclosure can construct contact constraints for the contact state between the end effector of the robot system 120 and the target object, thereby constraining the contact stability between the end effector and the target object during the completion of the task by the robot system 120.

[0074] In some embodiments, the electronic device 110 can determine the degree of deviation between the contact state between the end effector and the target object and the contact constraint based on the state parameters corresponding to the contact state. As an example, such state parameters may include, but are not limited to, visual parameters detected by a vision sensor, tactile parameters detected by a tactile sensor, and torque parameters detected by a torque sensor.

[0075] In some embodiments, the electronic device 110 can calculate the cumulative duration or average deviation of the actual state parameters of the failure trajectory exceeding the state constraints at multiple time points to determine the degree of deviation of the failure trajectory from the state constraints. As an example, to determine whether the contact state is stable (e.g., whether there are frequent switching between contact and separation states), the electronic device 110 can also quantify the degree of deviation corresponding to the contact state by detecting the oscillation frequency and amplitude of the force signal or distance signal.

[0076] In this way, the embodiments of this disclosure can use the state parameters corresponding to the contact state to determine the degree of deviation of the failure trajectory from the contact constraint, thereby ensuring the accuracy of the result of the degree of deviation of the failure trajectory from the contact constraint.

[0077] In some embodiments, state constraints may indicate the operational states that the robot system 120 must follow to successfully complete a task. Such constraint states may be associated with at least one operational metric of the robot system 120, describing whether the embodied robot's own execution state meets safety or task feasibility requirements. As an example, operational states may include the robot system 120's joint speeds, positions, angles, force feedback data, etc., not exceeding safety thresholds during task execution; the end effector or gripper being in the correct opening / closing range; the overall power consumption of the robot system 120 being below a power consumption threshold; and at least one stability metric associated with the robot system 120 or the embodied robot (e.g., the range of zero-torque points).

[0078] In some scenarios, the parameter thresholds associated with the constraint state can be obtained from the descriptive information or planning model associated with the task.

[0079] In this way, the embodiments of this disclosure can construct state constraints for the operating states that the robot system 120 needs to follow during the execution of tasks, thereby constraining the robot system 120 to maintain a stable operating state during the execution of tasks, so as to ensure the task execution capability of the robot system 120.

[0080] In some embodiments, the electronic device 110 can determine the degree of deviation between the failure trajectory and the state constraints based on the operating parameters corresponding to the operating state of the robot system 120 at multiple time points.

[0081] As an example, such operating parameters may include, but are not limited to: joint speed, angle, drive current, stability threshold, etc.

[0082] In this way, embodiments of the present disclosure can use the operating parameters corresponding to the operating state at multiple time points to determine the degree of deviation of the failure trajectory from the state constraints, thereby ensuring the accuracy of the deviation result.

[0083] In step 240, the electronic device 110 determines the training priority corresponding to the failure trajectory based on the degree of deviation between the failure trajectory and at least one of the above-mentioned task constraints.

[0084] In this scheme, the electronic device 110 can filter low-value failure trajectories by utilizing the degree of deviation between the failure trajectory and the task constraints, so as to select high-value failure trajectories for model training, thereby ensuring the relevance of model training and improving training efficiency.

[0085] As an example, electronic device 110 can utilize task constraints to filter the training value of key nodes corresponding to failed trajectories. Such key nodes can be determined based on the constraint information corresponding to the task constraints, or they can be determined based on the weight information corresponding to the task constraints.

[0086] In some embodiments, such weighting information may be preset during the task construction process or determined based on at least one reference trajectory for successfully completing the task.

[0087] As an example, after step 220, in response to determining that the task was successfully executed, electronic device 110 can execute step 222.

[0088] In step 222, the electronic device 110 may collect at least one reference trajectory indicating that the task was successfully executed, and then proceed to step 235.

[0089] In step 235, the electronic device 110 can analyze the constraint weights corresponding to at least one task constraint based on at least one reference trajectory. Then, in step 240, the electronic device 110 can determine the training priority corresponding to the failed trajectory based on the degree of deviation determined in step 230 and the constraint weights determined in step 235.

[0090] As an example, such constraint weights can be associated with points in time; that is, the constraint weights corresponding to the same task constraint may be different at different points in time.

[0091] In this way, the embodiments of this disclosure can effectively ensure the correlation between the training priority of the failed trajectory and the weight information of at least one constraint at different time points, thereby ensuring the accuracy of the training value of the determined failed trajectory.

[0092] In step 235 of some embodiments, the electronic device 110 may determine weight information of at least one constraint corresponding to multiple time points based on reference data corresponding to at least one reference trajectory at multiple time points. As an example, such reference data indicates deviation information of at least one reference trajectory corresponding to a preset offset at multiple time points, such preset offset may indicate interference parameters applied to the reference trajectory.

[0093] In some scenarios, for the reference trajectories collected in step 222, the electronic device 110 can add preset offsets to the task constraints corresponding to these reference trajectories at multiple time points, and obtain the running results after adding the preset offsets to determine the magnitude of the impact of the preset offsets on the running results at different time points, thereby determining the task weight of the task constraints at different time points. For example, for the same task constraint, if the preset offset has a smaller impact on the running results at time point t1 than at time point t2, the electronic device 110 can determine that the task weight of the task constraint at time point t2 is greater than the weight at time point t1; and for the first task constraint and the second task constraint at the same time point t3, if the preset offset has a greater impact on the running results of the first task constraint than on the running results of the second task constraint, the electronic device 110 can determine that the task weight of the first task constraint is greater than the task weight of the second task constraint at time point t3.

[0094] In some scenarios, if at a certain time point t4, the results (e.g., variance of spatial position offset) of multiple reference trajectories corresponding to the same task constraint (e.g., spatial constraint) are very small, the electronic device 110 can determine that the spatial constraint at this point is very strict, and the task weight corresponding to time point t4 should be set to a higher value. Conversely, if the variance is large, it indicates that the constraint is relatively loose, and the corresponding task weight can be set to a lower value.

[0095] In this way, the embodiments of this disclosure can determine the weight information of at least one constraint at multiple time points based on the deviation information of the robot system's successful task execution trajectory relative to a preset offset at multiple time points. This effectively ensures the accuracy of the weight information of at least one constraint at different time points, ensures the accuracy of the judgment of key time points, and thus ensures the accuracy of the training priority of historical trajectories, the accuracy of the training strategy, and the effectiveness of the training results.

[0096] In some embodiments, the electronic device 110 may determine the task weight of at least one constraint at multiple time points based on the matching degree between at least one collected reference trajectory and at least one preset event at multiple time points.

[0097] As an example, such a preset event may include, but is not limited to: first contact with the target object, discrete switching of the contact state, triggering action switching (such as reaching maximum pressure, starting to move or rotate, etc.).

[0098] In some scenarios, electronic device 110 can determine the reference state characteristics corresponding to each preset event by analyzing at least one reference trajectory. Then, electronic device 110 can calculate the actual state characteristics of the failed trajectory at the time point corresponding to the preset event, and calculate the matching degree between the actual state characteristics and the corresponding reference state characteristics. In time regions where the matching degree is below a threshold, electronic device 110 can determine that the robot system failed to correctly trigger or execute a critical event. In this case, the task weight of the task constraints (such as the contact constraints or spatial constraints required for the event) at that time point can be set to a higher value or increased by a preset amount.

[0099] As an example, corresponding to time point t, electronic device 110 can track the failure trajectory. The training priority is represented as follows:

[0100] ;

[0101] In the formula This represents the constraint weight of constraint k at time point t. Indicates the failure trajectory The state at time point t, Indicates the failure trajectory The action at time point t, Indicates the failure trajectory The degree of deviation from constraint k at time point t.

[0102] In this way, the embodiments of this disclosure can determine the weight information of at least one constraint at different time points based on the degree of matching between the reference trajectory of the robot system successfully completing the task and at least one preset event at multiple time points. This effectively ensures the accuracy of the weight information of at least one constraint at different time points, ensures the accuracy of the judgment of key time points, and thus ensures the accuracy of the training priority of failed trajectories, and ensures the pertinence of the training process and the accuracy of the training effect.

[0103] In step 250, the electronic device 110 constructs failure samples and a training strategy corresponding to the failure trajectory based on the training priority determined in step 240. Such a training strategy can instruct how to use the failure samples corresponding to the failure trajectory to perform the training process. As an example, such a training strategy can describe the order or importance (e.g., training weights) of the failure samples.

[0104] In some embodiments, the electronic device 110 can construct failure samples using failure trajectories based on information such as task constraints and data structures associated with the model. Then, based on the training priority corresponding to the failure trajectory, the training order of the failure samples is determined, wherein failure samples corresponding to greater deviations are preferentially used to train the control model. In some scenarios, the electronic device 110 can also determine the training weights corresponding to failure samples based on training priority, wherein training samples corresponding to greater deviations have higher training weights.

[0105] In this way, the embodiments of this disclosure determine the training order or training weight of multiple failed samples based on the training priority of multiple historical trajectories, so as to give priority to the training process using training samples with greater deviation, thereby effectively improving the model training effect based on failed samples and improving the training efficiency of the robot system.

[0106] In step 260, electronic device 110 trains control model 160 associated with robot system 120 based on the failure samples and training strategy constructed in step 250.

[0107] In some embodiments, the control model 160 can be a supervised learning-based model structure. The electronic device 110 can use failure samples as input to the model, and the correct action or action correction amount, adjusted based on at least one task constraint, as the label. The control model 160 is trained by minimizing the loss between the predicted execution action and the label. For failure samples corresponding to high-priority failure trajectories, the proportion in the loss function can be larger, or such failure samples can be prioritized for training.

[0108] In some embodiments, the control model 160 may be a reinforcement learning-based model structure. The electronic device 110 may use failure samples to build or supplement the experience replay buffer and may sample according to training priorities (prioritized experience replay) to accelerate the optimization process of the training policy.

[0109] The training process performed by the electronic device 110 using failed samples can be continuously iterated until the control model 160 achieves satisfactory performance on the validation set (which may include simulated successful and failed scenarios), or the success rate of reproducing failed tasks is significantly improved. Then, the electronic device 110 can reuse the trained control model 160 to execute the corresponding task to improve the accuracy and success rate of task execution.

[0110] In this manner, embodiments of the present disclosure, for the failure trajectory associated with the failed task of the robot system 120, determine the training priority of the failure trajectory by utilizing the degree of deviation of the failure trajectory from at least one constraint, and then construct failure samples and a training strategy for training the robot system 120 based on the training priority, and train the control model 160 associated with the robot system 120 accordingly. This effectively ensures the correlation between the training strategy and the degree of deviation of the failure trajectory from at least one constraint, guarantees the sample value of the failure trajectory used for training, and thereby improves the model training effect based on failure samples and improves the training efficiency of the robot system 120.

[0111] Figure 3 A flowchart illustrating an example process 300 of a training robot system according to some embodiments of the present disclosure is shown. Example process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 Let's describe example process 300.

[0112] like Figure 3 As shown, in step 310, the electronic device 110 acquires multiple historical trajectories associated with multiple failed tasks based on the historical task information 130 of the robot system. Each historical trajectory indicates the state sequence and action sequence of the robot system during the execution of the corresponding failed task. The state sequence indicates the control state of the robot system at multiple time points, and the action sequence indicates the control actions of the robot system at multiple time points.

[0113] In step 320, the electronic device 110 determines the degree of deviation of multiple historical trajectories from at least one constraint, which describes the conditions that the robot system must follow to successfully complete the task.

[0114] In step 330, the electronic device 110 determines the training priority of multiple historical trajectories based on the degree of deviation.

[0115] In step 340, the electronic device 110 determines a training strategy for the robot system based on training priorities. The training strategy indicates how to use at least one of a plurality of historical trajectories to perform the training process.

[0116] In step 350, the electronic device 110 trains a control model associated with the robot system based on a training strategy.

[0117] In this manner, embodiments of the present disclosure target a robot system with multiple historical trajectories associated with multiple failed tasks. The training priority of the multiple historical trajectories is determined by the degree of deviation of the multiple historical trajectories from at least one constraint. Then, based on the training priority, a training strategy for training the robot system is determined using the failure samples corresponding to the multiple historical trajectories. The control model associated with the robot system is then trained accordingly. This effectively ensures the correlation between the training strategy and the degree of deviation of the historical trajectories in the failure samples from at least one constraint, guarantees the sample value of the failure trajectories used for training, and thus improves the model training effect based on failure samples, thereby improving the training efficiency of the robot system.

[0118] In some embodiments, at least one constraint includes a spatial constraint, which indicates the pose conditions that the end effector of the robot system must follow in order to successfully complete the task.

[0119] In this way, the embodiments of this disclosure can construct the spatial constraints required to successfully complete the task based on the pose information of the end effector of the robot system, thereby ensuring the pose accuracy of the robot system during the task completion process.

[0120] In some embodiments, determining the degree of deviation of multiple historical trajectories from at least one constraint includes: the electronic device 110 determining the pose offset of the multiple historical trajectories relative to the spatial constraint at multiple points in time.

[0121] In this way, the embodiments of this disclosure can use the pose offset of multiple historical trajectories at multiple time points relative to spatial constraints to determine the degree of deviation of multiple historical trajectories, thereby ensuring that the robot system can be trained in a targeted manner using the pose offset of multiple historical trajectories in spatial constraints, effectively improving the training accuracy and efficiency of the model in spatial constraints.

[0122] In some embodiments, the pose offset includes at least one of the following: the distance offset between the end effector and the target position; the angular offset between the end effector and the target pose.

[0123] In this way, the embodiments of this disclosure can use the distance offset between the end effector and the target position and the angular offset between the end effector and the target posture to determine the degree of deviation between multiple historical trajectories and spatial constraints from the aspects of distance and angle, effectively ensuring the accuracy of the deviation corresponding to the spatial constraints, thereby ensuring the accuracy of the training priority of multiple historical trajectories and the accuracy of the training strategy.

[0124] In some embodiments, at least one constraint includes a timing constraint, which indicates the sequence of actions that the robot system must follow to successfully complete the task.

[0125] In this way, embodiments of the present disclosure can construct timing constraints for the sequence of actions that the robot system needs to follow, thereby ensuring the orderliness of the robot system's control actions during task completion.

[0126] In some embodiments, determining the degree of deviation of multiple historical trajectories from at least one constraint includes: the electronic device 110 determining the degree of deviation of multiple historical trajectories from the temporal constraint based on prior parameters corresponding to the action sequence.

[0127] In this way, the embodiments of this disclosure can use prior parameters of the action sequence to determine the degree of deviation of multiple historical trajectories from the temporal constraints, thereby ensuring that the robot system is trained in a targeted manner using multiple historical trajectories in terms of temporal constraints, effectively improving the training accuracy and efficiency of the model in terms of temporal constraints.

[0128] In some embodiments, at least one constraint includes a contact constraint, which indicates the contact conditions that the end effector of the robot system must follow in order to successfully complete the task, and the contact conditions indicate the contact state between the end effector and the target object.

[0129] In this way, the embodiments of this disclosure can construct contact constraints for the contact state between the end effector of the robot system and the target object, thereby ensuring the contact stability between the end effector and the target object during the completion of the task by the robot system.

[0130] In some embodiments, determining the degree of deviation of multiple historical trajectories from at least one constraint includes: the electronic device 110 determining the degree of deviation of the contact state between the end effector and the target object from the contact constraint based on the state parameters corresponding to the contact state.

[0131] In this way, the embodiments of this disclosure can use the state parameters corresponding to the contact state to determine the degree of deviation of multiple historical trajectories from the contact constraints, thereby ensuring that the robot system is trained in a targeted manner using multiple historical trajectories in terms of contact constraints, effectively improving the training accuracy and efficiency of the model in terms of contact constraints.

[0132] In some embodiments, at least one constraint includes a state constraint, which indicates the operating state that the robot system must follow to successfully complete the task.

[0133] In this way, embodiments of the present disclosure can construct state constraints for the operating states that the robot system needs to follow, thereby ensuring the stability of the operating state of the robot system in the process of completing the task.

[0134] In some embodiments, determining the degree of deviation of multiple historical trajectories from at least one constraint includes: the electronic device 110 determining the degree of deviation of multiple historical trajectories from state constraints based on operating parameters corresponding to the operating state at multiple time points.

[0135] In this way, the embodiments of this disclosure can use the operating parameters corresponding to the operating state at multiple time points to determine the degree of deviation of multiple historical trajectories from the state constraints, thereby ensuring that the robot system can be trained in a targeted manner using multiple historical trajectories in terms of state constraints, effectively improving the training accuracy and efficiency of the model in terms of state constraints.

[0136] In some embodiments, determining the training priority of multiple historical trajectories based on the degree of deviation includes: the electronic device 110 determining weight information of at least one constraint at multiple time points; and determining the training priority of multiple historical trajectories based on the weight information and the degree of deviation.

[0137] In this way, the embodiments of this disclosure can utilize the degree of deviation between multiple historical trajectories and at least one constraint, combined with the weight information of at least one constraint at multiple time points, to determine the training priority of multiple historical trajectories, thereby effectively ensuring the correlation between the training priority of historical trajectories and the weight information of at least one constraint, and thus ensuring the orderliness of the training strategy and the targeting and accuracy of training important nodes.

[0138] In some embodiments, determining the weight information of at least one constraint at multiple time points includes: the electronic device 110 determines the weight information of at least one constraint at multiple time points based on reference data corresponding to at least one reference trajectory in historical task information at multiple time points, wherein at least one reference trajectory indicates at least one execution trajectory corresponding to the successful completion of a task by the robot system, and the reference data indicates the deviation information of at least one execution trajectory corresponding to a preset offset at multiple time points.

[0139] In this way, the embodiments of this disclosure can determine the weight information of at least one constraint at multiple time points based on the deviation information of the robot system's successful task execution trajectory relative to a preset offset at multiple time points. This effectively ensures the accuracy of the weight information of at least one constraint at different time points, ensures the accuracy of the judgment of key time points, and thus ensures the accuracy of the training priority of historical trajectories, the accuracy of the training strategy, and the effectiveness of the training results.

[0140] In some embodiments, determining the weight information of at least one constraint at multiple time points includes: the electronic device 110 determining the weight information of at least one constraint at multiple time points based on the matching degree between at least one reference trajectory in historical task information and at least one preset event at multiple time points.

[0141] In this way, the embodiments of this disclosure can determine the weight information of at least one constraint at different time points based on the degree of matching between the reference trajectory of the robot system successfully completing the task and at least one preset event at multiple time points. This effectively ensures the accuracy of the weight information of at least one constraint at different time points, ensures the accuracy of the judgment of key time points, and thus ensures the accuracy of the training priority of historical trajectories, the accuracy of the training strategy, and the effectiveness of the training results.

[0142] In some embodiments, a training strategy for the robot system is determined based on training priority, including: the electronic device 110 determining the training order of multiple training samples corresponding to multiple failed tasks based on training priority, wherein training samples corresponding to greater deviation are preferentially used to train the control model; or determining the training weights of multiple training samples corresponding to multiple failed tasks based on training priority, wherein training samples corresponding to greater deviation have higher training weights.

[0143] In this way, the embodiments of this disclosure determine the training order or training weight of multiple failed samples based on the training priority of multiple historical trajectories, so as to give priority to the training process using training samples with greater deviation, thereby effectively improving the model training effect based on failed samples and improving the training efficiency of the robot system.

[0144] Example devices and equipment:

[0145] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example device 400 for a training robot system according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in electronic device 110. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0146] like Figure 4As shown, the device 400 includes: a trajectory acquisition module 410, a deviation determination module 420, a priority determination module 430, a strategy determination module 440, and a model training module 450. The trajectory acquisition module 410 is configured to acquire multiple historical trajectories associated with multiple failed tasks based on the robot system's historical task information. Each historical trajectory indicates the state sequence and action sequence of the robot system during the execution of the corresponding failed task. The state sequence indicates the control state of the robot system at multiple time points, and the action sequence indicates the control actions of the robot system at multiple time points. The deviation determination module 420 is configured to determine the degree of deviation between the multiple historical trajectories and at least one constraint, which describes the conditions required for the robot system to successfully complete the task. The priority determination module 430 is configured to determine the training priority of the multiple historical trajectories based on the degree of deviation. The strategy determination module 440 is configured to determine a training strategy for the robot system based on the training priority, the training strategy indicating how to use at least one of the multiple historical trajectories to perform the training process. The model training module 450 is configured to train a control model associated with the robot system based on the training strategy.

[0147] In some embodiments, at least one constraint includes a spatial constraint, which indicates the pose conditions that the end effector of the robot system must follow in order to successfully complete the task.

[0148] In some embodiments, the deviation determination module 420 is configured to determine the pose offset of multiple historical trajectories relative to spatial constraints at multiple time points.

[0149] In some embodiments, the pose offset includes at least one of the following: the distance offset between the end effector and the target position; the angular offset between the end effector and the target pose.

[0150] In some embodiments, at least one constraint includes a timing constraint, which indicates the sequence of actions that the robot system must follow to successfully complete the task.

[0151] In some embodiments, the deviation determination module 420 is configured to: determine the degree of deviation between multiple historical trajectories and temporal constraints based on prior parameters corresponding to the action sequence.

[0152] In some embodiments, at least one constraint includes a contact constraint, which indicates the contact conditions that the end effector of the robot system must follow in order to successfully complete the task, and the contact conditions indicate the contact state between the end effector and the target object.

[0153] In some embodiments, the deviation determination module 420 is configured to: determine the degree of deviation between the contact state and the contact constraint between the end effector and the target object based on the state parameters corresponding to the contact state.

[0154] In some embodiments, at least one constraint includes a state constraint, which indicates the operating state that the robot system must follow to successfully complete the task.

[0155] In some embodiments, the deviation determination module 420 is configured to: determine the degree of deviation between multiple historical trajectories and state constraints based on the running parameters corresponding to the running state at multiple time points.

[0156] In some embodiments, the priority determination module 430 is configured to: determine the weight information of at least one constraint at multiple time points; and determine the training priority of multiple historical trajectories based on the weight information and the degree of deviation.

[0157] In some embodiments, determining the weight information of at least one constraint at multiple time points includes: determining the weight information of at least one constraint at multiple time points based on reference data corresponding to at least one reference trajectory in historical task information at multiple time points, wherein at least one reference trajectory indicates at least one execution trajectory corresponding to the successful completion of a task by the robot system, and the reference data indicates deviation information of at least one execution trajectory corresponding to a preset offset at multiple time points.

[0158] In some embodiments, determining the weight information of at least one constraint at multiple time points includes: determining the weight information of at least one constraint at multiple time points based on the matching degree between at least one reference trajectory in historical task information and at least one preset event at multiple time points.

[0159] In some embodiments, the strategy determination module 440 is configured to: determine the training order of multiple training samples corresponding to multiple failed tasks based on training priority, wherein training samples corresponding to greater deviation are preferentially used to train the control model; or determine the training weights of multiple training samples corresponding to multiple failed tasks based on training priority, wherein training samples corresponding to greater deviation have higher training weights.

[0160] Figure 5 A block diagram is shown of a processing apparatus 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 5 The processing device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. As an example, Figure 5 The processing device 500 shown can be used to implement Figure 1 The electronic device 110 shown.

[0161] like Figure 5 As shown, the processing device 500 is in the form of a general-purpose electronic device. Components of the processing device 500 may include, but are not limited to, at least one processor or processing unit 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the processing device 500.

[0162] Processing device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to processing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within processing device 500.

[0163] The processing device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0164] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the processing device 500 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the processing device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0165] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Processing device 500 can also communicate as needed with one or more external devices (not shown) via communication unit 540. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with processing device 500, or with any device that enables processing device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).

[0166] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0167] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0168] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0169] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0171] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for training a robot system, characterized in that, The method includes: Based on the historical task information of the robot system, multiple historical trajectories associated with multiple failed tasks are obtained. Each historical trajectory indicates the state sequence and action sequence of the robot system during the execution of the corresponding failed task. The state sequence indicates the control state of the robot system at multiple time points, and the action sequence indicates the control actions of the robot system at the multiple time points. Determine the degree of deviation of the plurality of historical trajectories from at least one constraint, wherein the at least one constraint describes the conditions that the robot system must follow to successfully complete the task; Based on the degree of deviation, the training priority of the multiple historical trajectories is determined; Based on the training priority, a training strategy is determined for the robot system, the training strategy indicating how to perform the training process using at least one of the plurality of historical trajectories; and Based on the training strategy, a control model associated with the robot system is trained; The method of determining the training priority of the plurality of historical trajectories based on the degree of deviation includes: determining the weight information of the at least one constraint corresponding to the plurality of time points; and determining the training priority of the plurality of historical trajectories based on the weight information and the degree of deviation.

2. The method according to claim 1, characterized in that, The at least one constraint includes a spatial constraint, which indicates the pose conditions that the end effector of the robot system must follow in order to successfully complete the task.

3. The method according to claim 2, characterized in that, Determining the degree of deviation between the multiple historical trajectories and at least one constraint includes: Determine the pose offset of the plurality of historical trajectories relative to the spatial constraints at the plurality of time points.

4. The method according to claim 3, characterized in that, The pose offset includes at least one of the following: The distance offset between the end effector and the target position; The angular offset between the end effector and the target posture.

5. The method according to claim 1, characterized in that, The at least one constraint includes a timing constraint, which indicates the sequence of actions that the robot system needs to follow to successfully complete the task.

6. The method according to claim 5, characterized in that, Determining the degree of deviation between the plurality of historical trajectories and the at least one constraint includes: Based on the prior parameters corresponding to the action sequence, the degree of deviation between the multiple historical trajectories and the temporal constraints is determined.

7. The method according to claim 1, characterized in that, The at least one constraint includes a contact constraint, which indicates the contact conditions that the end effector of the robot system must follow in order to successfully complete the task, and the contact conditions indicate the contact state between the end effector and the target object.

8. The method according to claim 7, characterized in that, Determining the degree of deviation between the plurality of historical trajectories and the at least one constraint includes: Based on the state parameters corresponding to the contact state, the degree of deviation between the contact state between the end effector and the target object and the contact constraint is determined.

9. The method according to claim 1, characterized in that, The at least one constraint includes a state constraint, which indicates the operating state that the robot system needs to follow to successfully complete the task.

10. The method according to claim 9, characterized in that, Determining the degree of deviation between the plurality of historical trajectories and the at least one constraint includes: Based on the operating parameters corresponding to the operating state at the multiple time points, the degree of deviation between the multiple historical trajectories and the state constraints is determined.

11. The method according to claim 1, characterized in that, Determining the weight information of the at least one constraint at the plurality of time points includes: Based on the reference data corresponding to at least one reference trajectory in the historical task information at multiple time points, the weight information of at least one constraint corresponding to the multiple time points is determined, wherein the at least one reference trajectory indicates at least one execution trajectory corresponding to the successful completion of the task by the robot system, and the reference data indicates the deviation information of the at least one execution trajectory corresponding to a preset offset at multiple time points.

12. The method according to claim 1, characterized in that, Determining the weight information of the at least one constraint at the plurality of time points includes: Based on the matching degree between at least one reference trajectory in the historical task information and at least one preset event at multiple time points, the weight information corresponding to the at least one constraint at the multiple time points is determined.

13. The method according to claim 1, characterized in that, The step of determining a training strategy for the robot system based on the training priority includes: Based on the training priority, the training order of multiple training samples corresponding to the multiple failed tasks is determined, wherein training samples corresponding to greater deviations are preferentially used to train the control model; or Based on the training priority, training weights are determined for multiple training samples corresponding to the multiple failed tasks, wherein training samples corresponding to greater deviations have higher training weights.

14. An apparatus for training a robot system, characterized in that, The device includes: The trajectory acquisition module is configured to acquire multiple historical trajectories associated with multiple failed tasks based on the historical task information of the robot system. Each historical trajectory indicates the state sequence and action sequence of the robot system during the execution of the corresponding failed task. The state sequence indicates the control state of the robot system at multiple time points, and the action sequence indicates the control actions at the multiple time points. A deviation determination module is configured to determine the degree of deviation of the plurality of historical trajectories from at least one constraint, the at least one constraint describing the conditions that the robot system must follow to successfully complete the task; The priority determination module is configured to determine the training priority of the multiple historical trajectories based on the degree of deviation. A strategy determination module is configured to determine a training strategy for the robot system based on the training priority, the training strategy indicating how to perform the training process using at least one of the plurality of historical trajectories; and The model training module is configured to train a control model associated with the robot system based on the training strategy. The method of determining the training priority of the plurality of historical trajectories based on the degree of deviation includes: determining the weight information of the at least one constraint corresponding to the plurality of time points; and determining the training priority of the plurality of historical trajectories based on the weight information and the degree of deviation.

15. An electronic device, characterized in that, The electronic device includes: At least one processing unit; and At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 13 when executed by the at least one processing unit.

16. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that, The computer-executable instructions can be executed by a processing unit to implement the method according to any one of claims 1 to 13.

17. A computer program product, said computer program product being tangibly stored in a computer storage medium and comprising computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the device, cause the device to perform the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Integrated test management method and system

    CN121210299A