Method for collecting robot motion data, electronic device, and program product
Patent Information
- Application Number
- CN202611101959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]上述方案虽然能够为机器人模型训练提供数据来源,但普遍仍以人类示教或人类行为作为数据采集基础,存在工作量大、数据采集效率低、标准化程度低以及难以持续规模化采集等问题
Smart Images

Figure CN122807898A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method for acquiring robot motion data, an electronic device, a storage medium, and a program product. Background Technology
[0002] With the development of embodied intelligence, robot control models, and world models, there is a need to collect a large amount of robot motion data for model training. Currently, the main methods for collecting robot motion data include real-machine teleoperation, human first-person perspective behavior acquisition and motion capture, simulation environment data generation, and human-centered portable data acquisition.
[0003] Among these methods, remote operation of the robot is typically carried out by an operator using an exoskeleton, virtual reality device, or control device to remotely teach the robot, thereby collecting multimodal data such as visual information, end-effector pose, joint status, and tactile feedback during the robot's movement; first-person behavior acquisition and motion capture collect human movements and generate robot motion data using human-robot motion redirection technology; simulation environment data generation uses digital twins or robot simulation platforms to generate robot motion data; and human-centered data acquisition methods collect human movement information through devices such as haptic gloves and handheld grippers and convert it into data that the robot can learn.
[0004] While the aforementioned methods can provide data sources for robot model training, they generally rely on human instruction or human behavior as the basis for data collection, resulting in problems such as high workload, low data collection efficiency, low standardization, and difficulty in continuous large-scale data collection. Furthermore, due to the differences between human motion structures and robot body configurations, the motion generated by robot control models trained from human instruction data is usually not the robot's optimal motion, which can easily affect the accuracy, stability, and generalization ability of robot actions.
[0005] Furthermore, although simulation environments can quickly generate large amounts of data, there is a simulation-to-real gap between simulation and real physical environments. When simulation data is directly applied to real robots, the model performance is prone to decline due to differences in object shape, friction characteristics, and environmental factors.
[0006] Therefore, how to provide a method that can automatically collect high-quality motion data during the actual execution of tasks by robots, and further establish the correspondence between video information and robot motion state information, thereby generating a multimodal dataset suitable for training robot control models or world models, is an urgent technical problem to be solved. Summary of the Invention
[0007] This disclosure provides a method for acquiring robot motion data, an electronic device, a storage medium, and a program product.
[0008] According to one aspect of this disclosure, a method for acquiring robot motion data is provided, comprising: Obtain the position and orientation of the object being executed as identified by the camera device; The motion trajectory is determined by the motion planning algorithm based on the position, posture, and target task requirements of the object being executed; The robot is controlled to perform a target task according to a motion trajectory. During the execution of the target task, video data of the robot during the execution of the target task and motion execution data generated by the robot during the execution of the target task are acquired; the video data is collected by the camera device.
[0009] Optionally, the data acquisition method further includes: Establish the association between the video data and the motion execution data to generate multimodal motion data describing the robot's execution of the target task; the motion execution data includes at least motion state information characterizing the robot's execution process.
[0010] Optionally, acquiring video data of the robot performing the target task, and motion execution data generated by the robot while performing the target task, includes: While the robot is performing actions according to the motion trajectory, video data of the robot during the execution of the motion trajectory, as well as motion state information generated by the robot during the execution of the motion trajectory, are collected.
[0011] Optionally, establishing the association between the video data and the motion execution data includes: Obtain the video time information corresponding to each video frame in the video data; Obtain the motion state time information corresponding to each motion state information in the motion state information; Based on the correspondence between the video time information and the motion state time information, each video frame is associated with the motion state information at the corresponding time.
[0012] Optionally, establishing the association between the video data and the motion execution data includes: The acquired video dataset and motion state dataset are mapped using a post-processing algorithm, and the video data is associated with the motion state information to form the multimodal motion data.
[0013] Optionally, the motion state information includes: Robot joint angle sequence; robot end effector pose sequence; Robot motion trajectory sequence; The sensor data sequence of the end effector, wherein the sensor data includes at least one of pressure sensor data, visual-tactile sensor data, six-dimensional force sensor data, and gyroscope data; and at least one of the robot control command sequences generated by the motion planning algorithm; The motion state information is arranged in chronological order to form a sequence of motion state information describing the continuous motion process of the robot.
[0014] Optionally, the multimodal motion data includes: A sequence of video frames arranged in chronological order; And a sequence of motion state information arranged in chronological order, wherein each motion state information in the sequence corresponds to a video frame in the video frame sequence through the association relationship; The state of the robot when performing the same action is described by the video frames corresponding to the association relationship and the motion state information.
[0015] Optionally, the target task is an operational task that the robot autonomously performs based on the motion planning algorithm in an actual working environment.
[0016] Optionally, the method further includes: Robot training samples are generated based on the multimodal motion data; The video data serves as input data for the robot control model, and the motion state information in the motion execution data that has the correlation with the video data serves as training target data. This enables the robot control model to learn the correspondence between the video data and the robot's motion states.
[0017] Optionally, generating robot training samples includes: Based on video frames at multiple consecutive time points in the multimodal motion data and the motion state information corresponding to each video frame, a time-series training sample is generated to describe the complete motion process. The time-series training samples are used to train the robot control model to predict the robot's future motion state.
[0018] Optionally, generating robot training samples includes: The motion state information representing the execution result of the target task in the motion state information is used as the training target data; The video data that has the correlation with the training target data is used as the input data of the robot control model; This creates robot training samples that do not require manual action labeling.
[0019] Optionally, the video data is continuously acquired by a camera device in the form of a video stream, and the camera device is located at least one of the following: near the robot's end effector, above the working area, and to the side of the working area.
[0020] Optionally, the method further includes: Based on at least one of the changes in motion trajectory and the completion status of the target task, a valid motion data segment is determined from the multimodal motion data; Robot training samples are generated using the effective motion data fragments.
[0021] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, causing the processor to perform a robot motion data acquisition method according to any embodiment of this disclosure.
[0022] According to another aspect of this disclosure, a readable storage medium is provided, wherein execution instructions are stored therein, which, when executed by a processor, are used to implement a robot motion data acquisition method according to any embodiment of this disclosure.
[0023] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method for acquiring robot motion data according to any embodiment of this disclosure. Attached Figure Description
[0024] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0025] Figure 1 This is a flowchart illustrating a method for acquiring robot motion data according to one embodiment of this disclosure.
[0026] Figure 2 This is a schematic diagram illustrating the process of acquiring video data and motion execution data according to one embodiment of the present disclosure.
[0027] Figure 3 This is a schematic diagram illustrating the process of establishing the association between the video data and the motion execution data according to one embodiment of this disclosure.
[0028] Figure 4 This is a flowchart illustrating a method for acquiring robot motion data according to yet another embodiment of this disclosure.
[0029] Figure 5 This is a flowchart illustrating a method for acquiring robot motion data according to yet another embodiment of this disclosure.
[0030] Figure 6 This is a schematic block diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation
[0031] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.
[0032] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0033] Figure 1 This is a flowchart illustrating a method for acquiring robot motion data according to one embodiment of this disclosure.
[0034] refer to Figure 1 In some embodiments of this disclosure, the robot motion data acquisition method M100 of this disclosure includes: S102. Obtain the position and orientation of the object to be executed as identified by the camera device; S104. Determine the motion trajectory based on the position, posture, and target task requirements of the object being executed using a motion planning algorithm; S106. Control the robot to perform the target task according to the motion trajectory. During the robot's execution of the target task, acquire video data of the robot in the process of performing the target task, as well as motion execution data generated by the robot when performing the target task. The video data is collected by a camera device.
[0035] In some embodiments of this disclosure, the camera device may be an integrated configuration or a separate configuration.
[0036] When using an integrated configuration, the same camera device is configured to switch between 3D point cloud acquisition mode and video stream acquisition mode.
[0037] When performing step S102, the camera device operates in three-dimensional point cloud acquisition mode to identify the position and orientation of the object being executed (at this time it is used as a 3D camera).
[0038] When performing step S106, the camera device switches to two-dimensional video stream acquisition mode to acquire the video data.
[0039] When a split configuration is adopted, the recognition in step S102 is performed by the first camera device (e.g., a 3D camera), which is a camera capable of acquiring three-dimensional point clouds; and the video acquisition in step S106 is performed by the second camera device, which is a camera capable of acquiring video streams.
[0040] The first camera device (or the camera device that performs recognition in an integrated configuration) can also be placed near the robot's end effector to form an "eye on hand" recognition configuration.
[0041] In this disclosure, "robot" should be interpreted broadly, referring not only to robots with only robotic arms in industrial settings, but also to biomimetic robots, humanoid robots (such as wheeled humanoid robots and legged humanoid robots), etc. "Target task" refers to the operational task that the robot autonomously executes based on a motion planning algorithm in its actual working environment. Unlike related technologies where human operators drive robot movement through teaching, teleoperation, etc., the actions of the robot performing the target task in this disclosure are autonomously generated by the motion planning algorithm according to the task requirements, rather than being a record of human-taught actions.
[0042] In step S102, the position and orientation of the object to be executed can be identified and acquired through a camera device. The object to be executed is the object manipulated by the robot when performing the target task, such as a workpiece to be grasped, a material to be transported, or a component to be assembled. The camera device can be set on or near the robot's end effector, above, diagonally above, or to the side of the work area, etc., to collect three-dimensional information of the work scene, thereby identifying the position coordinates and orientation information of the object to be executed in space, providing basic data for subsequent motion planning.
[0043] In step S104, the motion planning algorithm calculates the motion trajectory required for the robot to move from its current position to the operating position and complete the target task based on the position and orientation of the object being executed and the requirements of the target task (such as the target placement position, operation path constraints, etc.). The motion trajectory specifies the spatial position and orientation that the robot should reach at each moment during the execution of the target task.
[0044] In step S106, the robot is controlled to execute the target task according to the motion trajectory. During the execution of the target task, video data of the robot during the execution of the target task, as well as motion execution data generated by the robot during the execution of the target task, are acquired. The motion execution data includes at least motion state information characterizing the robot's execution process.
[0045] Specifically, during the process of a robot performing a target task according to a motion trajectory, video data of the robot's actions can be collected by a camera device (e.g., a camera device mounted above or diagonally above the work area), and motion execution data generated by the robot can be obtained from the robot controller. Motion state information is used to describe the robot's motion state at various moments during execution, such as joint angles, end-effector pose, motion trajectory, or control commands.
[0046] When determining motion state information, various coordinate systems can be used for representation, such as base coordinate system, camera coordinate system, tool coordinate system, object coordinate system, joint coordinate system, etc. Motion state information in different coordinate systems can be selected or transformed according to the actual application scenario and subsequent model training requirements.
[0047] Continue to refer to Figure 1 In some embodiments of this disclosure, the robot motion data acquisition method M100 further includes: S108. Establish the correlation between video data and motion execution data to generate multimodal motion data that describes the robot's execution of the target task; the motion execution data shall include at least motion state information that characterizes the robot's execution process.
[0048] In step S108, a correlation is established between video data and motion execution data to generate multimodal motion data describing the robot's execution of the target task. Through this correlation, the visual information recorded in the video data corresponds to the motion state information in the motion execution data, thereby forming multimodal motion data that simultaneously contains visual information and motion state information, which can completely describe the process of the robot performing the target task.
[0049] This disclosure, through the aforementioned steps, simultaneously acquires video data and motion execution data during the robot's actual execution of the target task, and establishes a correlation between the two to form multimodal motion data. Since the robot's actions in performing the target task are autonomously generated by the motion planning algorithm, the collected motion execution data originates from the robot's actual movements and reflects the robot's true motion characteristics within its own configuration. The data acquisition process requires no human teaching, achieving automation and improving acquisition efficiency and standardization.
[0050] Figure 2 This is a schematic diagram illustrating the process of acquiring video data and motion execution data according to one embodiment of the present disclosure.
[0051] refer to Figure 2 In some embodiments of this disclosure, S106 described above, acquiring video data of the robot performing the target task and motion execution data generated by the robot while performing the target task, includes: S202. When the robot performs actions according to the motion trajectory, collect video data of the robot during the execution of the motion trajectory, as well as motion state information generated by the robot during the execution of the motion trajectory.
[0052] In step S202, when the robot performs the action according to the motion trajectory, video data of the robot during the execution of the motion trajectory and motion state information generated by the robot during the execution of the motion trajectory are collected.
[0053] The robot autonomously executes the target task according to the generated motion trajectory. During this execution, video data of the robot's actions is continuously acquired via a camera device, recording the robot's external visual appearance during movement. Simultaneously, motion state information generated by the robot during the execution of this trajectory is obtained from the robot controller, such as the angle change sequence of each joint, the pose change sequence of the end effector, and the actual motion trajectory sequence. This motion state information is generated in real time during the execution of the planned trajectory, reflecting the robot's actual action state while performing the target task.
[0054] In this disclosure, the robot's motion trajectory is autonomously generated by a motion planning algorithm based on the information of the object being executed and the task requirements in the actual work scenario, rather than being given by manual teaching. The video data and motion state information are collected during the process of the robot performing actions according to the autonomously generated motion trajectory, ensuring that the collected motion data comes from the robot's own actions when performing real work tasks. The data can truly reflect the motion characteristics of the robot's own configuration in the actual work scenario, providing an accurate and reliable data foundation for establishing the correlation between video data and motion execution data in the future.
[0055] Figure 3 This is a schematic diagram illustrating the process of establishing the association between video data and motion execution data according to one embodiment of this disclosure.
[0056] refer to Figure 3 In some embodiments of this disclosure, establishing the association between video data and motion execution data in step S108 described above includes: S302. Obtain the video time information corresponding to each video frame in the video data; S304. Obtain the motion state time information corresponding to each motion state information in the motion state information; S306. Based on the correspondence between video time information and motion state time information, associate each video frame with the motion state information at the corresponding time.
[0057] The video data is continuously captured by a camera device in the form of a video stream, with each frame corresponding to a capture time. Video timing information, such as a timestamp, identifies the capture time of each video frame. By reading the timestamp information in the video data, it is possible to determine when each video frame was captured.
[0058] Motion state information is generated and recorded in real time by the robot controller during the robot's execution of the target task. Each piece of motion state information also corresponds to a recording time, such as the time when a joint angle value was read or the time when an end-effector pose was sampled. Motion state time information is used to identify the recording time of each piece of motion state information, such as a timestamp. By reading the timestamp information in the motion state data, it is possible to determine the time when each set of motion state information was recorded.
[0059] After acquiring the video time information of each video frame and the motion state time information of each motion state information in the video data, video frames with the same time or a time difference within a preset threshold are associated with the motion state information, so that each video frame corresponds to a set of motion state information at the same moment. Therefore, for any given moment during the robot's execution of the target task, both video frames record the robot's external visual performance at that moment, and motion state information records the robot's motion state information at that moment; the two are linked through a temporal correspondence.
[0060] Through the above steps S302 to S306, this embodiment uses the acquisition timestamp of the video frame and the recording timestamp of the motion state information as the basis for association, and pairs the visual information at the same time or with the motion state information within a preset threshold, thereby forming multimodal motion data corresponding to the video frame and the motion state information.
[0061] In other embodiments of this disclosure, establishing the association between video data and motion execution data in step S108 described above includes: The acquired video dataset and motion state dataset are mapped using post-processing algorithms, and the video data is associated with the motion state information to form multimodal motion data.
[0062] Specifically, after the robot completes its target task, two datasets are obtained: a video dataset collected during the task's execution and a motion state dataset generated during the same process. The video dataset includes video data collected during the task's execution, while the motion state dataset includes motion state information generated during the task's execution. The video and motion state datasets are acquired independently and do not require real-time correlation during acquisition.
[0063] After obtaining the video dataset and motion state dataset, post-processing algorithms are used to process them accordingly, associating the relevant content in the video data with the corresponding motion state information, thereby establishing the relationship between the video data and the motion execution data, forming multimodal motion data.
[0064] In one optional implementation, the post-processing algorithm can employ a time-based data correspondence algorithm. Specifically, the video time information corresponding to each video frame in the video dataset and the motion state time information corresponding to each motion state information in the motion state dataset are read respectively, and the association between the video data and the motion state information is established based on the correspondence between the video time information and the motion state time information.
[0065] For example, when the video time information of a certain video frame is the same as the motion state time information corresponding to a certain motion state information, or when the time difference between the two is less than a preset threshold, the video frame can be matched with the motion state information. The preset threshold can be set according to factors such as video acquisition frequency, motion state sampling frequency, and data transmission latency.
[0066] When the video sampling frequency and the motion state sampling frequency are inconsistent, for each video frame, the time difference between the video time information corresponding to that video frame and the motion state time information corresponding to each motion state information can be calculated, and the motion state information with the smallest time difference can be selected as the motion state information corresponding to that video frame, thereby completing the correspondence between the video dataset and the motion state dataset.
[0067] In other embodiments, the post-processing algorithm can also be based on action feature matching, motion trajectory matching (which can match the position change features of the robot's end effector in the video data with the end pose change trajectory in the motion state data, and associate video segments with consistent motion change trends with corresponding motion state information), or other data association methods that can establish a correspondence between video data and motion state information, all of which fall within the protection scope of this disclosure.
[0068] In this implementation method, the acquisition of video data and motion state data are independent of each other, eliminating the need for real-time correlation processing during the acquisition process. This reduces the requirements for real-time processing capabilities during the data acquisition phase and improves the flexibility of the data acquisition process. Furthermore, for the already acquired video dataset and motion state dataset, post-processing algorithms can still establish the correlation between them to generate multimodal motion data.
[0069] For the above-described embodiments, the motion state information described above includes at least one of the following: Robot joint angle sequence; robot end effector pose sequence; Robot motion trajectory sequence; The sensor data sequence of the end effector, wherein the sensor data includes at least one of pressure sensor data, visual-tactile sensor data, six-dimensional force sensor data, and gyroscope data; A sequence of robot control commands generated by a motion planning algorithm.
[0070] The motion state information is arranged in chronological order to form a sequence of motion state information describing the continuous motion process of the robot.
[0071] Among them, the robot joint angle sequence is the information on the angle changes of each joint of the robot during the execution of actions, recorded in chronological order. For example, for a robot with multiple mechanical joints, the rotation angles corresponding to each joint can be acquired at a preset sampling frequency during the robot's movement, and arranged according to the acquisition time to form a joint angle sequence, so as to reflect the movement changes of each joint of the robot over time.
[0072] A robot end effector pose sequence is a record of the position and orientation information of a robot end effector in chronological order. The end effector can be a gripper, tool, suction cup mechanism, manipulator, or other actuator at the end of a robotic arm. During the execution of a target task, the robot controller can acquire the spatial position and orientation of the end effector at different points in time, such as the end effector's position coordinates and orientation information relative to a preset coordinate system. This pose information from multiple points in time is then arranged chronologically to form the end effector pose sequence.
[0073] A robot motion trajectory sequence is the spatial path information actually traversed by the robot's end effector or the robot body during the execution of a target task. When a robot performs actions according to the motion trajectory generated by the motion planning algorithm, it can record information such as the robot's spatial position, direction of movement, and trajectory points at each moment, and form a motion trajectory sequence in chronological order to describe the robot's actual motion path during execution.
[0074] The robot control command sequence generated by the motion planning algorithm is a sequence of instruction information used to control the robot to perform actions after the motion planning algorithm determines the robot's motion trajectory. For example, the motion planning algorithm can generate joint motion commands, end effector motion commands, or speed control commands at each moment according to the target task requirements. The robot controller controls the robot to perform corresponding actions according to the above control commands, and can record the above control commands in chronological order to form a control command sequence.
[0075] In this disclosure, the aforementioned motion state information can be used individually as motion execution data, or it can be combined to form motion execution data. For example, during the execution of a grasping task, changes in robot joint angles, end effector pose, and motion trajectory can be recorded simultaneously to obtain motion state information that can describe the robot's action process from different dimensions.
[0076] The motion state information is arranged chronologically to form a sequence describing the robot's continuous movements. Specifically, during the robot's execution of the target task, motion state information is continuously acquired at each moment according to a preset sampling period, and then sorted according to the time information corresponding to each motion state information, so that the motion state information at different time points forms a continuous data sequence. Therefore, the motion state information not only represents the robot's action state at a specific moment, but also reflects the robot's complete motion process from the start to the end of the task.
[0077] Through the above methods, this disclosure can obtain continuous motion state information of the robot during the execution of the target task, so that the motion execution data can accurately describe the changes in the robot's own actions.
[0078] For the above-described implementation methods, the multimodal motion data may include: The sequence of video frames arranged in chronological order, and the sequence of motion state information arranged in chronological order, with each motion state information in the motion state information sequence corresponding to a video frame in the video frame sequence through a correlation relationship.
[0079] Specifically, the state of the robot when performing the same action is described by corresponding video frames and motion state information based on the association relationship.
[0080] Each video frame in the video frame sequence corresponds to each motion state information in the motion state information sequence through the aforementioned association relationship. This association relationship establishes a match between a specific video frame and a specific motion state information, thereby enabling the combination of the robot's visual information when performing a certain action with the corresponding motion state information to jointly describe the robot's state when performing that action.
[0081] Thus, the resulting multimodal motion data can completely and continuously describe the process of the robot performing the target task from both visual and state dimensions, providing a structured data foundation for subsequent motion analysis or model learning.
[0082] The target task described above in this disclosure refers to the operation task that the robot autonomously executes based on motion planning algorithms in a real working environment; the operation task includes, but is not limited to, grasping task, handling task, assembly task, grinding task, spraying task, welding task, etc.
[0083] Grasping tasks involve the robot acquiring an object from the work area and transferring it to a designated location; handling tasks involve the robot moving an object from one workstation to another; and assembly tasks involve the robot combining multiple objects according to predetermined positions and orientations. In all of these tasks, the robot acquires the position and orientation information of the object through a camera device. A motion planning algorithm then generates a motion trajectory based on the object's position, orientation, and the target task requirements, and autonomously executes the corresponding operation according to this trajectory.
[0084] During the robot's execution of the aforementioned tasks, video data and motion execution data are collected simultaneously, and the correlation between the two is established to form multimodal motion data. Since the target task is a real-world operation in an actual working environment, the collected motion data originates from the robot's own movements when performing real tasks. The data accurately reflects the robot's motion characteristics in actual working conditions and is consistent with the actual application scenario.
[0085] Figure 4 This is a flowchart illustrating a method for acquiring robot motion data according to yet another embodiment of this disclosure.
[0086] refer to Figure 4 ,exist Figure 1 Based on the illustrated implementation, the robot motion data acquisition method M100 of this embodiment further includes: S110. Generate robot training samples based on multimodal motion data.
[0087] In this model, video data serves as input data for the robot control model, while motion state information in the motion execution data that is related to the video data serves as training target data, enabling the robot control model to learn the correspondence between video data and robot motion states.
[0088] Specifically, through steps S106 and S108, multimodal motion data corresponding to video data and motion state information has been obtained.
[0089] In multimodal motion data, each video frame or video segment is associated with its corresponding motion state information through a correlation relationship.
[0090] Based on the aforementioned correlations, corresponding video data and motion state information can be directly constructed as training samples: video data is input into the robot control model, and motion state information correlated with the video data is used as the training target data for the robot control model's expected output. By learning the correspondence between video data and motion state information in a large number of such training samples, the robot control model can establish a mapping capability from visual information to action states.
[0091] Since the training samples disclosed herein are directly derived from the multimodal motion data generated by the robot when performing the target task in the actual working environment, and the correspondence between video data and motion state information has been established in the data acquisition stage or post-processing stage, the training sample generation process does not require manual action annotation of video data, and can automatically obtain the pairing relationship between input data and training target data.
[0092] In some implementations, step S110 described above, generating robot training samples, may include: Based on video frames at multiple consecutive time points in the multimodal motion data and the motion state information corresponding to each video frame, temporal training samples are generated to describe the complete motion process; the temporal training samples are used to train the robot control model to predict the robot's future motion state.
[0093] In some implementations, generating robot training samples in step S110 may include generating time-series training samples to describe the complete motion process based on video frames at multiple consecutive time points in the multimodal motion data and the motion state information corresponding to each video frame.
[0094] For example, during a robot's grasping task, a camera continuously captures video data of the entire process from the robot's approach to the object to the completion of the grasp. The robot controller simultaneously records motion state information such as the end effector pose and joint angles at each moment. Through correlation, each video frame corresponds to the motion state information at the corresponding moment. From this multimodal motion data, video frames at multiple consecutive time points can be selected, such as N consecutive frames from the robot's end effector approaching the object to the gripper closing and completing the grasp. Simultaneously, N sets of end effector poses and joint angles corresponding to these N frames are acquired. Using the consecutive video frames as input data and the corresponding motion state information sequence as training target data, a time-series training sample is formed. This time-series training sample records the complete dynamic process of the robot performing the grasping action.
[0095] In some implementations, model training can be based on the temporal continuity of the training samples. Specifically, a sequence of video frames from multiple consecutive time points in the training samples can be used as input data, and the motion state information corresponding to subsequent time points can be used as training target data, enabling the robot control model to learn the ability to predict the robot's subsequent motion state based on continuous visual information.
[0096] For example, during the process of a robot performing a grasping task, video frames collected at multiple consecutive time points can be used as input, and the end effector pose, joint angle, or motion trajectory information corresponding to subsequent time points can be used as training targets. Through training with multiple time-series training samples, the robot control model can learn the correspondence between changes in video information over time and changes in the robot's motion state.
[0097] This implementation method utilizes video frames and motion state information from consecutive time points in multimodal motion data to construct temporal training samples, enabling the training samples to preserve the temporal continuity of the robot's actions. Compared to training based solely on data from a single time point, this implementation method allows the robot control model to learn the characteristics of the robot's motion state changing over time, thereby improving the robot's ability to predict future motion states based on continuous visual input.
[0098] In some implementations, the generated robot training samples described above include: The motion state information representing the execution result of the target task in the motion state information is used as the training target data; the video data that is related to the training target data in the video data is used as the input data of the robot control model; thus forming robot training samples that do not require manual annotation of action labels.
[0099] Specifically, when a robot performs a target task, its motion state information includes state data that reflects the result of the task execution. For example, in a grasping task, the pose of the end effector or the gripper's closed state at the moment of grasping is the motion state information indicating the result of the grasping task; in a handling task, the robot's end effector pose when the object reaches the target position is the motion state information indicating the result of the handling task; and in an assembly task, the robot's end effector pose or joint angle when parts are assembled in place is the motion state information indicating the result of the assembly task.
[0100] Since a correlation has been established between video data and motion state information in multimodal motion data, the motion state information representing the result of the target task can be used as the training target data. The corresponding video data is then determined through this correlation and used as the input data to form a set of training samples. These training samples take the visual information at the moment the task execution result is reached as input and the action state corresponding to the task execution result as the expected output. The action labels for the training samples are automatically provided by the motion state information, eliminating the need for manual annotation of the actions in the video data.
[0101] This implementation method automatically generates training samples by utilizing motion state information representing task execution results in multimodal motion data. This not only obtains high-quality task execution result samples but also achieves automatic labeling of training samples, eliminating the need for manual action labeling and improving the automation and efficiency of training sample generation.
[0102] In this disclosure, the video data described above can be continuously acquired by a camera device in the form of a video stream, and the camera device can be located at least one of the following: near the robot's end effector, above the work area, and to the side of the work area.
[0103] Figure 5 This is a flowchart illustrating a method for acquiring robot motion data according to yet another embodiment of this disclosure.
[0104] refer to Figure 5 ,exist Figure 1 Based on the illustrated implementation, the robot motion data acquisition method M100 of this embodiment further includes: S112. Determine valid motion data segments from multimodal motion data based on at least one of the changes in motion trajectory and the completion status of the target task; S114. Generate robot training samples using effective motion data fragments.
[0105] During the execution of the target task, the robot's camera continuously collects video data, and the robot controller continuously records motion state information. The resulting multimodal motion data encompasses all data from the start to the end of the task. However, not all data within this complete timeframe is directly related to the actions performed in the target task. For example, before the task begins, the robot may be in a standby state or preparing to move to the operating position; after the task is completed, the robot may be in a reset or transition phase, moving to the next task position. The data from these phases is less directly related to the actions performed in the target task itself.
[0106] To extract the parts directly related to the target task's execution actions from complete multimodal motion data, effective motion data segments can be determined based on at least one of the changes in motion trajectory and the completion status of the target task.
[0107] In some implementations, valid motion data segments are determined based on changes in the motion trajectory.
[0108] Specifically, during the robot's execution of the target task, the motion trajectory generated by the motion planning algorithm specifies the spatial position and posture the robot should reach at each moment. The motion state information contains the sequence of motion trajectories actually executed by the robot. When the motion trajectory begins to change regularly according to the requirements of the target task, it indicates that the robot has begun to execute the target task; when the motion trajectory reaches the target position and remains stable or begins to transition to the next stage, it indicates that the target task has been completed. Therefore, by analyzing the characteristics of the motion trajectory changes, the start and end times of the target task can be determined, and video data and motion state information between these two times can be extracted to form effective motion data segments.
[0109] For example, in a grasping task, the robot's end effector starts from a position close to the object being grasped, moves along a planned trajectory to the grasping point, and completes the grasping action. The motion trajectory of the end effector exhibits a change characteristic of first approaching and then closing. By identifying the moment when the end effector begins to move towards the object being grasped and the trajectory change after the grasping action is completed, the start and end positions of the effective motion data segments can be determined.
[0110] In other implementations, valid motion data segments are determined based on the completion status of the target task.
[0111] Specifically, the robot control system can acquire completion status information of the target task, such as the gripper closing position signal in a grasping task, the position signal of the object reaching the target location in a handling task, and the feedback signal of the completion of component assembly in an assembly task. This completion status information can directly indicate whether the target task has been completed.
[0112] Based on the task completion status information, the completion time of the target task can be determined, and by combining the changes in the motion trajectory, the start time of the task can be traced back, thereby determining the time range of the effective motion data segment.
[0113] By implementing the above methods, effective motion data segments that are directly related to the target task's execution actions can be selected from complete multimodal motion data. Then, training samples can be generated using these effective motion data segments. This reduces the interference of data unrelated to task execution on model training and improves the data quality and training efficiency of the training samples.
[0114] The execution subject of the robot motion data acquisition method in the specific embodiments of this disclosure can be an electronic device such as a computer.
[0115] Therefore, based on any of the above embodiments, this disclosure also provides an electronic device that can execute the robot motion data acquisition method of any of the embodiments described above.
[0116] Figure 6 This is a schematic block diagram of an electronic device 1000 according to one embodiment of the present disclosure.
[0117] The hardware architecture of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application of the hardware and overall design constraints. Bus 1100 connects various circuits, including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0118] Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, this diagram uses only one connection line, but this does not imply that there is only one bus or one type of bus.
[0119] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.
[0120] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.
[0121] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0122] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0124] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0126] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment / mode or example, which are included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0127] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0128] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.
Claims
1. A method for acquiring robot motion data, characterized in that, include: Obtain the position and orientation of the object being executed as identified by the camera device; The motion trajectory is determined by the motion planning algorithm based on the position, posture, and target task requirements of the object being executed; The robot is controlled to perform a target task according to a motion trajectory. During the execution of the target task, video data of the robot during the execution of the target task and motion execution data generated by the robot during the execution of the target task are acquired. The video data is collected by the camera device.
2. The method for acquiring robot motion data according to claim 1, characterized in that, The data acquisition method also includes: Establish the association between the video data and the motion execution data to generate multimodal motion data describing the robot's execution of the target task; the motion execution data includes at least motion state information characterizing the robot's execution process.
3. The method for acquiring robot motion data according to claim 1, characterized in that, The acquisition of video data of the robot performing the target task, and motion execution data generated by the robot while performing the target task, includes: While the robot is performing actions according to the motion trajectory, video data of the robot during the execution of the motion trajectory, as well as motion state information generated by the robot during the execution of the motion trajectory, are collected.
4. The method for acquiring robot motion data according to claim 2, characterized in that, Establishing the association between the video data and the motion execution data includes: Obtain the video time information corresponding to each video frame in the video data; Obtain the motion state time information corresponding to each motion state information in the motion state information; Based on the correspondence between the video time information and the motion state time information, each video frame is associated with the motion state information at the corresponding time.
5. The method for acquiring robot motion data according to claim 2, characterized in that, Establishing the association between the video data and the motion execution data includes: The acquired video dataset and motion state dataset are mapped using a post-processing algorithm, and the video data is associated with the motion state information to form the multimodal motion data.
6. The method for acquiring robot motion data according to claim 1, characterized in that, The motion state information includes: Robot joint angle sequence; robot end effector pose sequence; Robot motion trajectory sequence; The sensor data sequence of the end effector, wherein the sensor data includes at least one of pressure sensor data, visual-tactile sensor data, six-dimensional force sensor data, and gyroscope data; and at least one of the robot control command sequences generated by the motion planning algorithm; The motion state information is arranged in chronological order to form a sequence of motion state information describing the continuous motion process of the robot.
7. The method for acquiring robot motion data according to any one of claims 1 to 6, characterized in that, Optionally, the multimodal motion data includes: A sequence of video frames arranged in chronological order; And a sequence of motion state information arranged in chronological order, wherein each motion state information in the sequence corresponds to a video frame in the video frame sequence through the association relationship; The state of the robot when performing the same action is described by the video frames corresponding to the association relationship and the motion state information. Optionally, the target task is an operational task that the robot autonomously performs based on the motion planning algorithm in an actual working environment; Optionally, the method further includes: Robot training samples are generated based on the multimodal motion data; The video data serves as input data for the robot control model, and the motion state information in the motion execution data that has the correlation with the video data serves as training target data. This enables the robot control model to learn the correspondence between the video data and the robot's action states; Optionally, generating robot training samples includes: Based on video frames at multiple consecutive time points in the multimodal motion data and the motion state information corresponding to each video frame, a time-series training sample is generated to describe the complete motion process. The time-series training samples are used to train the robot control model to predict the robot's future motion state; Optionally, generating robot training samples includes: The motion state information representing the execution result of the target task in the motion state information is used as the training target data; The video data that has the correlation with the training target data is used as the input data of the robot control model; This creates robot training samples that do not require manual action labeling; Optionally, the video data is continuously acquired by a camera device in the form of a video stream, and the camera device is located at least one of the following: near the robot's end effector, above the working area, and to the side of the working area; Optionally, the method further includes: Based on at least one of the changes in motion trajectory and the completion status of the target task, a valid motion data segment is determined from the multimodal motion data; Robot training samples are generated using the effective motion data fragments.
8. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes the execution instructions stored in the memory, causing the processor to perform the robot motion data acquisition method according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium stores execution instructions, which, when executed by a processor, are used to implement the robot motion data acquisition method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for acquiring robot motion data as described in any one of claims 1 to 7.