A method, device and equipment for generating robot teaching data and a storage medium
Patent Information
- Application Number
- CN202611340828.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-01
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明提供了一种机器人示教数据的生成方法、装置、设备及存储介质,以解决机器人示教数据时序混乱以及存储成本高的问题
[0009]本发明提供的机器人示教数据的生成方案,通过获取包含不同时间戳的多源数据帧序列,并以RGB-D图像数据帧序列为基准,对剩余序列进行数据补全处理,从而获得精确对齐的对齐数据帧序列,有效克服了现有技术中因多路数据时序错乱而引入的噪声干扰,保证了“动作-视觉-状态”因果链条在时序上的一致性,显著提升了模仿学习或机器人策略训练过程中模型的收敛速度,并增强了模型在不同场景下的泛化性能。本方案中的数据压缩大幅减少了机器人示教数据的整体数据量,从而有效缓解了存储系统的空间压力,显著降低了硬件存储成本。同时,在压缩深度图像时,采用无损压缩策略,确保了深度数据中包含的空间几何信息和距离信息的完整性。
Smart Images

Figure CN122845950A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and storage medium for generating robot teaching data. Background Technology
[0002] Currently, existing robot teaching data processing solutions typically include three stages: data acquisition, data conversion, and model training. In the data acquisition stage, the system uses teleoperated devices, robot controllers, cameras, and status recording programs to acquire multimodal data such as images, joint states, gripper states, control commands, and task information, and records the corresponding timestamps. In the data conversion stage, the acquired data is usually formatted into a predetermined training data format. In the model training stage, the model to be trained reads data from the corresponding dataset for imitation learning or robot policy training.
[0003] However, existing robot teaching data typically consists of multiple streams of data. During subsequent data fusion, feature extraction, or model training, the noise introduced by these streams due to temporal discrepancies severely interferes with the model's learning of the "action-vision-state" causal relationship, reducing the model's convergence speed and generalization performance. Furthermore, existing robot teaching data often includes a large amount of image data, which is usually stored independently in the file system as raw pixels frame by frame, resulting in a large data volume and high storage costs. During the data preprocessing stage, the training framework needs to load the complete sequence or a large number of samples into memory simultaneously for batch normalization and tensor transformation, leading to excessive memory consumption, impacting processing efficiency, and even causing memory shortages. This results in high computational costs for training. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for generating robot teaching data, in order to solve the problems of disordered timing and high storage costs of robot teaching data.
[0005] In a first aspect, the present invention provides a method for generating robot teaching data, comprising: Obtain a multi-source data frame sequence, wherein the timestamps of data frames in any two types of data frame sequences in the multi-source data frame sequence are different; Based on the RGB-D image data frame sequence in the multi-source data frame sequence, data completion processing is performed on the remaining sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. Lossy image compression is performed on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, and lossless image compression is performed on the depth image data frame group acquired by each camera to obtain depth image byte blocks. Robot teaching data is generated based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
[0006] Secondly, the present invention provides a device for generating robot teaching data, comprising: The data acquisition module is used to acquire a multi-source data frame sequence, wherein the timestamps of the data frames in any two types of data frame sequences in the multi-source data frame sequence are different; The data supplementation module is used to perform data completion processing on the remaining sequence based on the RGB-D image data frame sequence in the multi-source data frame sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. The data compression module is used to perform lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks, and to perform lossless image compression on the depth image data frame groups acquired by each camera to obtain depth image byte blocks. The teaching data generation module is used to generate robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
[0007] Thirdly, the present invention provides an electronic device comprising: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to perform the method for generating robot teaching data described in the first aspect.
[0008] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to execute the method for generating robot teaching data as described in the first aspect.
[0009] The robot teaching data generation scheme provided by this invention acquires multi-source data frame sequences containing different timestamps and uses RGB-D image data frame sequences as a benchmark to perform data completion processing on the remaining sequences, thereby obtaining a precisely aligned data frame sequence. This effectively overcomes the noise interference introduced by the temporal disorder of multi-channel data in existing technologies, ensuring the temporal consistency of the "action-vision-state" causal chain. This significantly improves the convergence speed of the model during imitation learning or robot policy training and enhances the model's generalization performance in different scenarios. The data compression in this scheme greatly reduces the overall data volume of robot teaching data, thereby effectively alleviating the space pressure on the storage system and significantly reducing hardware storage costs. Simultaneously, when compressing depth images, a lossless compression strategy is adopted to ensure the integrity of the spatial geometric and distance information contained in the depth data.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a method for generating robot teaching data according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of a method for generating robot teaching data according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of a robot teaching data generation device according to Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. In the description of this invention, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0015] Example 1 Figure 1 This is a flowchart of a method for generating robot teaching data according to Embodiment 1 of the present invention. This embodiment can be applied to the generation of robot teaching data. The method can be executed by a robot teaching data generation device, which can be implemented in hardware and / or software. The robot teaching data generation device can be configured in an electronic device, which can be composed of two or more physical entities or a single physical entity.
[0016] like Figure 1 As shown, the method for generating robot teaching data provided in Embodiment 1 of the present invention specifically includes the following steps: S101. Obtain a multi-source data frame sequence, wherein the timestamps of data frames in any two types of data frame sequences in the multi-source data frame sequence are different.
[0017] In this embodiment, raw multi-modal data (i.e., multi-source data frame sequences) generated by the robot during teleoperation or teaching can be acquired. Specifically, this can include multiple RGB-D (Red, Green, Blue-Depth) images, robot body state and sensor feedback, control commands, task language information, and data source identifiers. The robot body state and sensor feedback include joint positions, velocities, currents, and torques; end effector pose; gripper or dexterous hand state; and force, torque, and tactile information. Data sources include model output and manual teleoperation. All of the above data types have corresponding timestamps and frame numbers.
[0018] S102. Based on the RGB-D image data frame sequence in the multi-source data frame sequence, perform data completion processing on the remaining sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera. The RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera.
[0019] RGB-D images, with their stable acquisition frequency and reliable timestamps, contain rich visual semantic information and precise spatial geometric information. They are central to the model's learning of the "vision-action" causal relationship and are suitable as a reference timeline for aligning multimodal teaching data. In this embodiment, to address the issue of temporal asynchrony among multi-source data, the sequences other than the RGB-D image data frame sequence in the multi-source data frame sequence can be completed (e.g., through time mapping, interpolation, extrapolation, and state reconstruction) and / or aligned with the RGB-D image data frame sequence to obtain an aligned data frame sequence. The time of the frame data in this aligned data frame sequence corresponds one-to-one with the sampling time of the RGB-D image data frame.
[0020] S103. Perform lossy image compression on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, and perform lossless image compression on the depth image data frame group acquired by each camera to obtain depth image byte blocks.
[0021] In this embodiment, to balance storage efficiency and feature preservation, different compression methods can be used for different image data. RGB images (color images) are mainly used to help the model "understand" scene content (such as object colors, textures, and boundaries). Visual models are not sensitive to minor color deviations in images; losing some unimportant details (i.e., lossy compression) has minimal impact on recognition results but can significantly save storage space. Lossy compression encoding formats include the H.26x series (such as H.264 / AVC) and the MPEG series (such as MPEG-2). Depth images are used to record the precise distance from each pixel to the camera (i.e., spatial geometric information). In robot operations, even a few millimeters of error in the depth value can cause the robotic arm to "miss" or "collide" with an object during grasping. Therefore, lossless compression can be performed on the depth image data frames acquired by each camera to ensure the integrity of all spatial geometric data. Lossless compression encoding formats include FFV1 and HuffYUV. The compressed byte blocks can be understood as binary data packets. When the training model reads data, it only needs to load this data block into memory once, instead of frequently opening and closing thousands of small files, greatly improving data I / O (input / output) efficiency. The RGB image byte blocks include motion residual information, which characterizes the motion changes between RGB image data frames in an RGB image data frame group. The motion residual information includes motion vectors and residual information. The motion vectors describe the motion changes between frames, while the residual information represents the difference between the predicted result and the original image.
[0022] It is worth noting that when compressing RGB image data frame groups, motion residuals are recorded in byte blocks, which record the direction and distance of movement of objects in the image. The lost "time series dynamic information" is preserved in the byte block in the form of "motion residual vector".
[0023] S104. Generate robot teaching data based on the RGB image byte block, the depth image byte block, and the aligned data frame sequence.
[0024] In this embodiment, robot teaching data can be generated using these RGB image byte blocks, depth image byte blocks, and aligned data frame sequences. For example, the RGB image byte blocks, depth image byte blocks, and aligned data frame sequences (including state, action, pose, and metadata, specifically including robot body state and sensor feedback including joint position, speed, current and torque, end effector pose, gripper or dexterous hand state, and force, torque, and tactile information, etc.) are encapsulated according to a unified time index to form robot teaching data containing vision, state, action, and metadata fields.
[0025] The technical solution of this invention acquires a multi-source data frame sequence containing different timestamps, and uses the RGB-D image data frame sequence as a benchmark to perform data completion processing on the remaining sequence, thereby obtaining a precisely aligned data frame sequence. This effectively overcomes the noise interference introduced by the temporal disorder of multi-channel data in the prior art, ensuring the temporal consistency of the "action-vision-state" causal chain, significantly improving the convergence speed of the model during imitation learning or robot policy training, and enhancing the generalization performance of the model in different scenarios. The data compression in this solution significantly reduces the overall data volume of robot teaching data, thereby effectively alleviating the space pressure on the storage system and significantly reducing hardware storage costs. At the same time, when compressing depth images, a lossless compression strategy is adopted to ensure the integrity of the spatial geometric information and distance information contained in the depth data.
[0026] Optionally, before performing lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks, the method further includes: performing various data format conversions on the aligned data frame sequence and the RGB-D image data frame sequence to obtain corresponding conversion results, and dividing the conversion results into multiple batches to obtain a second intermediate data frame sequence, wherein the data format includes data format and data type; and streaming the second intermediate data frame sequence into a preset storage space according to the data batches; wherein performing lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks includes: performing lossy image compression on the RGB image data frame groups acquired by each camera in the preset storage space to obtain RGB image byte blocks.
[0027] For example, the aligned data frame sequence and the RGB-D image data frame sequence can be represented collectively as D: D={D1,D2,…,D n} D i ={I i Z i ,q i ,a i ,P i C i ,m i} Where i is the frame count, i=1, 2, ..., n. I represents the RGB image, Z represents the depth image, q represents the joint state, a represents the control action or action source information, P represents the end effector or TCP (Tool Center Point) pose, C represents the camera pose or camera parameters, and m represents metadata such as timestamp, frame number, task information, and data source identifier.
[0028] Specifically, the aligned data frame sequence and the RGB-D image data frame sequence can be converted into different data formats to obtain corresponding conversion results. For example, they can be converted to HDF5 (Hierarchical Data Format version 5) and NPZ formats. Each conversion result can be divided into multiple batches to obtain a second intermediate data frame sequence. To reduce memory usage, a streaming write method can be used to read, process, and write the second intermediate data frame sequence in batches. During the writing process to the preset storage space, the image, depth, state, pose, and metadata in the second intermediate data frame sequence can be type-normalized and structured for storage. For example, depth data in meters can be converted to uint16 data in millimeters.
[0029] Example 2 Figure 2 This is a flowchart of a method for generating robot teaching data according to Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above optional technical solutions, and a specific method for generating robot teaching data is given.
[0030] Optionally, the step of performing data completion processing on the remaining sequence based on the RGB-D image data frame sequence in the multi-source data frame sequence to obtain an aligned data frame sequence includes: determining the RGB-D image data frame sequence in the multi-source data frame sequence as a time-anchored data frame sequence, wherein the multi-source data frame sequence includes multiple image frame sequences, robot body state and sensor feedback frame sequences, control command frame sequences, task language information frame sequences, and data source identifier frame sequences; mapping the remaining sequence to the time axis of the time-anchored data frame sequence, and performing data alignment and / or completion processing on the remaining sequence in the time axis according to the time-anchored data frame sequence to obtain a first intermediate data frame sequence; and completing the end effector spatial pose of the first intermediate data frame sequence according to the time-anchored data frame sequence to obtain a time-synchronized aligned data frame sequence.
[0031] Optionally, the step of performing lossy image compression on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, and performing lossless image compression on the depth image data frame group acquired by each camera to obtain depth image byte blocks, includes: encoding the RGB image data frame group acquired by each camera into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks, and encoding the depth image data frame group acquired by each camera into AVI byte blocks to obtain depth image byte blocks, wherein the AVI byte block is a data unit that uses FFV1 encoding to form an AVI file and conforms to the RIFF block structure.
[0032] Optionally, after generating robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence, the method further includes: performing data cleaning and semantic annotation on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence to obtain data to be verified, and receiving target frame information; if the data corresponding to the target frame information is compressed data in the data to be verified, then decoding the target window data corresponding to the target frame information from the compressed data to obtain decoded data, wherein the compressed data is data in the RGB image byte blocks, and the target window data includes the compressed data corresponding to the target frame information and multiple frames of data adjacent to the corresponding compressed data; visually outputting the decoded data to realize the playback verification of the data to be verified, and obtaining robot teaching data samples.
[0033] like Figure 2 As shown in Embodiment 2 of the present invention, a method for generating robot teaching data specifically includes the following steps: S201. Obtain the multi-source data frame sequence.
[0034] In this multi-source data frame sequence, the timestamps of data frames from any two types of data frame sequences are different.
[0035] S202. Determine the RGB-D image data frame sequence in the multi-source data frame sequence as a time-anchored data frame sequence; map the remaining sequence to the time axis of the time-anchored data frame sequence, and perform data alignment and / or completion processing on the remaining sequence in the time axis according to the time-anchored data frame sequence to obtain a first intermediate data frame sequence; according to the time-anchored data frame sequence, complete the end effector spatial pose of the first intermediate data frame sequence to obtain a time-synchronized aligned data frame sequence.
[0036] The multi-source data frame sequence includes a multi-channel image frame sequence, a robot body state and sensor feedback frame sequence, a control command frame sequence, a task language information frame sequence, and a data source identifier frame sequence.
[0037] Specifically, firstly, one RGB-D image sequence can be selected from the multi-source data frame sequences as a reference (i.e., as the anchor data frame sequence). Using the timestamps of the anchor data frame sequence as the scale, a standard timeline is established, with the time points on the timeline... .in, This represents the k-th RGB-D image data frame, and N represents the total number of RGB-D image data frames. Data from other RGB-D image sequences in the multi-source data frame sequence and the remaining sequences (such as joint states, gripper states, and control commands) need to be mapped onto this time axis. For other RGB-D image sequences in the multi-source data frame sequence, RGB-D images with existing correspondences (such as RGB-D images with the same frame number as the anchor data frame) can be directly aggregated into the time-anchored data frame sequence. For the remaining sequences, the time-anchored data frame sequence can be used as a template to perform data alignment and / or completion using methods such as linear interpolation to obtain an intermediate result with preliminary time alignment, i.e., the first intermediate data frame sequence.
[0038] Specifically, when the target time T0 corresponding to a certain time-anchored data frame is outside the adjacent sampling times T1 and T2 of a certain data path (such as joint states) in the remaining sequence, interpolation can be performed based on the joint states corresponding to T1 and T2 to estimate the joint state at time T0, so as to align it with the time-anchored data frame. When time T0 is located at the beginning or end of the joint state sequence, and it is not possible to obtain the joint states of the adjacent sampling times before and after simultaneously, the joint state at time T0 can be estimated using methods such as pre-value preservation, post-value preservation, nearest neighbor matching, trend extrapolation, or other preset estimation methods. Specifically: pre-value preservation uses the most recent joint state before T0 as the joint state at time T0; post-value preservation uses the most recent joint state after T0 as the joint state at time T0; nearest neighbor matching uses the joint state closest in time to T0 as the joint state at time T0; and trend extrapolation estimates the joint state at time T0 based on the changing trend of the joint states of adjacent sampling times.
[0039] Optionally, the time-anchored data frame sequence can also be a robot control cycle sequence, a joint state sequence, a depth camera frame sequence, an external trigger signal sequence, a control action sequence, or a system clock frame sequence, etc.
[0040] Then, for the precise position and orientation of the end effector in the three-dimensional space in the first intermediate data frame sequence, the spatial pose of the end effector in the first intermediate data frame sequence at each time step of the time-anchored data frame sequence can be calculated based on known parameters such as joint angles, robot kinematic model and camera mounting position. It is also necessary to calculate the transformation relationship between the camera coordinate system and the world coordinate system so as to associate the image pixel coordinates with the three-dimensional space coordinates, and obtain a time-synchronized and spatially complete multimodal sample sequence (i.e., aligned data frame sequence), so that the visual observation, robot state and end effector spatial information are consistent in time.
[0041] Furthermore, the step of completing the end effector spatial pose of the first intermediate data frame sequence based on the time-anchored data frame sequence to obtain a time-synchronized aligned data frame sequence includes: interpolating the joint states of the first intermediate data frames corresponding to the data frames in the time-anchored data frame sequence to obtain an interpolation result, and then performing forward kinematics calculations on the interpolation result to obtain a time-synchronized aligned data frame sequence; or, performing forward kinematics calculations on the joint states of the first intermediate data frames corresponding to the time-anchored data frame sequence to obtain the corresponding tool center point pose, and then interpolating the tool center point pose to obtain a time-synchronized aligned data frame sequence.
[0042] Specifically, for each sampling time of the time-anchored data frame sequence, the joint state of the first intermediate data frame corresponding to that sampling time can be interpolated. Based on the interpolation result, the TCP pose at that time can be calculated using forward kinematics, thus obtaining a time-synchronized aligned data frame sequence. Alternatively, forward kinematics can be calculated on the joint states of the adjacent frames before and after the first intermediate data frame corresponding to that sampling time to obtain the corresponding TCP pose. Then, interpolation can be performed on the TCP pose (where linear interpolation is used for translation, and spherical linear interpolation or other equivalent rotational interpolation methods are used for rotation) to obtain a time-synchronized aligned data frame sequence.
[0043] S203. Perform various data format conversions on the aligned data frame sequence and the RGB-D image data frame sequence to obtain corresponding conversion results, and divide the conversion results into multiple batches to obtain a second intermediate data frame sequence; stream the second intermediate data frame sequence into a preset storage space according to the data batches.
[0044] The data format includes data format and data type, and the RGB-D image data frame sequence includes RGB image data frame groups and depth image data frame groups acquired by at least one camera.
[0045] S204. In the preset storage space, the RGB image data frame group acquired by each camera is encoded into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks, and the depth image data frame group acquired by each camera is encoded into AVI byte blocks to obtain depth image byte blocks.
[0046] The AVI (Audio Video Interleave) byte block is a data unit that uses FFV1 encoding to form an AVI file and conforms to the RIFF block structure.
[0047] Specifically, H.264 / AVC stands for H.264 / MPEG-4 AVC (Advanced Video Coding).
[0048] For example, when the data format of the RGB image data frame group and the depth image data frame group is HDF5, H.264 / AVC encoding can be used to encode each RGB image data frame group acquired by the camera into MP4 byte blocks to obtain RGB image byte blocks, and each depth image data frame group acquired by the camera can be encoded into AVI byte blocks to obtain depth image byte blocks. After encoding, the frame-by-frame images in the original HDF5 are replaced with a single or a small number of video byte datasets.
[0049] Furthermore, the step of encoding the RGB image data frame groups acquired by each camera into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks includes: converting the RGB image data frame groups acquired by each camera into YUV format data frames to obtain YUV frame groups, and performing optimal motion vector search on the prediction units of the YUV frame groups to obtain optimal motion vectors; determining the spatial domain residual matrix using the optimal motion vectors, and obtaining a binary bitstream based on the spatial domain residual matrix through frequency domain transformation, quantization, and entropy coding; and encapsulating the binary bitstream to obtain RGB image byte blocks. The advantage of this setup is that it preserves key dynamic motion features (motion residuals) during the robot's action execution process during RGB image compression, allowing the compressed data to still provide rich kinematic information for subsequent model training.
[0050] Specifically, the entire process of obtaining RGB image byte blocks can be described as follows: 1) Pre-processing: Convert RGB image data frames to YUV (color encoding system) format to obtain YUV frame groups.
[0051] 2) Motion estimation: The YUV frames in the YUV frame group are divided into coding tree units (CTUs) or macroblocks, and further divided into smaller prediction units (PUs). An optimal motion vector search is performed on each PU to obtain the optimal motion vector.
[0052] 3) Motion compensation and residual generation: Based on the optimal motion vector, the corresponding predicted pixel block is extracted from the reference frame, and the predicted value is subtracted from the original pixel value of the predicted pixel block to obtain the residual matrix, which is the spatial domain residual matrix. The value of this matrix is the spatial domain representation of the motion residual (i.e., residual information).
[0053] 4) Transformation, quantization, and entropy coding: Transformation: The spatial domain residual matrix is transformed from the spatial domain to the frequency domain to obtain the transformation coefficient matrix.
[0054] Quantization: The transform coefficients in the transform coefficient matrix are divided by integers according to the quantization parameters to obtain the quantized residual coefficients. Entropy coding: Golomb coding or arithmetic coding is performed on the optimal motion vector and the quantized residual coefficients to generate a binary code stream.
[0055] 5) Encapsulated as MP4 byte blocks Finally, the binary stream is packaged into network abstraction layer units, and the necessary file headers are added according to the MP4 container format. These are then combined into the final output MP4 byte blocks, which are RGB image byte blocks.
[0056] S205. Based on the RGB image byte block, the depth image byte block, and the aligned data frame sequence, generate robot teaching data.
[0057] Optionally, before performing data cleaning and semantic annotation on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence, the following can be performed: data transformation on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence to obtain transformed RGB image byte blocks, transformed depth image byte blocks, and transformed aligned data frame sequences. Through this step, multiple training data versions can be generated from the same batch of processed data (i.e., RGB image byte blocks, depth image byte blocks, and aligned data frame sequences), such as data versions with different sampling frequencies, different action representations, different coordinate system representations, different historical observation lengths, or different future action window lengths, thereby reducing data transformation during the training phase and generating multiple reusable data versions.
[0058] Specifically, data transformations include, but are not limited to: adjusting the sampling frequency, pruning valid segments, converting absolute actions into action increments, converting actions into relative actions in TCP or end-effector coordinate systems, constructing historical observation blocks, constructing future action blocks, generating fill masks, and synchronously updating data fields, metadata, and statistics.
[0059] Wherein, the action increment Δa i It can be calculated from action 'a' in adjacent frames:
[0060] Among them, a i Indicates the action in frame i. This represents the action in frame i+1, where i is the frame count. For future action blocks, if the model to be trained predicts that the future action frame to be output exceeds the time boundary, the excess part can be masked to generate a corresponding padding mask.
[0061] S206. Perform data cleaning and semantic annotation on the RGB image byte block, the depth image byte block, and the aligned data frame sequence to obtain the data to be verified, and receive the target frame information; if the data corresponding to the target frame information is compressed data in the data to be verified, decode the target window data corresponding to the target frame information from the compressed data to obtain the decoded data; visualize and output the decoded data to realize the playback verification of the data to be verified and obtain robot teaching data samples.
[0062] The compressed data is data in the RGB image byte block, and the target window data includes compressed data corresponding to the target frame information and multiple frames of data adjacent to the corresponding compressed data.
[0063] Specifically, data cleaning and visualization checks can be performed on the RGB image byte blocks, depth image byte blocks, and aligned data frame sequences, respectively. Visual checks can be performed by loading RGB image byte blocks, depth image byte blocks, and aligned data frame sequences according to a timeline to synchronously display RGB-D images, robot status, end-effector trajectory, camera pose, motion information, and data source. This allows users to check data synchronization, image quality, trajectory continuity, and motion rationality. Data cleaning involves using preset algorithms to remove abnormal data frames, such as blurry images.
[0064] Then, semantic annotation is performed on the data segments (such as a single image frame) displayed after data cleaning and visualization to obtain the data to be verified. Annotation content includes, but is not limited to, keyframes, capture points, release points, success states, failure states, abnormal states, discard markers, and manually taken-over segments. For sequences containing data source identifiers, manually taken-over segments can be automatically filtered based on the source identifiers and confirmed or corrected in conjunction with the manually annotated information.
[0065] Next, after obtaining the data to be verified, the data can be replayed for verification. Through replay verification, it can be further confirmed that the image, state, action, pose, and annotation information are consistent in time.
[0066] For example, when playing back and verifying RGB image byte blocks, instead of expanding the RGB image byte blocks into a complete frame sequence all at once, the target window data corresponding to the target frame index is located and decoded from the RGB image byte blocks based on the received target frame index (i.e., target frame information) to obtain decoded data. The decoded data is cached, and when the number of cached data or cached bytes exceeds a threshold, historical frames are removed from the cache. The advantage of this setup is that it reduces the image data size while avoiding excessive memory consumption caused by generating complete video decoding during the visualization stage.
[0067] The robot teaching data generation method provided in this invention utilizes time-anchored data frame sequences to align multi-source data frame sequences, ensuring the temporal consistency of the final robot teaching data in terms of vision, state, and action. This method completes the spatial pose of the end effector, enabling the robot teaching data to simultaneously possess low-level joint states and high-level spatial pose representations. This method integrates data cleaning, semantic annotation, and playback verification into the same processing chain, allowing the processed data to undergo quality checks, segment selection, and annotation closure before entering the training set, reducing the inclusion of invalid or abnormal samples in the training set.
[0068] Example 3 Figure 3 This is a schematic diagram of a robot teaching data generation device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a data acquisition module 301, a data supplementation module 302, a data compression module 303, and a teaching data generation module 304, wherein: The data acquisition module is used to acquire a multi-source data frame sequence, wherein the timestamps of the data frames in any two types of data frame sequences in the multi-source data frame sequence are different; The data supplementation module is used to perform data completion processing on the remaining sequence based on the RGB-D image data frame sequence in the multi-source data frame sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. The data compression module is used to perform lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks, and to perform lossless image compression on the depth image data frame groups acquired by each camera to obtain depth image byte blocks. The teaching data generation module is used to generate robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
[0069] The robot teaching data generation device provided in this invention acquires multi-source data frame sequences containing different timestamps and performs data completion processing on the remaining sequences based on the RGB-D image data frame sequence, thereby obtaining a precisely aligned data frame sequence. This effectively overcomes the noise interference introduced by the temporal disorder of multi-channel data in the prior art, ensuring the temporal consistency of the "action-vision-state" causal chain, significantly improving the convergence speed of the model during imitation learning or robot policy training, and enhancing the generalization performance of the model in different scenarios. The data compression in this device significantly reduces the overall data volume of robot teaching data, thereby effectively alleviating the space pressure on the storage system and significantly reducing hardware storage costs. At the same time, the key dynamic motion features (motion residuals) during the robot's action execution process are preserved during the compression of RGB images, so that the compressed data can still provide rich kinematic information for subsequent models. When compressing depth images, a lossless compression strategy is adopted to ensure the integrity of the spatial geometric information and distance information contained in the depth data.
[0070] Optionally, the data supplementation module includes: The time anchor determination unit is used to determine the RGB-D image data frame sequence in the multi-source data frame sequence as the time anchor data frame sequence, wherein the multi-source data frame sequence includes a multi-channel image frame sequence, a robot body state and sensor feedback frame sequence, a control command frame sequence, a task language information frame sequence, and a data source identification frame sequence. The alignment and completion unit is used to map the remaining sequence to the time axis of the time anchored data frame sequence, and to perform data alignment and / or completion processing on the remaining sequence in the time axis according to the time anchored data frame sequence to obtain the first intermediate data frame sequence. The pose completion unit is used to complete the end effector spatial pose of the first intermediate data frame sequence according to the time-anchored data frame sequence, so as to obtain a time-synchronized aligned data frame sequence.
[0071] Furthermore, the step of completing the end effector spatial pose of the first intermediate data frame sequence based on the time-anchored data frame sequence to obtain a time-synchronized aligned data frame sequence includes: interpolating the joint states of the first intermediate data frames corresponding to the data frames in the time-anchored data frame sequence to obtain an interpolation result, and then performing forward kinematics calculations on the interpolation result to obtain a time-synchronized aligned data frame sequence; or, performing forward kinematics calculations on the joint states of the first intermediate data frames corresponding to the time-anchored data frame sequence to obtain the corresponding tool center point pose, and then interpolating the tool center point pose to obtain a time-synchronized aligned data frame sequence.
[0072] Optionally, the device may also include: The compression module is used to perform various data format conversions on the aligned data frame sequence and the RGB-D image data frame sequence before performing lossy image compression on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, so as to obtain corresponding conversion results respectively, and divide the conversion results into multiple batches to obtain a second intermediate data frame sequence, wherein the data format includes data format and data type; The writing module is used to stream the second intermediate data frame sequence into a preset storage space according to data batches. The data compression module includes: The lossy compression unit is used to perform lossy image compression on each group of RGB image data frames acquired by the camera in the preset storage space to obtain RGB image byte blocks.
[0073] Optionally, the data compression module includes: The compression unit is used to encode the RGB image data frame group acquired by each camera into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks, and to encode the depth image data frame group acquired by each camera into AVI byte blocks to obtain depth image byte blocks. The AVI byte blocks are data units that are encoded into AVI files using FFV1 and conform to the RIFF block structure.
[0074] Optionally, the device may also include: The first post-processing module is used to perform data cleaning and semantic annotation on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence respectively after generating robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence to obtain data to be verified, and to receive target frame information. The second post-processing module is used to decode the target window data corresponding to the target frame information from the compressed data if the data corresponding to the target frame information is compressed data in the data to be verified, so as to obtain decoded data. The compressed data is data in the RGB image byte block, and the target window data includes the compressed data corresponding to the target frame information and multiple frames of data adjacent to the corresponding compressed data. The sample generation module is used to visualize and output the decoded data in order to replay and verify the data to be verified, and obtain robot teaching data samples.
[0075] Furthermore, the step of encoding the RGB image data frame group acquired by each camera into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks includes: converting the RGB image data frame group acquired by each camera into YUV format data frames to obtain YUV frame groups, and performing optimal motion vector search on the prediction units of the YUV frame groups to obtain optimal motion vectors; using the optimal motion vectors to determine the spatial domain residual matrix, and based on the spatial domain residual matrix, obtaining a binary bitstream through frequency domain conversion processing, quantization processing, and entropy coding processing; and encapsulating the binary bitstream to obtain RGB image byte blocks.
[0076] The robot teaching data generation device provided in the embodiments of the present invention can execute the robot teaching data generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0077] Example 4 Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0078] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 and a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0079] Multiple components in electronic device 40 are connected to input / output (I / O) interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0080] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as methods for generating robot teaching data.
[0081] In some embodiments, the method for generating robot teaching data may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded into and / or installed on electronic device 40 via read-only memory (ROM) 42 and / or communication unit 49. When the computer program is loaded into random access memory (RAM) 43 and executed by processor 41, one or more steps of the method for generating robot teaching data described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the method for generating robot teaching data by any other suitable means (e.g., by means of firmware).
[0082] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0083] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0084] The computer equipment provided above can be used to execute the robot teaching data generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0085] Example 5 In the context of this invention, the computer-readable storage medium may be a tangible medium, and the computer-executable instructions, when executed by a computer processor, are used to perform a method for generating robot teaching data, the method comprising: Obtain a multi-source data frame sequence, wherein the timestamps of data frames in any two types of data frame sequences in the multi-source data frame sequence are different; Based on the RGB-D image data frame sequence in the multi-source data frame sequence, data completion processing is performed on the remaining sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. Lossy image compression is performed on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, and lossless image compression is performed on the depth image data frame group acquired by each camera to obtain depth image byte blocks. Robot teaching data is generated based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
[0086] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by, or in conjunction with, an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0087] The computer equipment provided above can be used to execute the robot teaching data generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0088] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the robot teaching data generation method as described in any of the above embodiments.
[0089] It is worth noting that in the embodiments of the robot teaching data generation device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0090] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for generating robot teaching data, characterized in that, include: Obtain a multi-source data frame sequence, wherein the timestamps of data frames in any two types of data frame sequences in the multi-source data frame sequence are different; Based on the RGB-D image data frame sequence in the multi-source data frame sequence, data completion processing is performed on the remaining sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. Lossy image compression is performed on the RGB image data frame group acquired by each camera to obtain RGB image byte blocks, and lossless image compression is performed on the depth image data frame group acquired by each camera to obtain depth image byte blocks. Robot teaching data is generated based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
2. The method according to claim 1, characterized in that, The step of performing data completion processing on the remaining sequence based on the RGB-D image data frame sequence in the multi-source data frame sequence to obtain an aligned data frame sequence includes: The RGB-D image data frame sequence in the multi-source data frame sequence is determined as the time-anchored data frame sequence, wherein the multi-source data frame sequence includes a multi-channel image frame sequence, a robot body state and sensor feedback frame sequence, a control command frame sequence, a task language information frame sequence, and a data source identification frame sequence. The remaining sequence is mapped to the time axis of the time-anchored data frame sequence, and the remaining sequence in the time axis is aligned and / or completed according to the time-anchored data frame sequence to obtain the first intermediate data frame sequence. Based on the time-anchored data frame sequence, the end effector spatial pose of the first intermediate data frame sequence is completed to obtain a time-synchronized aligned data frame sequence.
3. The method according to claim 2, characterized in that, The step of completing the end effector spatial pose of the first intermediate data frame sequence based on the time-anchored data frame sequence to obtain a time-synchronized aligned data frame sequence includes: Interpolate the joint state of the first intermediate data frame corresponding to the data frame in the time-anchored data frame sequence to obtain the interpolation result, and obtain the time-synchronized aligned data frame sequence by performing forward kinematics calculation on the interpolation result; or, Forward kinematics calculations are performed on the joint states of the first intermediate data frame corresponding to the time-anchored data frame sequence to obtain the corresponding tool center point pose. Then, by interpolating the tool center point pose, a time-synchronized aligned data frame sequence is obtained.
4. The method according to any one of claims 1-3, characterized in that, Before performing lossy image compression on each group of RGB image data frames acquired by the camera to obtain RGB image byte blocks, the method further includes: The aligned data frame sequence and the RGB-D image data frame sequence are subjected to various data format conversions to obtain corresponding conversion results. The conversion results are then divided into multiple batches to obtain a second intermediate data frame sequence. The data format includes data format and data type. The second intermediate data frame sequence is written into a preset storage space in batches. The step of performing lossy image compression on each group of RGB image data frames acquired by the camera to obtain RGB image byte blocks includes: Lossy image compression is performed on each group of RGB image data frames acquired by the camera in the preset storage space to obtain RGB image byte blocks.
5. The method according to claim 1, characterized in that, The process of performing lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks, and performing lossless image compression on the depth image data frame groups acquired by each camera to obtain depth image byte blocks, includes: Each camera's RGB image data frame group is encoded into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks, and each camera's depth image data frame group is encoded into AVI byte blocks to obtain depth image byte blocks. The AVI byte blocks are data units that are FFV1 encoded into depth images and constitute AVI files, conforming to the RIFF block structure.
6. The method according to claim 1, characterized in that, After generating robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence, the method further includes: Data cleaning and semantic annotation are performed on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence to obtain the data to be verified, and target frame information is received. If the data corresponding to the target frame information is compressed data in the data to be verified, then the target window data corresponding to the target frame information is decoded from the compressed data to obtain decoded data. The compressed data is data in an RGB image byte block, and the target window data includes the compressed data corresponding to the target frame information and multiple frames of data adjacent to the corresponding compressed data. The decoded data is visualized and output to enable playback verification of the data to be verified, thereby obtaining robot teaching data samples.
7. The method according to claim 5, characterized in that, The process of encoding each group of RGB image data frames acquired by each camera into MP4 byte blocks using H.264 / AVC to obtain RGB image byte blocks includes: The RGB image data frames acquired by each camera are converted into YUV format data frames to obtain YUV frame groups, and the prediction units of the YUV frame groups are searched for optimal motion vectors to obtain optimal motion vectors. The spatial domain residual matrix is determined using the optimal motion vector, and a binary code stream is obtained based on the spatial domain residual matrix through frequency domain conversion, quantization, and entropy coding. The binary bitstream is encapsulated to obtain RGB image byte blocks.
8. A device for generating robot teaching data, characterized in that, include: The data acquisition module is used to acquire a multi-source data frame sequence, wherein the timestamps of the data frames in any two types of data frame sequences in the multi-source data frame sequence are different; The data supplementation module is used to perform data completion processing on the remaining sequence based on the RGB-D image data frame sequence in the multi-source data frame sequence to obtain an aligned data frame sequence. The remaining sequence is the sequence in the multi-source data frame sequence other than the RGB-D image data frame sequence. The RGB-D image data frames are acquired by at least one camera, and the RGB-D image data frame sequence includes a group of RGB image data frames and a group of depth image data frames acquired by the at least one camera. The data compression module is used to perform lossy image compression on the RGB image data frame groups acquired by each camera to obtain RGB image byte blocks, and to perform lossless image compression on the depth image data frame groups acquired by each camera to obtain depth image byte blocks. The teaching data generation module is used to generate robot teaching data based on the RGB image byte blocks, the depth image byte blocks, and the aligned data frame sequence.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for generating robot teaching data according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for generating robot teaching data according to any one of claims 1-7.