Data acquisition system of robot

By integrating multiple sensors and using unified timestamp and coordinate conversion processing, the robot data acquisition system is solved, and the problem of low perception reliability of traditional robots in complex environments is realized, efficient fusion of multimodal data and environmental modeling is achieved, and the robot's robustness and decision-making capabilities are improved.

CN120245000APending Publication Date: 2025-07-04SHANGHAI ROBOT IND TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674277.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional robots rely on a single sensor for environmental perception, which is susceptible to noise interference and data loss in complex dynamic scenarios, resulting in a decrease in perception reliability. How to achieve multimodal data fusion is an urgent technical problem.

Method used

It provides a robot data acquisition system, integrating cameras, tactile sensors, force sensors, temperature sensors and audio acquisition equipment, integrating and annotating multimodal data through control modules, using unified timestamps and coordinate conversion processing to achieve spatial and temporal consistency of multimodal data, and generating high-precision environmental images and state labels.

Benefits of technology

It improves the robustness and decision-making reliability of the robot in complex scenarios, supports real-time optimization of task strategies, and enhances the scenario adaptability of target recognition and obstacle avoidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245000A_ABST
    Figure CN120245000A_ABST
Patent Text Reader

Abstract

The invention provides a data acquisition system of a robot, which is characterized by being applied to the robot, the robot comprises a mobile platform and a mechanical arm, the mechanical arm is mounted on the mobile platform, and the data acquisition system comprises a data acquisition module mounted on the mobile platform and the mechanical arm and used for acquiring different types of original data; the control module is used for controlling the mobile platform and the mechanical arm to execute corresponding operation according to a preset work task and acquiring different types of original data through the data acquisition module; the control module is also used for integrating different types of original data to generate an integrated data set; and the control module is also used for marking the integrated data set through the label identification model to generate a marked data set. Through the data acquisition system of the robot provided by the invention, data acquired by the robot can be fused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robots, and in particular to a data acquisition system for robots. Background Art

[0002] The widespread application of robotics in the fields of industry, medicine and services has put forward higher requirements for its environmental perception capabilities. Traditional robots rely on a single sensor (such as vision or lidar) for environmental perception, but are easily restricted by noise interference, data loss and other problems in complex dynamic scenes, resulting in reduced perception reliability.

[0003] To improve robustness, multimodal perception technology achieves multi-dimensional modeling of the environment by fusing heterogeneous sensor data such as vision, touch, and force. In multimodal perception technology, the diversity of sensor types enables robots to understand the environment from different angles. However, how to achieve multimodal data fusion is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The purpose of the present invention is to provide a data acquisition system for a robot, which can fuse the data collected by the robot.

[0005] In order to solve the above technical problems, the present invention is achieved through the following technical solutions:

[0006] The present invention provides a data acquisition system for a robot, which is applied to the robot. The robot comprises a mobile platform and a mechanical arm, and the mechanical arm is installed on the mobile platform. The data acquisition system comprises:

[0007] A data acquisition module, installed on the mobile platform and the mechanical arm, for collecting different types of raw data;

[0008] A control module, used to control the mobile platform and the mechanical arm to perform corresponding operations according to a preset work task, and to obtain the different types of raw data through the data acquisition module; the control module is also used to integrate the different types of raw data to generate an integrated data set;

[0009] The control module is also used to label the integrated data set through a label recognition model to generate a labeled data set.

[0010] In one embodiment of the present invention, the data acquisition module includes:

[0011] A camera, mounted on the mobile platform and the mechanical arm, for collecting image data;

[0012] A tactile sensor, mounted on the robotic arm, for collecting tactile data;

[0013] A force sensor, installed on the robotic arm, for collecting force data;

[0014] A temperature sensor, installed on the mobile platform, for collecting temperature data;

[0015] An audio acquisition device, installed on the mobile platform, for collecting audio data.

[0016] In an embodiment of the present invention, the camera includes:

[0017] At least one internal camera, installed on the end effector of the robotic arm, for collecting interactive image data of the end effector performing tasks;

[0018] At least two external cameras, installed on the mobile platform, for collecting surrounding image data;

[0019] Wherein, the control module is further configured to perform coordinate transformation processing on the interactive image data and the surrounding image data, and generate environmental image data according to the transformed interactive image data and surrounding image data.

[0020] In an embodiment of the present invention, the control module performs coordinate transformation processing on the interactive image data and the surrounding image data according to the following steps:

[0021] Collect image data of a calibration board with a preset checkerboard pattern through the internal camera and the external cameras, and obtain corresponding calibration image data;

[0022] Extract the checkerboard corner coordinates according to the calibration image data, and calculate the internal parameter matrices and distortion parameters of the corresponding internal camera and external cameras according to the checkerboard corner coordinates;

[0023] Collect image data of the calibration board in a preset reference coordinate system to obtain standard image data;

[0024] According to the overlapping part of the standard image data and the calibration image data, and the internal parameter matrices and distortion parameters corresponding to the internal camera and the external cameras, calculate the rotation matrix and translation vector of the coordinate system to be adjusted corresponding to the calibration image data of the internal camera and the external cameras relative to the reference coordinate system;

[0025] According to the rotation matrix and translation vector, perform coordinate transformation processing on the coordinate system to be adjusted corresponding to the interactive image data and the surrounding image data, and generate the interactive image data and surrounding image data after coordinate transformation.

[0026] In an embodiment of the present invention, the control module generates environmental image data according to the following steps:

[0027] Perform grayscale conversion and downsampling on the converted interactive image data and the surrounding image data, extract the low-frequency components through discrete cosine transform, and generate corresponding perceptual hash values;

[0028] According to the perceptual hash values, calculate the similarity between two adjacent frames of the converted interactive image data, and calculate the similarity between two adjacent frames of the converted surrounding image data;

[0029] Judge the similarity between two adjacent frames of the converted interactive image data: when the similarity is greater than the corresponding threshold, discard one frame of the converted interactive image data; otherwise, retain the two adjacent frames of the converted interactive image data;

[0030] Judge the similarity between two adjacent frames of the converted surrounding image data: when the similarity is greater than the corresponding threshold, discard one frame of the converted surrounding image data; otherwise, retain the two adjacent frames of the converted surrounding image data;

[0031] Generate environmental image data according to the retained converted interactive image data and surrounding image data.

[0032] In an embodiment of the present invention, the control module is used for:

[0033] Perform time alignment on the collected raw data of different types according to the timestamps;

[0034] Extract from the aligned raw data of different types according to a preset extraction frequency, and integrate to obtain each data in the dataset.

[0035] In an embodiment of the present invention, the control module is further used for:

[0036] Input each data in the dataset into the label recognition model to obtain corresponding label information;

[0037] Perform labeling processing on the corresponding data in the dataset according to the label information to generate a labeled dataset.

[0038] In an embodiment of the present invention, the control module is further used for:

[0039] Input the environmental image data in the data into the target recognition sub-model of the label recognition model for target recognition processing, and generate corresponding instance segmentation results and target detection results;

[0040] Input the audio data in the data into the semantic recognition sub-model of the label recognition model for semantic recognition to generate corresponding semantic label results;

[0041] Obtain the state label of the robot according to the tactile data, force sense data and temperature data in the data;

[0042] Comprehensively analyze the instance segmentation result, the object detection result, the semantic label result and the state label, and perform labeling processing on the corresponding data.

[0043] In an embodiment of the present invention, the robot further includes a remote operation module, and the remote operation module is used to generate corresponding operation instructions according to the operations of the operator and send them to the control module; the control module controls the robotic arm to execute corresponding work tasks according to the operation instructions.

[0044] In an embodiment of the present invention, at least one of 4G, 5G, and Lora is used for communication between the remote operation module and the control module.

[0045] As described above, the present invention provides a robot. Through the high-precision synchronous acquisition mechanism of multiple types of sensors, combined with unified timestamps and trigger-based data capture, the problem of poor spatio-temporal consistency of multi-modal data is solved, and the environmental modeling dimension is extended to the fusion of three-dimensional space and physical attributes; through the adaptive fusion of multi-modal data, the robustness is significantly improved compared with traditional single-modal perception systems in complex scenarios, supporting the robot to optimize task strategies in real time according to environmental changes, especially showing stronger scene adaptability and decision-making reliability in tasks such as target recognition and obstacle avoidance.

[0046] Of course, it is not necessary for any product implementing the present invention to achieve all the above advantages at the same time. Brief Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be generated based on these drawings without creative efforts.

[0048] Figure 1 It is a schematic diagram of a robot in an embodiment of the present invention.

[0049] In the figure: 100, mobile platform; 200, robotic arm; 300, data acquisition module; 400, control module; 500, remote operation module. Detailed Embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments generated by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] Please refer to Figure 1 , the present invention provides a data acquisition system for a robot. The data acquisition system can be applied to the robot, and the robot can perform preset work tasks in a complex environment. A data acquisition system for a robot can collect different types of raw data when the robot is performing work tasks, and process different types of raw data to obtain a processed data set. The robot can include a mobile platform 100, a robotic arm 200, a remote operation module 500, and a data acquisition system. The data acquisition system can include a data acquisition module 300 and a control module 400.

[0052] In some embodiments, the mobile platform 100 can be a wheeled standing desk, and the mobile platform 100 can support height adjustment and device movement. Through the wheeled design, the robot can quickly move to different work scenarios or experimental environments, reducing the time for reconfiguring equipment, and is particularly suitable for multi-task and multi-scenario environments. The height adjustment function enables the robotic arm to adapt to workbenches of different heights, providing greater flexibility, especially when facing objects with large height differences.

[0053] In some embodiments, the robotic arm 200 can be mounted on the mobile platform 100. The model of the robotic arm 200 can be a Franka Emika Panda 7-degree-of-freedom robotic arm. The 7-degree-of-freedom design of the robotic arm 200 enables it to perform complex and delicate operation tasks, such as multi-object grasping and task sequence operations, and is more flexible than traditional 6-degree-of-freedom robotic arms. The built-in force feedback system of the robotic arm 200 can sense the force and position of external objects in real time, adjust the grasping and operation actions, and ensure the stability and accuracy of the operation, especially in handling and assembly tasks. The robotic arm 200 can support standardized operation interfaces, which are easy to integrate with other devices, improving the scalability and flexibility of the system.

[0054] In some embodiments, the end effector of the robotic arm 200 can be a Robotiq 2F-85 gripper, which has functions such as high grasping force and intelligent control. The maximum grasping force of the Robotiq 2F-85 gripper is 85N, which can handle objects of various shapes and sizes and is suitable for various tasks such as handling, assembling, sorting, and packaging. The precise electric control system of the Robotiq 2F-85 gripper can adjust the grasping force according to the different materials, shapes, and weights of the objects to ensure that the objects are not damaged. The Robotiq 2F-85 gripper can be applicable to a variety of operating scenarios and can quickly adjust the shape and grasping angle of the jaws to adapt to different types of object grasping tasks.

[0055] In some embodiments, the data acquisition module 300 can be installed on the mobile platform 100 and the robotic arm 200 for acquiring different types of raw data. The data acquisition module 300 can include cameras, tactile sensors, force sensors, temperature sensors, audio acquisition devices, etc.

[0056] In some embodiments, cameras can be installed on the mobile platform 100 and the robotic arm 200 for acquiring image data. The cameras can include internal cameras and external cameras. The internal camera can be installed on the end effector of the robotic arm 200 for acquiring interactive image data of the end effector performing tasks. The external camera can be installed on the mobile platform 100 for acquiring surrounding image data.

[0057] In some embodiments, the model of the internal camera can be a Zed-Mini camera. The Zed-Mini camera can acquire interactive image data when the end effector performs work tasks, such as detailed information during operations such as grasping, placing, and moving.

[0058] In some embodiments, the model of the external camera can be a Zed 2 camera, and the number of them is at least two. The Zed 2 cameras can be installed on adjustable tripods on the mobile platform 100. The Zed 2 cameras can acquire image data of the robot's surrounding environment and provide environmental information with a global perspective. The Zed 2 cameras can provide real-time stereoscopic image streams, capturing details of the robot and its surrounding environment from multiple perspectives. By adopting stereoscopic vision technology, the Zed 2 cameras support 3D space reconstruction and high-precision depth perception, and are suitable for real-time perception of dynamic scenes and complex environments. By installing the Zed 2 cameras on adjustable tripods, it can quickly adapt to scene changes and provide accurate depth information.

[0059] In some embodiments, the Zed 2 camera captures the dynamic changes of the robot and its surrounding environment from a global perspective, including object positions, environmental dynamics, etc. The Zed-Mini camera captures the details of the interaction between the end effector and the object from a local perspective, such as grasping force, contact points, etc. The two Zed 2 cameras and the Zed-Mini camera work together to comprehensively capture the interaction between the robot and the environment, including robot movements, object position changes, environmental dynamics, etc.

[0060] In some embodiments, the Zed-Mini and Zed 2 cameras can simultaneously collect depth images and RGB images through the built-in binocular stereo vision system. The depth images and RGB images are aligned and fused at the pixel level in real time, so as to synchronously obtain the three-dimensional spatial position information and appearance visual features of the object.

[0061] In some embodiments, the depth image provides the distance information from each pixel point in the scene to the camera. Based on the depth information, the three-dimensional spatial information, contour shape, and spatial distribution of the object can be accurately modeled. Through the depth information, obstacles in the scene can be detected to support the robot's navigation. Based on the three-dimensional spatial information, a safe movement path can be planned. The depth image helps the robot understand the spatial layout of the scene and improve its environmental perception ability.

[0062] In some embodiments, the RGB image provides rich two-dimensional visual features such as color, texture, edges, etc. Based on the RGB image, the object can be accurately classified and recognized. The RGB image supports the continuous tracking of the object and helps the robot understand the movement trajectory of the object.

[0063] In some embodiments, the depth image and the RGB image can be aligned at the pixel level to ensure the consistency of the spatial position and visual features of each pixel point. By jointly analyzing the depth image and the RGB image, the three-dimensional spatial information and two-dimensional visual features can be utilized simultaneously, significantly enhancing the target detection and environmental perception capabilities.

[0064] In some embodiments, the tactile sensor can be installed on the end effector of the robotic arm 200 to collect tactile data. The tactile data can include information such as the texture, hardness, temperature, and contact force of the object surface. The model of the tactile sensor can be the XELA Robotics array-type triaxial force tactile sensor Uskin. The tactile sensor can provide high-resolution tactile information, accurately sense the texture, hardness, and contact force of the object surface, and can simultaneously sense the forces in three directions (X, Y, Z axes) to provide more comprehensive contact force information.

[0065] In some embodiments, a force sensor can be installed on the end effector of the robotic arm 200 to collect force data. The force data can include the interaction force (three-dimensional force) and torque (three-dimensional torque) between the robot end effector and the object. The model of the force sensor can be the OnRobot HEX-E / H six-axis force sensor. The force sensor can simultaneously sense forces in three directions (X, Y, Z axes) and torques in three directions (rotational torques about the X, Y, Z axes), providing comprehensive force information. The force sensor can help the robot achieve high-precision operation control, such as adjusting the force when grasping fragile objects or ensuring precise docking in assembly tasks. The force sensor can detect unexpected collisions between the robot's end effector and objects or the environment in real time, avoiding equipment damage or task failure.

[0066] In some embodiments, a temperature sensor can be installed on the mobile platform 100 to collect temperature data. The temperature data can be used to sense temperature changes on the object surface or in the environment. The model of the temperature sensor can be the WRNK-191 thermocouple temperature sensor. When the robot is handling high-temperature objects, the temperature sensor can monitor the object temperature in real time to ensure operation safety.

[0067] In some embodiments, an audio acquisition device can be installed on the mobile platform 100 to collect audio data. The audio data can be used to implement functions such as speech recognition, environmental sound analysis, and abnormal sound source localization. The audio acquisition device can assist in identifying background noise, detecting emergencies, and providing the robot with more comprehensive environmental perception capabilities. The model of the audio acquisition device can be the Hikvision DS-VM1 microphone. By collecting voice data, the audio acquisition device can achieve human-robot interaction, such as voice command recognition and voice feedback. The audio acquisition device can analyze the sound characteristics in the environment to help the robot understand the surrounding environment, such as identifying machine operation sounds, vehicle driving sounds, etc. In case of emergencies, the audio acquisition device can detect abnormal sounds and locate the sound source, such as fire alarms, glass breaking sounds, etc., providing early warning information for the robot.

[0068] In some embodiments, the control module 400 can be used to control the mobile platform 100 and the robotic arm 200 to perform corresponding operations according to a preset work task and obtain different types of raw data. The control module 400 is also used to integrate and process different types of raw data to generate an integrated data set.

[0069] In some embodiments, the control module 400 may control the mobile platform 100 and the robotic arm 200 to perform corresponding operations according to preset work tasks (such as grasping, placing, pushing, or assembling, etc.), and collect raw data from a variety of sensors (such as cameras, tactile sensors, force sensors, temperature sensors, audio acquisition devices, etc.). The control module 400 may include a Polymetis controller and a NUC server (Next Unit of Computing), and the Polymetis controller and the NUC server may be respectively responsible for the real-time control of the robotic arm and the acquisition, processing, and storage of data.

[0070] In some embodiments, the Polymetis controller can transmit motion commands in real time at a control frequency of 15 Hz to ensure the motion stability and accuracy of the robotic arm. The Polymetis controller can precisely synchronize the actions of each joint of the robotic arm to ensure the synchronization of task execution. The Polymetis controller can support parallel execution of multiple tasks, such as simultaneously performing a grasping task and a path planning task, to improve operation efficiency. The Polymetis controller can cooperate with the NUC server to immediately process operation commands and feedback, and provide real-time task execution information.

[0071] In some embodiments, the NUC server can receive data streams from multiple aspects such as cameras, robot control systems, and task feedback, and perform efficient data fusion and analysis. The NUC server can be responsible for storing the processed data to provide support for subsequent analysis and optimization. The NUC server can support a wide range of interface standards and can be seamlessly connected to other computing platforms or devices, facilitating the integration of multiple hardware and sensors.

[0072] In some embodiments, a variety of sensors (such as cameras, tactile sensors, force sensors, temperature sensors, audio acquisition devices, etc.) may collect data asynchronously, resulting in data inconsistency in the time dimension. Data misalignment will affect subsequent data processing, fusion, and analysis, and reduce the accuracy and reliability of the system. The control module 400 can perform timing alignment processing on different types of raw data through timestamps to ensure data consistency in the time dimension.

[0073] In some embodiments, the control module 400 can attach timestamps to each frame of image data or other sensor data through a timestamp mechanism, serving as the time reference for subsequent data processing and fusion. The control module 400 can adopt an accurate clock synchronization mechanism (such as CLOCK_REALTIME) to ensure that all sensor drivers uniformly obtain the system-level high-precision time. Based on the timestamps, the control module 400 can perform timing alignment on data from different sensors to generate an aligned dataset.

[0074] In some embodiments, taking image data as an example for illustration, when multiple cameras simultaneously collect images, due to acquisition latency or asynchronous operations, there may be a time misalignment in the image frames of different sensors. The control module 400 can attach a system-level high-precision timestamp during the acquisition of each frame of image. The timestamp can serve as the time reference for subsequent data processing and fusion, ensuring that the image frames of different cameras are aligned in the time dimension. The temporally aligned image data can accurately reflect the motion trajectory and behavioral changes of objects. The temporally aligned multi-view image data can support high-precision 3D reconstruction.

[0075] In some embodiments, in a robot system, the internal camera and the external camera may use different coordinate systems, resulting in the inability to directly fuse the data. Through coordinate transformation processing, it is possible to align the coordinate systems of the image data collected by different cameras, enabling the fusion and analysis of the image data in a unified coordinate system.

[0076] In some embodiments, the control module 400 can be used to perform coordinate transformation processing on the interactive image data and the surrounding image data, and generate environmental image data based on the interactive image data and the surrounding image data after coordinate transformation.

[0077] In some embodiments, the steps of the coordinate transformation processing may include: acquiring corresponding calibration image data by using the internal camera and the external camera to capture images of a calibration board with a preset checkerboard pattern. The calibration board refers to using a calibration board with a preset checkerboard pattern as the calibration target.

[0078] In some embodiments, the steps of the coordinate transformation processing may include: extracting the checkerboard corner coordinates based on the calibration image data, and calculating the internal parameter matrices and distortion parameters of the corresponding internal camera and external camera according to the checkerboard corner coordinates.

[0079] In some embodiments, corner extraction refers to extracting the corner coordinates of the checkerboard from the calibration image data. According to the corner coordinates, the internal parameter matrices K (such as focal length, center coordinates) and distortion parameters D (such as radial distortion, tangential distortion) of the internal camera and the external camera can be calculated. Among them, the internal parameter matrix K can be expressed as f x and f y refer to the equivalent focal lengths in the x and y directions of the image, whose unit is pixel, reflecting the scaling relationship between the camera imaging plane and the pixel coordinates; (c x , c y ) refers to the center coordinates of the image data.

[0080] In some embodiments, the steps of the coordinate transformation processing may include: capturing an image of the calibration board in a preset reference coordinate system to obtain standard image data.

[0081] In some embodiments, according to the overlapping part of the standard image data corresponding to the reference coordinate system and the calibrated image data, as well as the internal parameter matrix and distortion parameters corresponding to the internal camera and the external camera, calculate the rotation matrix and translation vector of the coordinate system to be adjusted corresponding to the calibrated image data of the internal camera and the external camera with respect to the reference coordinate system.

[0082] In some embodiments, the internal parameter matrix and distortion parameters corresponding to the reference coordinate system can be represented as K1 and D1 respectively, and the internal parameter matrix and distortion parameters corresponding to the coordinate system to be adjusted can be represented as K2 and D2 respectively. Through the internal parameters (K1, D1, K2, D2) of the two and the corner point correspondence relationship of the overlapping area, the rotation matrix R and translation vector T between the two coordinate systems can be obtained. The rotation matrix describes the rotation relationship of the coordinate system, and the translation vector describes the translation relationship of the coordinate system. The two together realize the conversion of the coordinate system.

[0083] In some embodiments, according to the rotation matrix and translation vector, perform coordinate transformation processing on the corresponding coordinate system to be adjusted to generate the interactive image data and surrounding image data after coordinate transformation. The coordinate transformation processing can be expressed as: P ref = R×P 待调 +T, where P ref and P 待调 respectively refer to the coordinates of the same point in the reference coordinate system and the coordinate system to be adjusted.

[0084] In some embodiments, after the control module obtains the original data such as image data, tactile data, force sense data, temperature data, audio data, etc., for the original data, data cleaning and data enhancement steps can be adopted to improve the quality of the original data.

[0085] In some embodiments, the data streams of multiple sensors (such as vision, touch, and force sense) are misaligned in the time axis due to different acquisition frequencies or delays. At this time, a global timestamp can be adopted, and based on the Linux system time or an accurate synchronization protocol (such as the PTP protocol), the clock sources of all sensors can be unified. Subsequently, different modalities of data (such as images, action instructions) can be aligned to the same time reference according to the timestamp to ensure the synchronization of cross-modal data in the time dimension.

[0086] In some embodiments, there may be duplicate frames in the continuously acquired image data (such as the images when the robot is stationary). For the image data with duplicate frames, the hash value of the image can be calculated to quantify the image similarity; if the similarity is higher than a threshold (such as 95%), it is determined as a duplicate frame and removed. Among them, the threshold can be adjusted according to the work task requirements (a high threshold retains more data, and a low threshold improves data diversity).

[0087] In some embodiments, mutations may occur in the robotic arm motion data (such as jumps caused by sensor noise). By calculating the joint angle change rate through a sliding window, if the change rate exceeds the physical limit (such as the maximum acceleration of the robotic arm), it is determined as an outlier and removed.

[0088] In some embodiments, the sensor may have data missing due to temporary failures, communication interruptions, or occlusions. For short-term missing data repair, methods such as linear interpolation, spline interpolation, and historical trajectory filling can be used for supplementation. Linear interpolation is applicable to smoothly changing data (such as temperature, joint angles in uniform motion). Spline interpolation is applicable to non-linearly changing data (such as trajectories in the acceleration / deceleration phase). Historical trajectory filling can be based on the time series model of historical data (such as ARIMA) to predict the missing segment. For long-term missing data processing, it can be directly marked as invalid data to avoid introducing noise.

[0089] In some embodiments, data augmentation can be achieved by methods such as synthetic noise injection, random masking, temporal perturbation, and spatial transformation. After data augmentation of the original data, the augmented data can be used to train the neural network model to improve the robustness of the trained robot.

[0090] In some embodiments, synthetic noise injection refers to simulating noise interference in the real environment to improve the anti-interference ability of the model. For example, for image data, Gaussian noise, motion blur, or brightness perturbation can be added. Another example is that for action data, random noise (such as small offsets of joint angles) can be superimposed on the operation instructions.

[0091] In some embodiments, the purpose of random masking is to prevent the model from overfitting specific regions and enhance the robustness to local features. Random masking can be classified into image region masking and action sequence masking. Image region masking refers to randomly occluding some regions in the image (such as 20% of the pixels). Action sequence masking refers to randomly discarding the action data of some time steps.

[0092] In some embodiments, the purpose of temporal perturbation is to enhance the adaptability of the model to changes in the time scale. Temporal perturbation can be classified into time scaling and random temporal offset. Time scaling refers to resampling the action sequence by accelerating or decelerating. Random temporal offset refers to fine-tuning the starting point of the data segment on the time axis.

[0093] In some embodiments, the purpose of spatial transformation is to enhance the robustness of the model to geometric deformations. For example, for image data, operations such as random rotation (±10°), translation (±5%), and scaling (90%-110%) can be performed. Another example is that for action data, the target pose can be randomly perturbed within the workspace of the robotic arm.

[0094] In some embodiments, when there are duplicate frames in the continuously acquired image data, the control module generates environmental image data according to the following steps: perform grayscale conversion and downsampling on the converted interactive image data and the converted surrounding image data, extract the low-frequency components through discrete cosine transform, and generate corresponding perceptual hash values.

[0095] In some embodiments, the purpose of grayscale conversion and downsampling is to reduce the computational complexity and retain the key structural information of the image. Grayscale conversion refers to converting an RGB image into a single-channel grayscale image to reduce color interference. Downsampling refers to reducing the image resolution (e.g., from 1920×1080 to 32×32) to reduce the amount of data while retaining low-frequency features (such as object contours). By retaining the key structural information of the image, it can provide a lightweight input for subsequent feature extraction and improve the processing efficiency.

[0096] In some embodiments, DCT low-frequency component extraction refers to converting the image data from the spatial domain to the frequency domain, separating the high-frequency (detail noise) and low-frequency (main structure) components. DCT low-frequency component extraction can retain only the low-frequency components (such as the upper left 8×8 coefficients), filter out high-frequency noise, and enhance the robustness of the hash to illumination and minor deformations. Subsequently, the hash value (such as a 64-bit binary code) can be calculated based on the DCT low-frequency coefficients, and a binary sequence can be generated by quantifying the coefficient mean.

[0097] In some embodiments, the steps of generating environmental image data may further include: calculating the similarity between the corresponding adjacent two frames of the converted interactive image data according to the perceptual hash value, and calculating the similarity between the corresponding adjacent two frames of the converted surrounding image data.

[0098] In some embodiments, the similarity can be calculated by comparing the Hamming distance (i.e., the number of different bits in the binary code) of adjacent two frames of image data. Specifically, A = 1 - (B / C), where A can represent the similarity, B can represent the Hamming distance, and C can represent the total number of hash bits. For example, if the Hamming distance is 4 and the total number of bits is 64, the similarity is 1 - 4 / 64 = 93.75%.

[0099] In some embodiments, the steps of generating environmental image data may further include: judging the similarity between the adjacent two frames of the converted interactive image data: when the similarity is greater than the corresponding threshold, removing one frame of the converted interactive image data; otherwise, retaining the adjacent two frames of the converted interactive image data.

[0100] In some embodiments, the steps of generating environmental image data may further include: judging the similarity between the adjacent two frames of the converted surrounding image data: when the similarity is greater than the corresponding threshold, removing one frame of the converted surrounding image data; otherwise, retaining the adjacent two frames of the converted surrounding image data.

[0101] In some embodiments, the size of the threshold can be set according to actual requirements and will not be limited here. For example, the threshold for interactive image data is relatively high (such as 95%), which can avoid accidentally deleting key frames of operations (such as the grasping moment). Another example is that the threshold for surrounding image data is relatively low (such as 90%), which can allow moderate changes in the environmental background and retain more dynamic information. If the similarity between two adjacent frames of interactive image data > the threshold (such as 95%), it is considered that the robotic arm is in a static or slightly moving state, and one of the frames is removed (such as retaining the first frame), which can reduce duplicate operation data (such as static images when continuously gripping an object). If the similarity between two adjacent frames of interactive image data > the threshold (such as 90%), it is considered that there is no significant change in the environment, and one of the frames is removed.

[0102] In some embodiments, the step of generating environmental image data may further include: generating environmental image data based on the retained interactive image data and surrounding image data.

[0103] In some embodiments, the filtered interactive image data (from the perspective of the robotic arm) and the surrounding image data (from the environmental perspective) can be aligned according to timestamps. The fusion method can be multi-view stitching and temporal superposition. Multi-view stitching refers to generating a panoramic environmental map (such as a close view of the robotic arm + a distant view of the environment); temporal superposition refers to constructing a dynamic environmental change sequence (such as the movement trajectory of an object).

[0104] In some embodiments, the control module 400 is further configured to perform time alignment on the collected different types of raw data according to timestamps; extract from the aligned different types of raw data at a preset extraction frequency and integrate to obtain each data in the dataset.

[0105] In some embodiments, the environmental image data, tactile data, force sense data, temperature data, and audio data can be subjected to temporal alignment processing according to timestamps to generate temporally aligned environmental image data, tactile data, force sense data, temperature data, and audio data.

[0106] In some embodiments, the preset extraction frequency means that every time a preset time period passes, extraction needs to be performed from the aligned different types of raw data. For example, the temporally aligned environmental image data, tactile data, force sense data, temperature data, and audio data can be segmented according to preset time nodes and preset time periods, so as to generate corresponding environmental image data, tactile data, force sense data, temperature data, and audio data at each preset time period; at each preset time period, the corresponding environmental image data, tactile data, force sense data, temperature data, and audio data are integrated to obtain each data in the dataset.

[0107] In some embodiments, for example, the preset time node can be 12:00, and the preset duration can be 10 minutes. At this time, the first data in the dataset can include all environmental image data, tactile data, force sense data, temperature data, and audio data within the range of 12:00 to 12:10, and the second data can include all environmental image data, tactile data, force sense data, temperature data, and audio data within the range of 12:10 to 12:20.

[0108] In some embodiments, the control module 400 is further configured to input each data in the dataset into the label recognition model respectively to obtain the corresponding label information; perform labeling processing on the corresponding data in the dataset according to the label information to generate a labeled dataset. The label recognition model can include a target recognition sub-model, a semantic recognition sub-model, and the like.

[0109] In some embodiments, the control module 400 is further configured to input the environmental image data in the data into the target recognition sub-model of the label recognition model for target recognition processing to generate corresponding instance segmentation results and target detection results.

[0110] In some embodiments, the control module 400 can complete the automatic annotation of environmental perception and action modalities through the target recognition sub-model, and finally output environmental perception annotation data and action modality annotation data. The environmental perception annotation data refers to the instance segmentation / target detection results (object category, position, contour). The action modality annotation data refers to the action classification label and abnormal state marker (normal / abnormal, action category).

[0111] In some embodiments, the target recognition sub-model can be selected from YOLOv8, SAM (Segment Anything Model), Mask2Former, etc. YOLOv8 is a target detection model that can balance speed and accuracy, and outputs the object category and bounding box (bbox). SAM is a general instance segmentation model that supports zero-shot segmentation. Mask2Former refers to a segmentation model based on Transformer, which is suitable for high-precision segmentation in complex scenarios. The output format of the target recognition model can be COCO format, LabelMe format, etc. The COCO format can record standardized fields ("class_id", "bbox", "segmentation", "score"). The LabelMe format can be compatible with custom annotation tools and support polygon vertex coordinates ("points").

[0112] In some embodiments, when obtaining action modality annotation data, the input data can be the time series of the joint positions or velocities of the robotic arm. Through a sliding window and a pattern matching algorithm, action modality annotation data can be output. The window size of the sliding window can be set according to the action duration (e.g., 0.5 seconds); the step size can determine the overlap rate of the action segments (e.g., 50% overlap to capture the boundaries). The pattern matching algorithm can include LSTM, Gated Transformer, etc. LSTM can capture temporal dependencies and classify action categories (such as "grasp", "place", "rotate", etc.). Gated Transformer can identify long-range associations through the self-attention mechanism and is suitable for complex action sequences. By automatically detecting parameters such as mutation values, zero drifts, NaN / saturation values, etc., abnormal labels can be marked for the data at the corresponding time points. The mutation value refers to calculating the derivative of the joint velocity (Δv / Δt), and if it exceeds the physical limit, it is marked as abnormal. The zero drift refers to the joint velocity being close to zero within a duration but the command not being stationary (such as sensor failure). The NaN / saturation value refers to the interruption of the data stream or the sensor exceeding the range (such as abnormal voltage).

[0113] In some embodiments, the control module 400 is further configured to input the audio data in the data into the semantic recognition sub-model of the label recognition model for semantic recognition, and generate corresponding semantic label results. Among them, the semantic recognition sub-model can include a speech conversion sub-model and an intent recognition sub-model. According to the pre-trained speech conversion sub-model, text conversion processing is performed on the audio data to generate corresponding text data; according to the pre-trained intent recognition sub-model, intent recognition processing is performed on the text data to generate corresponding semantic label results. The speech conversion model can be Whisper, Kaldi, etc. The speech conversion sub-model can segment the audio data into short-time frames (such as 25 ms / frame), extract Mel Frequency Cepstral Coefficients (MFCC) or log-Mel spectrograms as input features, and output text data. The intent recognition sub-model can remove stop words (such as "please", "of") from the text data and correct spelling mistakes (such as correcting "qingli zhuomian" to "clean the tabletop"). The intent recognition model can be a Large Language Model.

[0114] In some embodiments, the control module 400 is further configured to obtain the state label of the robot according to the tactile data, force data, and temperature data in the data. The state label can include the working state, working environment, etc. of the robot. For example, when the robotic arm of the robot grasps an object, the state label can include grasping success, grasping failure, etc. Another example is that when the robot is in a high-temperature environment, the state label can include high temperature, etc.

[0115] In some embodiments, the control module 400 is further configured to comprehensively analyze the instance segmentation result, the object detection result, the semantic label result, and the status label, and perform labeling processing on the corresponding data. Through spatio-temporal alignment and structured description, the control module 400 integrates the scattered multi-modal data (vision, touch, sound, etc.) in the robot task, the operation instructions, and the action trajectories into a complete event stream (Episode) with causal relationships along the time axis. The generated labeled data set not only supports task review and fault diagnosis, but also provides a high-quality data source for the "perception-decision-action" closed-loop training of the embodied intelligence model.

[0116] In some embodiments, the remote operation module 500 can be used to generate corresponding operation instructions according to the operations of the operator and send them to the control module 400. The control module 400 can control the robotic arm 100 to execute the corresponding work tasks according to the operation instructions. At least one of 4G, 5G, and Lora can be used for communication between the remote operation module 500 and the control module 400.

[0117] In some embodiments, the remote operation module 500 can be an Oculus Quest 2 device. The sensors and gesture tracking system built into the Oculus Quest 2 device can capture the natural motion states of the upper body and hands of the human body in real time, and transmit the pose information (including position, orientation, gestures, etc.) to the control module 400 wirelessly. After receiving the pose information, the control module 400 maps the human actions into control instructions that conform to the workspace constraints of the robotic arm 100 based on the coordinate mapping algorithm and the pose solution model, and then drives the robotic arm 100 to complete the corresponding actions. The remote operation module 500 can convert the natural actions of the human operator into precise motion instructions for the robotic arm through virtual reality interaction and low-latency transmission, while the control module 400 ensures the safety and feasibility of the actions through real-time perception and constraint mapping.

[0118] It can be seen that in the above solution, through the high-precision synchronous acquisition mechanism of various types of sensors, combined with unified timestamps and trigger-based data capture, the problem of poor spatio-temporal consistency of multi-modal data is solved, and the environmental modeling dimension is extended to the integration of three-dimensional space and physical attributes; through the adaptive fusion of multi-modal data, the robustness is significantly improved compared with traditional single-modal perception systems in complex scenarios, supporting the robot to optimize the task strategy in real time according to environmental changes, especially showing stronger scene adaptability and decision reliability in tasks such as target recognition and obstacle avoidance.

[0119] The embodiments of the present invention disclosed above are only used to help illustrate the present invention. The embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A data acquisition system for a robot, characterized in that, Applied to a robot, the robot includes a mobile platform and a robotic arm, the robotic arm is installed on the mobile platform, and the data acquisition system includes: A data acquisition module, installed on the mobile platform and the robotic arm, for acquiring different types of raw data; A control module, for controlling the mobile platform and the robotic arm to perform corresponding operations according to a preset work task, and obtaining the different types of raw data through the data acquisition module; The control module is further used for integrating and processing different types of raw data to generate an integrated data set; The control module is further used for performing labeling processing on the integrated data set through a label recognition model to generate a labeled data set.

2. The data acquisition system of the robot according to claim 1, characterized in that, The data acquisition module includes: A camera, installed on the mobile platform and the robotic arm, for acquiring image data; A tactile sensor, installed on the robotic arm, for acquiring tactile data; A force sensor, installed on the robotic arm, for acquiring force data; A temperature sensor, installed on the mobile platform, for acquiring temperature data; An audio acquisition device, installed on the mobile platform, for acquiring audio data.

3. The data acquisition system of the robot according to claim 2, wherein The camera includes: At least one internal camera, installed on the end effector of the robotic arm, for acquiring interactive image data of the end effector performing a task; At least two external cameras, installed on the mobile platform, for acquiring surrounding image data; Wherein, the control module is further used for performing coordinate transformation processing on the interactive image data and the surrounding image data, and generating environmental image data according to the transformed interactive image data and surrounding image data.

4. The data acquisition system of the robot according to claim 3, characterized in that The control module performs coordinate transformation processing on the interactive image data and the surrounding image data according to the following steps: Performing image acquisition on a calibration board with a preset checkerboard pattern through the internal camera and the external cameras to obtain corresponding calibration image data; Extracting checkerboard corner coordinates according to the calibration image data, and calculating the internal parameter matrix and distortion parameters of the corresponding internal camera and external cameras according to the checkerboard corner coordinates; Performing image acquisition on the calibration board in a preset reference coordinate system to obtain standard image data; Calculating the rotation matrix and translation vector of the coordinate system to be adjusted corresponding to the calibration image data of the internal camera and the external cameras relative to the reference coordinate system according to the overlapping part of the standard image data and the calibration image data, and the corresponding internal parameter matrix and distortion parameters of the internal camera and the external cameras; Performing coordinate transformation processing on the coordinate system to be adjusted corresponding to the interactive image data and the surrounding image data according to the rotation matrix and translation vector to generate the coordinate-transformed interactive image data and surrounding image data.

5. The data acquisition system of the robot according to claim 3, characterized in that, The control module generates environmental image data according to the following steps: Performing grayscale conversion and downsampling processing on the transformed interactive image data and surrounding image data, extracting low-frequency components through discrete cosine transform, and generating corresponding perceptual hash values; Calculate the similarity between the converted interactive image data of two adjacent frames corresponding to the perceptual hash value, and calculate the similarity between the converted surrounding image data of two adjacent frames corresponding to the perceptual hash value; Judge the similarity between the converted interactive image data of two adjacent frames: when the similarity is greater than the corresponding threshold, eliminate the converted interactive image data of a certain frame; otherwise, retain the converted interactive image data of the two adjacent frames; Judge the similarity between the converted surrounding image data of two adjacent frames: when the similarity is greater than the corresponding threshold, eliminate the converted surrounding image data of a certain frame; otherwise, retain the converted surrounding image data of the two adjacent frames; Generate environmental image data according to the retained converted interactive image data and surrounding image data.

6. The data acquisition system of the robot according to claim 1, characterized in that The control module is used for: Perform time alignment on the collected raw data of different types according to the time stamps; Extract from the aligned raw data of different types at a preset extraction frequency and integrate to obtain each data in the dataset.

7. The data acquisition system of the robot according to claim 1, characterized in that, The control module is also used for: Input each data in the dataset into the label recognition model respectively to obtain the corresponding label information; Perform labeling processing on the corresponding data in the dataset according to the label information to generate a labeled dataset.

8. The data acquisition system of the robot according to claim 7, characterized in that The control module is also used for: Input the environmental image data in the data into the target recognition sub-model of the label recognition model for target recognition processing to generate corresponding instance segmentation results and target detection results; Input the audio data in the data into the semantic recognition sub-model of the label recognition model for semantic recognition to generate corresponding semantic label results; Obtain the state label of the robot according to the tactile data, force sense data and temperature data in the data; Comprehensively analyze the instance segmentation result, the target detection result, the semantic label result and the state label, and perform labeling processing on the corresponding data.

9. The data acquisition system of the robot according to claim 1, wherein, The robot further includes a remote operation module, and the remote operation module is used to generate corresponding operation instructions according to the operations of the operator and send them to the control module; the control module controls the robotic arm to execute corresponding work tasks according to the operation instructions.

10. The data acquisition system of the robot according to claim 9, characterized in that, The remote operation module and the control module communicate with each other by at least one of 4G, 5G, and Lora.