Body intelligence multi-source data quality evaluation and verification method, device, medium and product

By converting multi-source data into a standardized format and performing integrity, availability, and consistency checks, combined with URDF models and large-scale model analysis, the challenges of standardization and validation of embodied intelligence data are solved, improving data management efficiency and model training effectiveness.

CN121188440BActive Publication Date: 2026-04-14SHANGHAI COOPERS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511734791.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-04-14
Estimated Expiration
2045-11-25

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of unified standards for multi-source data for embodied intelligence, a lack of data quality assessment systems, and difficulties in data visualization and physical verification, resulting in chaotic data management, low efficiency, and affecting model training performance.

Method used

Multi-source data is converted into a standardized format, and an integrity, availability, and consistency check framework is adopted. Combined with the URDF model, action and vision are synchronized and played back. The physical safety and high-level semantic correctness of the action are analyzed through a large model.

Benefits of technology

It enables efficient management and standardization of multi-source data, improves the objectivity and visualization capabilities of data quality assessment, ensures data reliability and consistency, and enhances the performance and stability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121188440B_ABST
    Figure CN121188440B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the field of information technology, and discloses a body intelligent multi-source data quality evaluation and verification method, equipment, medium and product, the method comprising: converting body intelligent multi-source data into standardized format data; quality evaluation is carried out on the standardized format data, the quality evaluation comprises at least one of integrity check, availability check and consistency check; a robot URDF model corresponding to the standardized format data is loaded, the URDF model is driven to move according to joint data, and corresponding visual data is played back synchronously, so as to realize linkage playback of model action and visual picture; the linkage playback is analyzed through a first large model and / or a second large model; wherein the first large model is used for judging the physical safety of robot action, and the second large model is used for judging the high-level semantic correctness of robot action and task annotation, so that a high-quality and reliable body intelligent data set is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to a method, device, medium and product for assessing and verifying the quality of embodied intelligent multi-source data. Background Technology

[0002] In recent years, embodied intelligence has been rapidly developing as a cutting-edge field at the intersection of artificial intelligence and robotics. Embodied intelligence emphasizes that intelligent agents achieve autonomous learning and evolution through dynamic interaction between their bodies and the environment, with its core being the deep integration of perception, action, and cognition. The visual-language-action (VLA) model, as a core technology of embodied intelligence, can effectively integrate visual information, language commands, and action decisions, and has become a research focus in both academia and industry.

[0003] The development of vision-language-action models heavily relies on high-quality, large-scale, and diverse robot operation data. However, building such a high-quality corpus faces numerous severe challenges in practice, significantly limiting the iteration speed and performance ceiling of embodied intelligence technology. These challenges are mainly reflected in the following three aspects:

[0004] The lack of standardization for multi-source data: Embody intelligent machine data comes from a wide range of sources, covering different types of robot platforms, diverse sensors, and various task scenarios. This heterogeneous nature leads to significant differences in data format, structure, naming conventions, coordinate system definitions, and timestamp accuracy. Currently, the industry lacks unified data standards and specifications, resulting in severe data fragmentation and hindering efficient data utilization and sharing.

[0005] A lack of data quality assessment systems is a significant issue: data quality determines the upper limit of model performance. Currently, the industry lacks a systematic, multi-dimensional data quality assessment framework. Existing data quality inspection methods mostly rely on manual sampling or simple script rules, which are inefficient and fail to comprehensively cover potential quality issues, with limited attention paid to deeper dimensions such as data integrity, usability, and multimodal consistency.

[0006] Data visualization and physical verification are difficult: Traditional datasets typically only provide raw videos and separate joint data files. Researchers find it difficult to intuitively verify whether the joint data matches the movements in the video, nor can they determine whether the movements are reasonable and effective in the physical world, such as whether the movements are smooth or conform to kinematic constraints. This "blind" data processing method may lead to erroneous or dangerous motion data being mixed into the training set, seriously "polluting" the model training process.

[0007] Therefore, there is an urgent need in this field for a systematic solution that can unify data standards, evaluate data quality from multiple perspectives, and provide intuitive verification methods to support the large-scale construction of high-quality embodied corpora. Summary of the Invention

[0008] One objective of this application is to provide a method, device, medium, and product for assessing and verifying the quality of embodied intelligent multi-source data, at least to address the problems of lack of standardization, deficiencies in data quality assessment systems, and difficulties in verifying the physical and dynamic nature of data in the prior art.

[0009] To achieve the above objectives, some embodiments of this application provide the following aspects:

[0010] This application provides a method for assessing and verifying the quality of embodied intelligence multi-source data, the method comprising:

[0011] The embodied intelligence multi-source data is converted into standardized format data; the standardized format data includes visual data, joint data, task annotation, camera parameters, and calibration information corresponding to the robot's URDF model;

[0012] The standardized format data is subjected to a quality assessment, which includes at least one of integrity checks, usability checks, and consistency checks.

[0013] After the quality assessment is completed, the robot URDF model corresponding to the standardized format data is loaded, and the URDF (Universal Robot Description Format) model is driven to move according to the joint data. The corresponding visual data is played synchronously to realize the linkage playback of model actions and visual images.

[0014] The linked playback is analyzed using a first major model and / or a second major model; wherein, the first major model is used to determine the physical safety of the robot's actions, and the second major model is used to determine the high-level semantic correctness of the robot's actions and the task annotations.

[0015] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.

[0016] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method described above.

[0017] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.

[0018] Compared with related technologies, the solution provided in this application successfully solves the problem of chaotic multi-source data management by defining a strict set of general data naming and storage format specifications. This greatly improves the management efficiency and standardization level of multi-source data, making the data storage structure clear, easy to access programmatically, and manage on a large scale. The proposed three-dimensional quality inspection framework of "integrity, availability, and consistency" transforms data quality assessment from inefficient and subjective manual sampling to an efficient and objective automated process. The quality inspection results are quantifiable and traceable, ensuring that the data used for model training is reliable and of high quality, which helps to improve the performance and stability of the final model. The data in this application is linked with the URDF playback function, transforming static and separate data files into dynamic and interactive 3D visualization scenes. Combined with action recognition models and multimodal large models, automated discrimination of action sequences that do not conform to physical laws or pose safety hazards is achieved. This greatly enhances the visualization and action verification capabilities of the data and improves the overall quality of the dataset. Thus, it provides a replicable template for building high-quality and reliable embodied intelligence datasets, strongly supporting the quality and safety testing and assessment of embodied intelligent robots throughout their entire lifecycle. Attached Figure Description

[0019] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0020] Figure 1 A flowchart illustrating an embodied intelligence multi-source data quality assessment and verification method provided as an exemplary embodiment of this disclosure;

[0021] Figure 2 An illustration of the linked playback function in an embodied intelligent data acquisition method provided as an exemplary embodiment of this disclosure;

[0022] Figure 3 An exemplary structural diagram of the electronic device provided for some embodiments of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Figure 1A flowchart illustrating an embodied intelligence multi-source data quality assessment and verification method provided as an exemplary embodiment of this disclosure, the method comprising:

[0025] S101. Convert the embodied intelligence multi-source data into standardized format data; the standardized format data includes visual data, joint data, task annotation, camera parameters, and calibration information corresponding to the robot's URDF model.

[0026] Specifically, the standardized format data has a standardized data directory structure, which can be stored with a complete robot operation trajectory (episode) as the basic unit. Each episode folder is named with a UUID to ensure uniqueness. The episode directory contains various modal data: camera data, joint data (proprio_stats), camera parameter files (parameters), etc. Outside the episode directory, corresponding directories are stored according to different semantic levels, including task, sub-scene, scene, task annotation information directory (task_info), and top-level series directory named "robot body model-end effector type-scene name".

[0027] The standardized format data includes visual data, joint data, task annotations, camera parameters, and calibration information corresponding to the robot's URDF model.

[0028] The visual data includes multiple RGB video streams (.mp4 encoded in h264 format) and depth map videos (.mkv compressed in ffv1 format), corresponding to different viewpoints (such as left hand, right hand, head, etc.).

[0029] Joint data is stored in the HDF5 format file proprio_stats.hdf5, which stores action commands and sensor state information, covering the position, velocity, torque and other information of joints such as end effectors, head, arms, legs and waist. Each joint data field contains names corresponding to the URDF model.

[0030] The camera parameters include multiple intrinsic and extrinsic parameter files (stored in .json format). The intrinsic parameters include focal length, optical center coordinates, distortion coefficients, etc., while the extrinsic parameters include parent / child coordinate system definitions and translation / rotation parameters (stored in quaternion format).

[0031] Task annotations are stored in a .json file in the task_info directory, containing metadata such as episode identifier, scene description, and subject model; additionally, it contains sub-task annotation information for each step, detailing the start and end frames, action description, and skill type of each atomic skill.

[0032] The robot model and calibration information are stored in the URDF directory, which includes the URDF model file corresponding to the data and the robot calibration information (robot_calibration.json). Key calibration parameters such as the homing_offset and limit range (range_min, range_max) of each joint are clearly defined for physical verification.

[0033] S102. Perform a quality assessment on the standardized format data, wherein the quality assessment includes at least one of integrity check, usability check, and consistency check.

[0034] Specifically, this step involves automatically and quantitatively checking the data that has been converted to a unified format for three dimensions (or at least one of them): integrity, availability, and consistency. This efficiently identifies and filters out "dirty data" in the dataset that is missing, invalid, misaligned, or does not conform to physical constraints, thereby ensuring the high quality and reliability of the data used for subsequent model training.

[0035] S103. After the quality assessment is completed, load the robot URDF model corresponding to the standardized format data, drive the URDF model to move according to the joint data, and synchronously play the corresponding visual data to realize the linkage playback of model actions and visual images.

[0036] Specifically, in this step, firstly, a Unified Robot Description Format (URDF) model file corresponding to the current data to be verified is loaded. This URDF file serves as the robot's physical skeleton and kinematic model, containing information such as geometry, mass, inertia, and limits for all links and joints. Then, joint data (including position, velocity, and torque) is extracted from the standardized data and used as input to drive the loaded URDF model to move in a virtual or simulation environment. The model's motion is played synchronously with the originally acquired visual data (such as RGB video), thus transforming static robot trajectory data into a dynamic and intuitive 3D visualization scene. Through this synchronized playback, users can directly observe whether the virtual robot model's movements are consistent with the movements in the actual scene video, overcoming the shortcomings of traditional methods that lack intuitive verification of the physical properties of the data.

[0037] S104. Analyze the linkage playback using the first major model and / or the second major model; wherein, the first major model is used to determine the physical safety of the robot's actions, and the second major model is used to determine the high-level semantic correctness of the robot's actions and the task annotations.

[0038] Specifically, the linked playback generates a video stream that can be processed by the model. This video stream is then fed into two pre-trained key models for analysis: the first model focuses on determining the physical safety of the robot's actions, such as identifying dangerous movements like high-frequency tremors, violent shaking, excessive joint confinement, self-collision, or uncontrolled arm swinging; the second model is responsible for determining the high-level semantic correctness of the robot's actions relative to the original task annotations, analyzing the degree of semantic matching between the model's output action descriptions and the task objectives, thereby verifying the data's logical correctness within the task. Through the collaborative analysis of these two models, this method achieves automated assisted discrimination of data at both the physical and task-level dimensions, greatly improving the efficiency and objectivity of data verification, and automatically completing the verification without manual analysis.

[0039] In the above embodiments, by defining a strict set of general data naming and storage format specifications, the problem of chaotic multi-source data management was successfully solved, greatly improving the management efficiency and standardization level of multi-source data, making the data storage structure clear, easy to access programmatically, and manage on a large scale. The proposed three-dimensional quality inspection framework of "integrity, availability, and consistency" transforms data quality assessment from inefficient and subjective manual sampling to an efficient and objective automated process. The quality inspection results are quantifiable and traceable, ensuring that the data used for model training is reliable and of high quality, which helps to improve the performance and stability of the final model. Furthermore, static and separate data files are transformed into dynamic and interactive 3D visualization scenes. Combined with action recognition models and multimodal large models, automated discrimination of action sequences that do not conform to physical laws or pose safety hazards is achieved. This greatly enhances the visualization and action verification capabilities of the data and improves the overall quality of the dataset. Thus, it provides a replicable template for building high-quality and reliable embodied intelligence datasets, and strongly supports the quality and safety testing and assessment of embodied intelligent robots throughout their entire life cycle.

[0040] In one embodiment, the integrity check specifically includes verifying the directory hierarchy, naming format, and existence of key files of the standardized format data through path recursion traversal and / or regular expression matching.

[0041] Specifically, the directory hierarchy, naming format, and existence of key files in the standardized format data are automatically verified through path recursion traversal and / or regular expression matching techniques. For example, it verifies whether each episode folder is named with a UUID and whether it contains key subdirectories and files such as camera, proprio_stats.hdf5, and parameters. Regular expressions ensure that filenames and directory names strictly adhere to preset naming conventions, preventing subsequent processing interruptions due to missing files or naming errors.

[0042] For example, if the directory hierarchy has a nested relationship of "series→task_info→scene→subscene→task→episode", regularization matching and rule validation are used to check whether the naming of each directory level conforms to the preset format; the series level must match the naming pattern of "robot model-end effector-scene name"; the task level must contain information such as "action name-size_number_duration". If they do not conform, they are marked as format abnormal.

[0043] For example, verify that each episode (single data directory) must contain four subdirectories: camera / (including the depth / video subdirectory), parameters / , proprio_stats / , and audio / , and that key files (such as depth video, RGB video, camera parameter JSON, HDF5 joint data, and WAV audio) exist in each subdirectory, and that their names match the camera / sensor type (e.g., the prefix of the camera parameter file must match the corresponding camera name). If they do not meet the requirements, mark them as missing key files.

[0044] In one embodiment, the integrity check further includes performing content integrity checks on the joint data and the task annotations through field parsing and mapping verification.

[0045] Specifically, this validation targets the internal structure of the data file (such as HDF5 or JSON), ensuring that all expected key fields (such as joint position, velocity, torque, camera intrinsic parameters, mission start and end frames, etc.) exist, and that their data type, dimensions, and quantity match the standardization requirements. In particular, the field names of the joint data are validated to ensure accurate mapping and alignment with the corresponding URDF model joint names, guaranteeing data parsability. These two levels of validation effectively filter out incomplete or incorrectly formatted data, laying the foundation for subsequent usability and consistency checks.

[0046] For example, the h5py library is used to load robot joint data proprio_stats.hdf5 and proprio_stats_original.hdf5, and the existence of preset fields (such as joint position, velocity, torque, timestamp, etc.) is checked. If they do not exist, they are marked as missing joint data. Then, the corresponding field data is parsed, and the field dimensions are checked to see if they match the robot hardware configuration (e.g., the action / joint / position fields of the control arms of a 7-axis dual-arm robot should have 14 dimensions in radians). If they do not match, they are marked as abnormal joint data.

[0047] Parse the JSON tag file under task_info / and check whether the annotation fields (such as episode_id, start_frame, end_frame, action_label, etc.) of each data episode are complete. If they are not complete, mark them as missing annotations. At the same time, verify whether episode_id is completely consistent with the corresponding data's UUID directory name. If they are not complete, mark them as tags and data mismatch.

[0048] In the above embodiments, by combining path traversal, regular expression matching, and field parsing mapping verification, comprehensive automated verification of embodied intelligent multi-source data at both the structural and content levels is achieved. This solves fundamental problems easily overlooked in traditional data quality inspection, such as missing files, non-standard naming, and missing key metadata fields. It greatly improves the efficiency and automation of data preprocessing, reducing reliance on manual file-by-file checks; it ensures that all data passing this round of checks is highly consistent and complete in format and structure, providing a reliable and standardized input prerequisite for subsequent more complex usability and consistency checks; by verifying the mapping relationship between joint data fields and URDF model names, it fundamentally guarantees data parsability, laying a solid foundation for subsequent URDF linkage verification and avoiding interruptions to the entire verification process due to data structure errors.

[0049] In one embodiment, the availability check specifically includes:

[0050] Invalid frames, stuttering, or depth anomalies in the visual data are identified through color histogram analysis, inter-frame difference calculation, or depth value range statistics.

[0051] Specifically, to address the availability of visual data, various image processing methods are employed to identify common anomalies: color histogram analysis is used to identify potentially black or white frames (i.e., invalid exposure), inter-frame difference calculation is used to detect whether there are stuttering or frozen frames in the video that have not changed for a long time, and depth value range statistics are used to detect whether there are a large number of invalid depth pixels in the depth image that exceed the physically reasonable range, thereby ensuring the input quality of visual data.

[0052] For example, color histogram analysis; statistically analyzing the distribution of RGB three-channel pixel values ​​for each frame. A normal video's three-channel histogram should show a continuous distribution. Abnormal data includes black and white screens where R=G=B and are concentrated at 0 (black) or 255 (white); green screens where the G channel pixel value is close to 255 and the R / B channel is close to 0. For example, if a video has five consecutive frames where the G channel pixel value is around 255 and the R / B channel values ​​are all 0±1, it is marked as an invalid video.

[0053] Stuttering detection: Calculate the pixel difference between adjacent frames (e.g., mean absolute difference). Normal frames usually have a larger difference, while stuttering or repetitive frames are generally less than a set threshold (e.g., 5). If the pixel difference of 5 consecutive frames is less than the set threshold, the video is judged to be stuttering.

[0054] Depth map anomaly detection: Combining threshold filtering and geometric fitting to identify invalid values ​​and noise.

[0055] Invalid value detection: Set a reasonable physical depth range (e.g., 0.3m-10m), and count the percentage of pixels that exceed the range. If the percentage of pixels that exceed the range is consistently greater than the threshold (e.g., 30%) in multiple consecutive frames of data, it is determined that the depth data is abnormal.

[0056] Noise Detection: Within the effective depth range, the interquartile range (IQR) method is used to statistically analyze the proportion of outlier depth values. If the proportion of outliers in multiple consecutive frames exceeds a threshold (e.g., 30%), the depth data is considered abnormal. For example, if the effective depth values ​​of a certain segment of depth video data are mainly concentrated between 0.6m and 10m, but IQR analysis shows that outliers (e.g., pixels that abruptly change to 0.2m or 12m) account for 38% of the data segment (exceeding the 30% threshold), the depth data is considered abnormal.

[0057] In one embodiment, the availability check further includes performing a motion limitation check by comparing the joint data with the calibration information.

[0058] In one embodiment, the availability check further includes detecting instantaneous jumps or noise anomalies in the speed or torque of the joint based on the joint data through time-series analysis and / or frequency domain analysis.

[0059] Specifically, regarding the availability of joint data, motion limit checks and / or instantaneous anomaly detection can be performed. The motion limit check involves extracting the joint data and comparing it with the calibration information (such as range_min and range_max) to confirm that the robot has not exceeded its preset physical motion limits at any point in the trajectory. This is crucial for ensuring the safety of robot operation.

[0060] For example, read the robot calibration information (robot_calibration.json) and obtain the range of motion of each joint (e.g., the range of motion of the left arm joint 1 is between -1.5 and 1.5). Iterate through the various joint data fields in the robot joint data proprio_stats.hdf5 (e.g., the observation data of the two arm joints "state / joint / position", the observation data of the head joint "state / head / position", etc.). If there are values ​​that exceed the range, it is determined that the joint is out of limit.

[0061] The instantaneous anomaly detection: Based on the joint data, through time-series analysis (such as calculating the velocity and torque change rate of adjacent frames) and / or frequency domain analysis (such as performing FFT on the sequence to detect high-frequency noise components), it is possible to detect whether there are abnormal instantaneous jumps or continuous high-frequency noise in the velocity or torque of the joint, effectively eliminating "dirty data" caused by sensor failure or physical interference.

[0062] In one embodiment, the availability check further includes checking the invertibility of the intrinsic parameter matrix of the camera parameters, the normalization degree of the extrinsic quaternions, and the validity of the intrinsic and extrinsic parameters.

[0063] Specifically, this embodiment is used to check the camera parameter file, verify the mathematical invertibility of the intrinsic parameter matrix, and verify whether the extrinsic quaternion satisfies the normalization condition. At the same time, it checks whether the values ​​of the intrinsic and extrinsic parameters are within a reasonable range of validity.

[0064] This can be achieved by verifying whether the intrinsic parameter matrix of each camera (such as K in hand_right_intrinsic_params.json) is invertible (determinant ≠ 0). For example, determine whether the principal point (cx, cy) is in the center region of the image, with a deviation of no more than 6% of the image width / height (e.g., for a 1920×1080 resolution image, cx should be between 900-1020 and cy between 500-600). If the calculated value does not meet the conditions, mark the intrinsic parameters as invalid. Verify whether the extrinsic quaternions of each camera (e.g., rotation in hand_right_extrinsic_params.json) are normalized (magnitude ≈ 1). If the value does not meet the conditions, mark the extrinsic parameters as invalid. Use OpenCV's undistort function to distort the RGB image. If the image edges are distorted after distortion removal (e.g., straight lines become curves), the intrinsic parameters are invalid. Visualize the pose of the camera relative to the robot link using Open3D (transform the camera coordinate system to the robot link coordinate system through extrinsic parameters). If the camera position conflicts with the physical installation position (e.g., the head camera and the two hand cameras are at the same height), mark the extrinsic parameters as invalid.

[0065] In this embodiment, signal processing techniques such as frequency domain analysis are introduced to efficiently and automatically remove high-frequency noise caused by sensor malfunctions from joint data. Invalid frames and stuttering in visual data are identified and filtered through color histogram / inter-frame difference calculations, thus completely solving the problem of difficulty in detecting noisy data in traditional methods. Simultaneously, by rigorously comparing the joint data with the motion limits of the URDF model, this embodiment ensures that all robot operation trajectories are physically safe and reasonable. Furthermore, the validity verification of camera parameters (such as intrinsic reversibility and extrinsic parameter normalization) mathematically guarantees the accuracy of all subsequent multimodal alignment and 3D spatial calculations, improving the training value and reliability of the dataset and ensuring high-quality training usability.

[0066] In one embodiment, the consistency check specifically includes:

[0067] Temporal consistency check: Determine the alignment degree between the number of visual data frames and the number of joint data frames;

[0068] Spatial consistency check: Determine the spatial deviation between the end position in the visual data and the end position in the joint data through geometric calculations;

[0069] Semantic consistency check: The video content is semantically validated by calling the third major model to determine the degree of semantic matching between the video content and the task-annotated text.

[0070] Specifically, a time consistency check can be performed first; this check is used to determine the alignment of the frame sequence of the visual data with the frame sequence of the joint data in terms of total duration and timestamps. The timestamps of the two modal data are compared to identify whether there are any asynchronicities, data loss, or time drift during acquisition, ensuring that each movement of the robot precisely corresponds to the correct visual image.

[0071] For example, count the total number of frames in RGB video (such as hand_right_color.mp4) and the total number of frames in depth video (such as hand_right_depth.mkv), and compare them with the length of "timestamps" in the joint data proprio_stats.hdf5 (i.e., the number of joint data frames). If the difference is greater than 5 frames, it is marked as a multimodal duration inconsistency.

[0072] Alternatively, the unified timestamp (timestamps field) of the joint data can be compared with the original timestamps of each sensor (e.g., hand_right_color_mp4_timestamps) to calculate the time difference between corresponding frames (e.g., the difference between the joint timestamp of frame i and the RGB timestamp of frame i). If the difference exceeds a set threshold (±20ms), it is marked as a multimodal time deviation inconsistency. For example, if the joint timestamp of frame 500 is 1755829013007, and the corresponding RGB timestamp is 1755829013037, the deviation is 30ms, which is greater than the set threshold and is marked as a multimodal time deviation inconsistency.

[0073] Secondly, a spatial consistency check can be performed; through complex geometric calculations, the spatial deviation between the robot end effector position identified or located in the visual data and the robot end effector position recorded in the joint data can be determined. If this deviation is too large, it indicates that there is a serious calibration error or kinematic inconsistency in the data.

[0074] Finally, a semantic consistency check can be performed. To verify the high-level semantic correctness of the data, the video content is semantically validated by calling a third major model, namely the Visual-Language Model (VLM). For example, firstly, the start and end frames (e.g., start_frame=200, end_frame=300) and action descriptions (e.g., "grab the package and put it in the box") of a certain subtask are extracted from the JSON annotations of the data; then, the 200-300 frame segment of the head-mounted camera video is extracted and input into the VLM model (e.g., Qwen3-VL), and the question "What is the core action of the robot in the video?" is asked; if the semantic similarity between the model output and the annotation description is greater than 80% (e.g., "put the package in the box" matches "grab the package and put it in the box"), then the annotation is consistent with the video; otherwise, it is marked as inconsistent with the video content.

[0075] In this embodiment, firstly, a time consistency check ensures precise synchronization of data streams from different sensors, completely eliminating the "spatiotemporal misalignment" problem caused by asynchronous acquisition and guaranteeing the accuracy of the data in terms of time sequence. Secondly, by utilizing camera parameters and geometric calculations, 3D spatial alignment verification of visual perception results and robot kinematic records is achieved. This high-precision physical verification capability can effectively identify and eliminate kinematically inconsistent data caused by calibration, model, or sensor errors, which is difficult to achieve with traditional methods. Finally, a multimodal large model is introduced for semantic consistency checking, enabling this method to verify the logical matching degree between the robot's actual actions and the task description from a high-level cognitive dimension, thereby transforming subjective and time-consuming manual annotation and review into objective and efficient machine semantic verification. In summary, this embodiment ensures high accuracy, high reliability, and high semantic value of data at the multimodal fusion level, fundamentally improving the effectiveness of embodied intelligence model training data.

[0076] In one embodiment, the step of analyzing the linked playback using the first major model and / or the second major model specifically includes:

[0077] The aforementioned linked playback process is recorded.

[0078] The screen recording data is input into the first large model, which is a motion recognition large model that identifies dangerous actions such as high-frequency shaking, joint self-collision, or uncontrolled arm swinging, and is used to determine physical safety.

[0079] The screen recording data and the task tags in the original annotations are input into the second large model, which is a multimodal large model. The semantic similarity between the action description output by the multimodal large model and the task tags is determined.

[0080] Specifically, firstly, after the linked playback (S103) is completed, the linked playback process is recorded. The dynamic, real-time 3D simulation environment and the synchronized raw visual data are solidified into a unified, standard screen recording data stream (such as a high frame rate video file), which serves as a stable input for subsequent model analysis. This screen recording data contains all the motion details and visual context of the robot model when executing the raw joint data.

[0081] Secondly, the screen recording data is input into the first major model, specifically a motion recognition model. This model, professionally trained, is capable of recognizing the underlying physical motion characteristics displayed in the video. It can accurately identify a series of dangerous actions such as high-frequency jitter, self-collision caused by joints exceeding limits, or uncontrolled arm swinging due to loss of control. Based on the recognition results, the model outputs a quantified physical safety score or binary judgment, thereby automatically determining the physical safety of the operation trajectory and effectively preventing erroneous motion data that does not conform to kinematic or dynamic constraints from entering the training set.

[0082] Finally, the screen recording data is simultaneously input into the second large model, specifically a multimodal large model (VLM), with the task labels (i.e., text instructions) from the original annotations also serving as input. This VLM leverages its powerful video understanding and text reasoning capabilities to perform detailed semantic analysis on the robot's action sequence in the screen recording and outputs a description of the actions. This model-generated description is then compared to the original task label text using semantic similarity calculations (e.g., employing text embedding methods such as BERT or Sentence-BERT) to determine the similarity score. If the score is below a preset threshold, it indicates that the robot's actual actions do not match the semantics of the task label (e.g., the instruction is "pick up the water cup," but the actual action is "put down the water bottle"), and the data is marked as semantically incorrect.

[0083] by Figure 2 For example, Figure 2 The image shows the effect of the data and URDF linkage visualization playback function. From left to right, it shows the linkage effect of multimodal data and URDF model in the "robot picking express packages and placing them in the designated box".

[0084] This embodiment can be further developed based on the open-source URDF visualization tool urdf-loaders, extending its functions to include video and joint data loading, multi-view video synchronous playback, dynamic mapping of joint trajectories and URDF joints, and joint playback. It also further supports recording the playback process as a standardized video stream. Based on this, an action recognition model or a visual-language large model (VLM) is introduced to automatically analyze the screen recording video, determining whether the robot exhibits dangerous actions (such as high-frequency shaking, joint self-collision, uncontrolled arm swinging, etc.) and whether the actions are consistent with the actual semantics, thereby achieving automatic data quality assessment without human intervention. This function provides an intuitive and dynamic means of verifying data quality and transforms the visualization results into automatically assessable quality signals, thereby reducing manual review costs. Specifically, it includes:

[0085] Load the zero and limit parameters from the URDF model and robot calibration information, and initialize the joint configuration of the robot model;

[0086] Upload and parse multimodal data in an episode directory, and extract the joint data corresponding to each time step;

[0087] Drive the URDF model to move according to joint data, and play the visual data with the corresponding timestamp at the same time to realize the linkage display of "model action - visual image", and support automatic screen recording of the playback process;

[0088] Visualized screen recording and automated quality assessment: The screen recording process of the linkage playback is recorded, and two-stage automated assessment is performed through action recognition model and multimodal large model to ensure that the corresponding action takes into account both the underlying physical security and the correctness of the high-level task.

[0089] Dangerous Action Detection and Analysis: Input screen recording data into a motion recognition model (such as mmaction2). The model automatically performs time-series detection and behavior analysis, automatically identifying dangerous or non-physical behaviors such as high-frequency shaking, joint self-collision, and uncontrolled arm swinging. If any anomaly is detected, the data segment is immediately marked as low-quality data.

[0090] Semantic consistency verification: Videos that pass the first-stage screening are further input into an open-source multimodal large model (such as Qwen3-VL). Combined with task labels from the original JSON annotations (e.g., "grab the package and put it in the box"), prompts (e.g., "Please describe what task the robot completed") guide the model to understand the semantics of the actions performed by the URDF model in the video. If the action description output by the model is significantly semantically inconsistent with the task label (e.g., it should be "grab and place" but is identified as "grab failed"), it is judged as a high-level semantic bias and marked as invalid data.

[0091] In this embodiment, by introducing advanced dual-model collaborative analysis, the physical safety blind spots and high-level semantic verification challenges in embodied intelligence data verification are completely solved. First, by utilizing an action recognition model to analyze linked playback screen recordings, this embodiment achieves automated, non-contact judgment of the physical safety of robot actions. It can accurately identify dangerous actions such as high-frequency jitter and self-collision, effectively preventing data with safety hazards from polluting the training set and greatly improving the safety and robustness of future models. Second, by introducing a multimodal large model to compare and analyze screen recordings and task labels, this embodiment successfully achieves objective and efficient verification of the high-level semantic correctness of data, transforming the originally time-consuming and subjective manual annotation and review into automated machine verification. This dual mechanism, by elevating data verification capabilities to both the physical and semantic levels, fundamentally ensures the high reliability of data in terms of functionality and security, providing an advanced verification method for the construction of large-scale, high-quality embodied intelligence corpora.

[0092] In one embodiment, the step of detecting instantaneous jumps or noise anomalies in the velocity or torque of the joint based on the joint data through time-series analysis and / or frequency domain analysis specifically includes:

[0093] The time-series analysis calculates the joint angular velocity and torque change rate using adjacent frames, and compares the change rate with a preset first threshold to determine the instantaneous jump anomaly in the joint data;

[0094] The frequency domain analysis involves performing a fast Fourier transform on the joint position or torque sequence, analyzing the energy proportion of high-frequency components, and comparing the energy proportion with a preset second threshold to determine noise anomalies in the joint data.

[0095] Specifically, firstly, a time-series analysis method is used to detect instantaneous jump anomalies in joint data: the rate of change of joint angular position (i.e., angular velocity) and the rate of change of torque are calculated using the difference between adjacent frames. Then, the calculated rates of change of angular velocity and torque are compared with a preset first threshold. If the calculated rates of change consistently exceed this first threshold, it indicates that the joint data exhibits an anomaly that does not conform to physical norms, usually caused by occasional sensor malfunctions or strong external impacts, and the data will be marked as unusable. For example, if the instantaneous change exceeds a physically reasonable threshold (e.g., angular velocity jump > 0.5 rad / s, torque jump > 10 N·m), it is marked as an anomaly in joint data jump.

[0096] Secondly, for noise anomaly detection in joint data, a frequency domain analysis method is employed: a Fast Fourier Transform (FFT) is performed on the joint position or torque sequence to transform it from the time domain to the frequency domain. In the frequency domain, the energy proportion of high-frequency components is analyzed. This energy proportion reflects the intensity of random noise mixed in the signal. Subsequently, the high-frequency energy proportion is compared with a preset second threshold. If it exceeds the threshold, it indicates that the joint data is continuously contaminated by high-frequency noise, which may originate from quality problems of the sensor itself or continuous electromagnetic interference, and the data will be identified as noise anomaly. For example, the energy of normal operation data is mainly concentrated in the low-frequency band (e.g., less than 10Hz), corresponding to the frequency of normal human or robot movements; if the energy proportion of high-frequency components (e.g., greater than 40Hz) exceeds the threshold (e.g., 30%), sensor noise may be present; if the high-frequency components of the joint count exceed the set threshold for 10 consecutive frames, it is marked as joint data noise anomaly.

[0097] In this embodiment, time-series analysis is employed to calculate the rate of change of velocity or torque between adjacent frames and compare it with a first threshold. This efficiently and accurately captures and eliminates abrupt changes in data caused by occasional sensor malfunctions or instantaneous physical impacts, ensuring the physical continuity and rationality of the motion trajectory. More innovatively, frequency domain analysis (FFT) is introduced. By quantifying the energy proportion of high-frequency components and comparing it with a second threshold, the automated identification and filtering of persistent and highly concealed sensor-inherent noise contamination is successfully achieved. This dual mechanism of time and frequency domains fundamentally ensures the high purity and reliability of joint data, preventing noise and anomalous data from misleading and degrading the training of the embodied intelligent model, and greatly improving the stability and accuracy of the final trained model.

[0098] In one embodiment, the specific method for performing the spatial consistency check includes:

[0099] Using the intrinsic parameter matrix in the camera parameters and the depth value in the visual data, the coordinates of the end pixel located in the visual data are back-projected into 3D points in the camera coordinate system;

[0100] The 3D points are converted to the robot link coordinate system using the extrinsic parameters in the camera parameters to obtain the visually calculated end position coordinates P_vis.

[0101] Extract the end-joint position recorded in the joint data;

[0102] Calculate the Euclidean distance of ||P_vis-P_joint|| and compare the Euclidean distance with a preset third threshold to determine the spatial deviation.

[0103] Specifically, firstly, using the intrinsic parameter matrix (K) in the camera parameters and the depth value (d) in the visual data, the pixel coordinates (u,v) of the robot's end effector located in the visual data are back-projected to calculate the 3D spatial coordinates (Xc,Yc,Zc) of that point in the camera coordinate system. Then, using the extrinsic parameters (including the rotation matrix R and translation vector T) in the camera parameters, the 3D point (Xc,Yc,Zc) in the camera coordinate system is transformed to the robot link coordinate system (or world coordinate system) to obtain the visually transformed end effector position coordinates P_vis. Simultaneously, the end effector position coordinates P_joint recorded by the robot's kinematic model are directly extracted from the joint data. P_vis and P_joint are thus mapped to the same 3D coordinate system. Finally, the Euclidean distance ||P_vis-P_joint|| is calculated and compared with a preset third threshold to determine if there is a significant spatial deviation. If the deviation exceeds the threshold, it indicates a serious inconsistency in calibration, kinematics, or sensor performance, and the data should be marked as unacceptable.

[0104] For example, the first step is to locate the pixel coordinates (u, v) of the end of both arms (such as the right claw) from the RGB image, combine them with the depth value d of the corresponding depth map, and back-project them into 3D points (Xc, Yc, Zc) in the camera coordinate system using the camera intrinsic parameter matrix K: Xc=(u−cx)×d / fx, Yc=(v−cy)×d / fy, Zc=d; the second step is to transform (Xc, Yc, Zc) into the link coordinate system using the camera extrinsic parameters (the transformation matrix from camera to link), to obtain the end position coordinates P_vis calculated from the video, for example (0.3, 0.2, 0.1); finally, extract the end position coordinates P_joint of "state / end / position" in the joint data, for example (0.32, 0.21, 0.1), and calculate the Euclidean distance between them ||P_vis-P_joint||. If this example exceeds a set threshold (such as 20mm), it is marked as a spatial consistency anomaly. In this example, the Euclidean distance between them is 22mm.

[0105] In this embodiment, by utilizing camera intrinsic and extrinsic parameters and depth information for precise geometric backprojection and coordinate transformation, high-precision alignment of the visual perception results (P_vis) and the robot kinematic record (P_joint) is achieved. The core advantage of this mechanism lies in its ability to quantitatively identify and eliminate spatially inconsistent data caused by factors such as sensor installation, inaccurate camera calibration, or errors in the robot's URDF model. This completely solves the deficiency of traditional methods that can only perform subjective comparisons at the 2D image level, ensuring the accurate correspondence of all multimodal information in the training data in 3D physical space, and greatly improving the accuracy and generalization ability of the model's perception-action closed loop.

[0106] Furthermore, in one embodiment, the standardized format data also includes audio data. The audio data is audio data collected during the data acquisition process, stored in .wav format, and typically includes human voice commands and ambient sounds.

[0107] Furthermore, in one embodiment, the audio data quality check specifically includes the following steps:

[0108] First, a signal-to-noise ratio (SNR) evaluation is performed: the acquired audio data undergoes signal processing analysis, and the SNR is determined by calculating the ratio of the overall power of the audio data to the background noise power. This SNR is then compared to a preset threshold (e.g., 15 dB). If the SNR is below the threshold, it indicates that the voice command is severely contaminated by noise, and the audio data is marked as substandard and should not be used for training models that rely on voice commands.

[0109] Secondly, a voice command validity check is performed: For audio data containing human voice commands, an ASR (Automatic Speech Recognition) model can be invoked to transcribe it into text. The transcribed text is then compared with the voice command text in the task annotation file to determine the content consistency or semantic similarity between the two. If the consistency or similarity score is lower than a preset threshold, it indicates that the audio command was recorded incorrectly or that the command does not match the actual task, thus marking the audio data as invalid.

[0110] Through the above processing, this embodiment ensures that the audio data input as embodied intelligent commands has sufficient clarity and accuracy, further enhancing the overall quality of multimodal data.

[0111] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0112] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0113] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.

[0114] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0115] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0116] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.

[0117] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0118] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0119] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0120] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0121] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0122] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0123] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0124] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0125] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for assessing and verifying the quality of embodied intelligence multi-source data, characterized in that, The method includes: The embodied intelligence multi-source data is converted into standardized format data; the standardized format data includes visual data, joint data, task annotation, camera parameters, and calibration information corresponding to the robot's URDF model; The standardized format data is subjected to a quality assessment, which includes at least one of integrity checks, usability checks, and consistency checks. After the quality assessment is completed, the robot URDF model corresponding to the standardized format data is loaded, the URDF model is driven to move according to the joint data, and the corresponding visual data is played synchronously to realize the linkage playback of model actions and visual images. The linked playback is analyzed using a first major model and / or a second major model; wherein, the first major model is used to determine the physical safety of the robot's actions, and the second major model is used to determine the high-level semantic correctness of the robot's actions and the task annotations.

2. The method according to claim 1, characterized in that, The integrity check specifically includes: The directory hierarchy, naming format, and existence of key files of the standardized format data are verified by recursive path traversal and / or regular expression matching. And / or perform content integrity verification on the joint data and the task annotations through field parsing and mapping verification.

3. The method according to claim 1, characterized in that, The availability check specifically includes: By analyzing color histograms, calculating inter-frame differences, or statistically analyzing depth value ranges, invalid frames, stutters, or depth anomalies in the visual data are identified. And / or by comparing the joint data with the calibration information, a motion limit check can be performed; And / or based on the joint data, detect instantaneous jumps or noise anomalies in the joint's velocity or torque through time-series analysis and / or frequency domain analysis; And / or check the invertibility of the intrinsic parameter matrix of the camera parameters, the normalization degree of the extrinsic quaternions, and the validity of the intrinsic and extrinsic parameters.

4. The method according to claim 1, characterized in that, The consistency check specifically includes: Temporal consistency check: Determine the alignment degree between the number of visual data frames and the number of joint data frames; Spatial consistency check: Determine the spatial deviation between the end position in the visual data and the end position in the joint data through geometric calculations; Semantic consistency check: The video content is semantically validated by calling the third major model to determine the degree of semantic matching between the video content and the task-annotated text.

5. The method according to claim 1, characterized in that, The steps of analyzing the linked playback using the first major model and / or the second major model specifically include: The screen recording process of the aforementioned linkage playback is performed; The screen recording data is input into the first large model, which is a motion recognition large model that identifies high-frequency shaking, joint self-collision, or uncontrolled arm swinging to determine physical safety. The screen recording data and the task tags in the original annotations are input into the second large model, which is a multimodal large model. The semantic similarity between the action description output by the multimodal large model and the task tags is determined.

6. The method according to claim 3, characterized in that, The step of detecting instantaneous jumps or noise anomalies in the velocity or torque of the joint based on the joint data through time-series analysis and / or frequency domain analysis specifically includes: The time-series analysis calculates the joint angular velocity and torque change rate using adjacent frames, and compares the change rate with a preset first threshold to determine the instantaneous jump anomaly in the joint data; The frequency domain analysis involves performing a fast Fourier transform on the joint position or torque sequence, analyzing the energy proportion of high-frequency components, and comparing the energy proportion with a preset second threshold to determine noise anomalies in the joint data.

7. The method according to claim 4, characterized in that, The specific methods for spatial consistency checking include: Using the intrinsic parameter matrix in the camera parameters and the depth value in the visual data, the coordinates of the end pixel located in the visual data are back-projected into 3D points in the camera coordinate system; The 3D points are converted to the robot link coordinate system using the extrinsic parameters in the camera parameters to obtain the visually calculated end position coordinates P_vis. Extract the end-joint position recorded in the joint data; Calculate the Euclidean distance of ||P_vis-P_joint|| and compare the Euclidean distance with a preset third threshold to determine the spatial deviation.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent data processing method and equipment for body

    CN120147918A

  • Reliability analysis method and system of body model, medium, equipment and product

    CN120449922A