Abnormality detection and usability verification method, device, medium and product of body data

CN122241155BActive Publication Date: 2026-09-15SHANGHAI COOPERS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610569046.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-09-15
Estimated Expiration
2046-04-28

AI Technical Summary

Technical Problem

[0004]现有技术中,对具身数据的质量验证通常仍停留在文件存在性检查、简单格式校验或人工抽检层面,缺乏对多模态具身数据进行统一、系统、自动化质量检测与可用性验证的技术手段

Benefits of technology

[0014]Compared with related technologies, the solution provided in this application uses timestamps as a unified time reference to perform in-depth time alignment analysis on multimodal perception data such as red-green-blue video and depth video, as well as robot state and motion data, and accurately extracts the response correlation between state data and changes in motion data. This allows for the keen detection of temporal hidden anomalies, such as modal misalignment and delayed command response, from massive amounts of raw collected samples—anomalies that are extremely difficult to detect using conventional inspection methods. Furthermore, by analyzing the matching degree of change trends between state and motion data, combined with the detection of the existence and continuity of multimodal data, deeper defects such as abnormal state jumps, high-frequency actuator jitter, and sampling interruptions are further identified. Based on this, the anomaly detection processing results, covering multiple dimensions including temporal consistency, rationality of state and action, and data integrity, are weighted and fused for calculation, enabling fully automated output of quantitative comprehensive evaluation scores and accurate usability judgments. This not only significantly reduces the time and manpower costs required for manual screening of abnormal data, but also ensures the reliability of the data entering the database from the source. It enables the automated cleaning and transformation from raw collection logs containing a large number of defects to high-quality, high-security embodied intelligent model (such as VLA model) training datasets, providing data assurance for the large-scale generalization of embodied intelligent entity robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122241155B_ABST
    Figure CN122241155B_ABST
Patent Text Reader

Abstract

The application discloses a body data anomaly detection and availability verification method, device, medium and product, comprising: acquiring body data to be verified, the body data at least including multi-modal perception data, state data and action data associated with a timestamp; taking the timestamp as a unified time reference, time aligning the multi-modal perception data, the state data and the action data, and extracting the time sequence correspondence relationship between the data; based on the time sequence correspondence relationship and the association relationship between the action data change and the response change of the state data within a preset time range, performing anomaly detection processing on the body data, the anomaly detection processing including: video quality detection, data integrity detection, time sequence synchronization detection and state-action rationality detection; fusing the results of the anomaly detection processing to generate a comprehensive evaluation result, and determining the availability determination result of the body data based on the comprehensive evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular to a method, device, medium and product for anomaly detection and usability verification of embodied data. Background Technology

[0002] Embodied AI aims to enable robots or intelligent agents to perform tasks in real or simulated environments through perception, understanding, and action. As embodied AI models develop, the training, evaluation, and deployment processes rely on large-scale embodied data containing visual, depth, and temporal information about states and actions. This data typically includes RGB video, depth video, and state, action, and timestamp data stored in related files. These data are used to characterize scene changes, spatial structure changes, robot state changes, and control output processes during task execution.

[0003] In actual data acquisition, embodied data often originates from real robot operations, remote teaching, automated acquisition systems, or simulated migration processes. Due to the diversity of acquisition equipment types, inconsistent sampling frequencies, differences in sensor performance, and the complexity of task execution, the acquired data is prone to various anomalies. For example, at the video level, there may be issues such as black and white frames or freezes; at the data integrity level, there may be issues with missing modalities or key fields; at the temporal level, there may be inconsistencies in correspondences or modal misalignments; and at the state and action level, there may be abnormal jumps or motion jitter.

[0004] In current technologies, the quality verification of embodied data typically remains at the level of file existence checks, simple format verification, or manual sampling, lacking unified, systematic, and automated technical means for quality inspection and usability verification of multimodal embodied data. On the one hand, simply checking the existence of files is insufficient to identify video quality anomalies and abnormal state / action behavior; on the other hand, even if all modal files exist, issues such as time asynchrony between modalities may still exist, resulting in data that cannot accurately and effectively reflect the robot's task execution process. Directly using such anomalous data for embodied model training, verification, or evaluation can easily affect model learning effectiveness, evaluation accuracy, and system stability. Summary of the Invention

[0005] One objective of this application is to provide a method, device, medium, and product for anomaly detection and usability verification of embodied data, at least to address the problem that existing technologies lack unified, systematic, and automated quality inspection methods for multimodal embodied data, making it difficult to comprehensively identify implicit anomalies across multiple dimensions such as video, time series, and state actions. This easily leads to defective embodied data directly entering the model training stage, severely impacting the learning effect and system stability of embodied intelligent models. To achieve the above objective, some embodiments of this application provide the following aspects:

[0006] This application provides a method for anomaly detection and usability verification of embodied data, the method comprising:

[0007] Obtain the embodied data to be verified, which includes multimodal perception data as well as state data and action data associated with timestamps;

[0008] Using the timestamp as a unified time reference, the multimodal perception data, the state data, and the action data are time-aligned, and the temporal correspondence between the data is extracted.

[0009] Based on the temporal correspondence and the correlation between the changes in the action data and the response changes in the state data within a preset time range, anomaly detection processing is performed on the embodied data. The anomaly detection processing includes at least one of the following: video quality detection based on image pixel features, data integrity detection based on data existence and continuity, temporal synchronization detection based on time deviation, and state-action rationality detection based on the relationship between state changes and action changes.

[0010] The results of the anomaly detection processing are fused to generate a comprehensive evaluation result, and based on the comprehensive evaluation result and preset judgment rules, the usability judgment result of the embodied data is determined.

[0011] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.

[0012] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method described above.

[0013] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.

[0014] Compared with related technologies, the solution provided in this application uses timestamps as a unified time reference to perform in-depth time alignment analysis on multimodal perception data such as red-green-blue video and depth video, as well as robot state and motion data, and accurately extracts the response correlation between state data and changes in motion data. This allows for the keen detection of temporal hidden anomalies, such as modal misalignment and delayed command response, from massive amounts of raw collected samples—anomalies that are extremely difficult to detect using conventional inspection methods. Furthermore, by analyzing the matching degree of change trends between state and motion data, combined with the detection of the existence and continuity of multimodal data, deeper defects such as abnormal state jumps, high-frequency actuator jitter, and sampling interruptions are further identified. Based on this, the anomaly detection processing results, covering multiple dimensions including temporal consistency, rationality of state and action, and data integrity, are weighted and fused for calculation, enabling fully automated output of quantitative comprehensive evaluation scores and accurate usability judgments. This not only significantly reduces the time and manpower costs required for manual screening of abnormal data, but also ensures the reliability of the data entering the database from the source. It enables the automated cleaning and transformation from raw collection logs containing a large number of defects to high-quality, high-security embodied intelligent model (such as VLA model) training datasets, providing data assurance for the large-scale generalization of embodied intelligent entity robots. Attached Figure Description

[0015] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0016] Figure 1 A flowchart of an anomaly detection and usability verification method for embodied data provided as an exemplary embodiment of this disclosure;

[0017] Figure 2 An exemplary structural diagram of the electronic device provided for some embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] Figure 1 An exemplary embodiment of this disclosure provides a method for anomaly detection and usability verification of embodied data, the method comprising:

[0020] S101. Obtain the embodied data to be verified, the embodied data including multimodal perception data and state data and action data associated with timestamps.

[0021] Specifically, the system reads the input original embodied record file, extracts the red, green and blue video frame sequences and the depth video frame sequences as the multimodal perception data, synchronously reads and records the state data of the robot body posture or joint features, and records the action data of the underlying control commands, and at the same time obtains the timestamp that uniformly records the time when the above information occurs.

[0022] S102. Using the timestamp as a unified time reference, perform time alignment on the multimodal perception data, the state data, and the action data, and extract the temporal correspondence between each data.

[0023] Specifically, during time alignment processing, the system extracts a preset reference time series as a unified time axis, which consists of multiple consecutive timestamps. For any target timestamp in the reference time series, the system determines the corresponding matching data item based on the principle of minimum time distance in the red-green-blue video frame sequence and depth video frame sequence, the state data sequence, and the motion data sequence of the multimodal perception data, thereby establishing the alignment relationship of each modal data on the unified time axis.

[0024] Based on this, the system calculates the time deviation of each data item relative to the target timestamp and retains the time deviation as the basis data for subsequent time-series synchronization analysis.

[0025] S103. Based on the temporal correspondence and the correlation between the changes in the action data and the response changes in the state data within a preset time range, perform anomaly detection processing on the embodied data. The anomaly detection processing includes at least one of the following: video quality detection based on image pixel features, data integrity detection based on data existence and continuity, temporal synchronization detection based on time deviation, and state-action rationality detection based on the relationship between state changes and action changes.

[0026] Specifically, for each detection dimension of the anomaly detection process, the system can adopt a variety of different implementation methods according to the actual business scenario and computing power configuration.

[0027] For the video quality detection, the system performs image anomaly detection on the image sequence in the multimodal perception data based on image pixel features to identify abnormal situations such as exposure anomalies, blurred frames, image freezes, and loss of depth information.

[0028] For the data integrity detection, the system detects the existence and continuity of the multimodal perception data, the state data, the action data, and the timestamps to identify modality missing, field missing, or sampling interruption anomalies.

[0029] For the timing synchronization detection, the system calculates the alignment error between each modal data based on the time deviation, and determines whether there is a modal misalignment or response lag anomaly based on the time difference between the change in the action data and the change in the response of the state data.

[0030] For the rationality detection of the state and action, the system performs differential analysis based on the relationship between the changes in the state data and the action data, and identifies abnormal situations such as abnormal state jumps, action jitters, or long periods of stagnation based on the magnitude, direction, or frequency of change.

[0031] S104. The results of the anomaly detection processing are fused to generate a comprehensive evaluation result, and the usability determination result of the embodied data is determined based on the comprehensive evaluation result and the preset judgment rules.

[0032] Specifically, the system generates corresponding individual evaluation scores for each anomaly ratio and judgment result generated in the video quality detection, data integrity detection, temporal synchronization detection, and state action rationality detection. The system then performs a weighted fusion calculation on these individual evaluation scores to obtain a comprehensive score representing the overall evaluation result. When the comprehensive score is greater than or equal to a first threshold, the usability determination result of the embodied data is determined to be usable; when the comprehensive score is less than a second threshold, the usability determination result is determined to be unusable.

[0033] In this embodiment, by using timestamps as a unified benchmark to extract and analyze the temporal correspondence and change matching relationships of multimodal perception data, state data, and action data, the comprehensiveness and accuracy of multi-dimensional anomaly identification are significantly improved. This effectively detects latent anomalies such as modal misalignment and response lag. Subsequently, by fusing and calculating the results of various anomaly detection processes, a comprehensive data quality score can be automatically completed, generating a scientifically reliable usability assessment result. Thus, a highly efficient, highly automated, and multimodal unified embodied data quality screening and verification method is established, achieving automated conversion and cleaning from mixed and defective raw data collection to high-quality embodied intelligent model training data.

[0034] In one embodiment, the timing synchronization detection in the anomaly detection process includes:

[0035] The overall alignment error is calculated based on the alignment deviations of the multimodal sensing data, the state data, and the action data relative to the timestamp.

[0036] The state-action synchronization error is calculated based on the correspondence between the change in state data at adjacent time points and the action data at the previous time point.

[0037] When the overall alignment error or the state action synchronization error is greater than a preset error threshold, it is determined that there is a synchronization anomaly.

[0038] Specifically, the system extracts a unified timestamp sequence as a timeline reference, and reads the recording time of each frame of visual image in the multimodal perception data, the sampling time of each ontological feature in the state data, and the issuance time of each control command in the action data. Subsequently, the system calculates the difference sequence between the actual recording time of each of the above data items and the corresponding timestamp to obtain the alignment deviation for each modality. The system performs statistical analysis on the extracted alignment deviation, such as calculating the mean squared error or the maximum deviation value, thereby quantifying the overall alignment error that characterizes the consistency of the overall timeline.

[0039] Next, the system calculates the state-action synchronization error based on the correspondence between the changes in state data at adjacent time points and the action data at the previous time point. In the specific calculation process, the system extracts the state data from the current time point and the previous time point, performs a difference operation to obtain the actual change in the entity's state over time. Simultaneously, the system extracts the action data issued at the previous time point and parses the expected state response action corresponding to the action data. The system compares the degree of difference between the actual change and the expected state response action to quantify the state-action synchronization error, which reflects the physical execution coherence. Finally, the system compares the calculated value with a set standard. If either the overall alignment error or the state-action synchronization error exceeds a preset error threshold, a synchronization anomaly is determined. The preset error threshold specifically includes an alignment tolerance threshold set for the overall alignment error and a response tolerance threshold set for the state-action synchronization error. When any of the actually calculated error values ​​exceeds the corresponding threshold limit, the system outputs an anomaly judgment conclusion indicating a timing inconsistency, marks the specific segment location where the anomaly occurred, and summarizes the judgment results in the subsequent fusion stage.

[0040] Furthermore, in one embodiment, the system uses timestamps as a unified time reference to perform time alignment on red-green-blue video frames, depth video frames, state data, and motion data. This time alignment can be specifically implemented using a nearest neighbor matching algorithm.

[0041] After the alignment operation is completed, the system checks whether the red-green-blue video frame (RGB video), the depth video frame, the state data, and the action data can all maintain a consistent time correspondence with the timestamp, and further checks whether the state data responds within a reasonable time range after the action data changes.

[0042] For the overall alignment between each independent modality and the timestamp, the system calculates the overall alignment error based on the time deviation after alignment. Let the first... Each timestamp is , and the first The actual sampling times of the red-green-blue video frames, depth video frames, state data, and motion data matched with each timestamp are respectively... , , and Then each mode in the th The time relative to the timestamp The alignment deviations can be expressed as follows:

[0043]

[0044]

[0045]

[0046]

[0047] in, , , and They represent the first The red-green-blue video frames, depth video frames, state data, and motion data at each moment are relative to the first... Alignment deviation of timestamps.

[0048] Furthermore, to eliminate the influence of the deviation direction on the overall error statistics, the system takes the absolute value of the above alignment deviation, and the overall alignment error can be expressed as:

[0049]

[0050] in, This indicates the total number of timestamps obtained within the current detection window or evaluation period. This represents the sequence number or index variable in the timestamp sequence.

[0051] The larger the overall alignment error, the more inconsistent the time correspondence between each modality and the timestamp.

[0052] Regarding the synchronization relationship between the state data and the action data, the system measures the timing response error by comparing the changes in adjacent state data with the correspondence of the action data at the previous moment. This error can be expressed as:

[0053]

[0054] in, Indicates the first The state data at each time point. This represents the action data from the previous moment. This is the proportionality coefficient.

[0055] when or If the value exceeds a preset threshold, a timing synchronization anomaly is determined.

[0056] After completing the above detection, the system obtains the overall alignment error result between each modality and the timestamp, as well as the synchronization error result between the state data and the action data.

[0057] Then, the system calculates the timing synchronization score based on the above abnormal results.

[0058] Furthermore, in one embodiment, the timing synchronization score can be expressed as:

[0059]

[0060] The larger the error, the lower the corresponding synchronization score.

[0061] In some embodiments, when or When the score exceeds the preset threshold, the data can be directly marked as having a timing synchronization anomaly; when the synchronization score is lower than the preset threshold, the data can be determined as unusable or pending review.

[0062] Finally, output the timing synchronization score. And score the timing synchronization. This information is provided for subsequent comprehensive scoring steps to determine usability.

[0063] In this embodiment, by calculating the overall alignment error based on the timestamp digital benchmark and the state-action synchronization error based on the physical motion logic, the system can not only accurately identify timeline misalignment anomalies at the data level generated during the recording of the underlying hardware acquisition device, but also deeply discover deep-seated defects such as lag or execution gaps between instruction issuance and physical response during robot entity operation. This dual temporal verification mechanism, which combines data alignment and physical response patterns, overcomes the drawback of conventional verification methods that rely solely on timestamp numerical comparisons while ignoring entity execution characteristics. It greatly improves the detection rate of implicit synchronization anomalies, ensuring the rigor and authenticity of the selected embodied data in the temporal causal logic of "perception-decision-execution," and providing temporal dimension guarantees for high-quality training of embodied intelligent models.

[0064] In one embodiment, the state action rationality detection in the anomaly detection process includes:

[0065] If the change in the state data between two adjacent time points is greater than the state transition threshold, it is determined that there is an abnormal state transition.

[0066] When the motion data at adjacent time points exhibits repeated directional changes and the magnitude of these changes exceeds the motion jitter threshold, motion jitter is determined to exist.

[0067] If the change in the action data at multiple consecutive time points is less than the action stillness threshold and the change in the state data is less than the state stillness threshold, it is determined that there is a stagnant segment.

[0068] Specifically, the system extracts the state data from two adjacent time points and performs differential operations to calculate the amplitude of state change. Then, the system compares the amplitude of state change with a preset state transition threshold. If the amplitude of the state data from two adjacent time points exceeds the state transition threshold, an abnormal state transition is determined. Further, the system statistically analyzes all times when abnormal state transitions occur to calculate the abnormal state transition rate. Regarding the rationality of the action data, the system evaluates the continuity of action based on the same differential analysis approach. The system extracts the action data from adjacent time points and calculates the continuity and change in the direction of movement. If the direction of the action data from adjacent time points is repeated and the amplitude of change exceeds the action jitter threshold, action jitter is determined. The system traverses the action sequence, calculates the proportion of segments that meet the action jitter conditions, and thus obtains the action jitter ratio. For stagnation during execution, the system identifies it based on continuous low-change intervals. The system extracts the action data and state data frame by frame. When the change amplitude of the action data is less than the action stillness threshold for multiple consecutive time intervals, and the change amplitude of the state data is simultaneously less than the state stillness threshold, a stagnant segment is determined to exist. Further, the system summarizes and statistically analyzes the duration of all stagnant segments to determine the proportion of stagnant segments. After completing the above detection, the system calculates a state-action rationality score based on the state abnormal transition rate, the action jitter ratio, and the proportion of stagnant segments, combined with preset weight parameters, and outputs the state-action rationality score to subsequent stages.

[0069] Furthermore, in one embodiment, if the change in the state data between two adjacent time points significantly exceeds a preset threshold, it is determined to be an abnormal state jump; if the action data frequently changes direction repeatedly and with a large change in amplitude within a short period of time, it is determined to be action jitter; if the action data is close to zero for a long time and the change in the state data is also minimal, it is determined to be a stagnant segment.

[0070] The state transition rate can be expressed as:

[0071]

[0072] in, This is the state transition threshold. Indicates the first State data at time 1 With the State data at time 1 The difference between them is the change in state data between two adjacent time points.

[0073] Furthermore, in one embodiment, the system can also statistically analyze the action data (action) sequence to determine the proportion of action jitter.

[0074] When the action data (action) shows directional repetition between adjacent moments, and the amplitude of the change exceeds a preset threshold, it is determined that there is motion jitter at the corresponding location. The motion jitter ratio can be expressed as:

[0075]

[0076] in, This is the threshold for motion jitter. Indicates the first Action data at each moment With the State data at time 1 Similarly, the difference between them, Indicates the first Action data at each moment With the State data at time 1 The difference between them.

[0077] Furthermore, in one embodiment, for stagnant segments, the system identifies them based on continuous low-change intervals. When multiple consecutive moments simultaneously satisfy the condition that the action data is close to zero and the state data changes extremely little, the continuous interval is determined to be a stagnant segment. Specifically, when the following formula is true:

[0078]

[0079] Then determine the first At some point, the situation is at a standstill; among them... Indicates the threshold for motion stillness. This represents the state stagnation threshold. Furthermore, the system counts the duration of each stagnant segment and obtains the percentage of stagnant segments:

[0080]

[0081] in, This represents the total number of moments that were determined to be in a stationary state. This represents the total number of moments.

[0082] After completing the above tests, the system obtained the abnormal state transition rate, motion jitter ratio, and percentage of stagnant segments, respectively.

[0083] Then, the system calculates the rationality score of the state action based on the above abnormal results.

[0084] In one embodiment, the rationality score of the state action can be expressed as:

[0085]

[0086] in, , and Let these represent the weights corresponding to abnormal state transitions, motion jitter, and static segments, respectively, and satisfy the following conditions:

[0087]

[0088] The higher the proportion of the above-mentioned anomalies, the lower the corresponding state and action rationality score.

[0089] In some embodiments, when , or If any indicator exceeds the preset threshold, the data can be directly marked as having an abnormal state / action rationality; if the state / action rationality score is lower than the preset threshold, the data can be determined as unusable or pending review.

[0090] Finally, this step outputs the abnormal state transition rate, motion jitter result, motion jitter ratio, stagnation segment result, stagnation segment ratio, and state motion rationality score, and provides the state motion rationality score to the subsequent comprehensive scoring step for usability determination.

[0091] In this embodiment, by performing temporal difference analysis on state and motion data and setting multi-dimensional thresholds for jumps, jitters, and stillness, the system can not only accurately identify numerical abrupt changes caused by underlying sensor failures or communication packet loss, but also deeply uncover hidden abnormal execution phenomena such as motion jitters and prolonged pauses during task execution. This verification mechanism based on kinematic physical correlation overcomes the technical limitations of conventional data integrity checks in assessing the rationality of underlying execution logic. From the dimensions of control commands and the evolution of the entity's posture, it ensures the smoothness and rationality of the collected embodied data in the physical world, thereby effectively avoiding the misleading and interference of erroneous or rigid teaching actions on the subsequent training of the embodied intelligent decision-making model.

[0092] In one embodiment, the data integrity detection in the anomaly detection process includes:

[0093] Detect whether the multimodal perception data, the state data, the action data, and the timestamp are missing;

[0094] The sampling interruption rate is calculated based on the time interval between adjacent timestamps;

[0095] When the sampling interruption rate exceeds the time interruption threshold, a sampling interruption anomaly is determined to exist.

[0096] Specifically, the system traverses the aligned embodied data sequence, checking each frame to ensure that the red-green-blue video frame sequence, the depth video frame sequence, the state data reflecting the embodied posture, and the action data representing control commands completely contain the corresponding data files or field information. If, at a specific recording node, the data file for any modality is not generated, the key field content is empty, or the timestamp itself is missing, the system determines that there is data loss at the corresponding node and marks the modality type and specific time location where the loss occurred.

[0097] Next, the system calculates the sampling interruption rate based on the time interval between adjacent timestamps. In the specific calculation process, the system extracts the continuously recorded timestamp sequence and sequentially calculates the actual physical time span between two adjacent timestamps. The system compares the actual physical time span with a preset theoretical standard sampling period, filters out abnormal time spans that significantly exceed the normal fluctuation range, and calculates the proportion of the cumulative duration or frequency of these abnormal time spans in the overall sequence, thereby quantifying the sampling interruption rate.

[0098] Finally, the system performs threshold interception, determining a sampling interruption anomaly when the sampling interruption rate exceeds the time interruption threshold. The system rigorously compares the calculated sampling interruption rate with the pre-set time interruption threshold. Once the calculated value exceeds the time interruption threshold, the system determines that a serious network disconnection or sensor failure has occurred during the acquisition or communication transmission of the data. It then generates a data integrity detection result including the aforementioned missing identifier and the sampling interruption anomaly, and outputs the summarized judgment data to the subsequent fusion stage.

[0099] Furthermore, in one embodiment, the system can detect whether a sampling interruption exists based on the time interval between adjacent timestamps. The sampling interruption rate can be expressed as:

[0100]

[0101] in, Indicates the first A timestamp, Indicates the time interruption threshold. This is an indicator function.

[0102] When the sampling interruption rate exceeds the threshold, a sampling interruption anomaly is determined; when any modality is missing, a modality missing anomaly is determined; when any key field is missing, a field missing anomaly is determined.

[0103] After completing the above tests, the system obtained the modality missing results, field missing results, and sampling interruption rate, respectively.

[0104] Then, the system calculates the integrity score based on the above abnormal results.

[0105] In one embodiment, the integrity score can be expressed as:

[0106]

[0107] in, This indicates a missing modality; it is indicated when any necessary modality is missing. =1, otherwise =0; This indicates a missing field flag; it is triggered when any key field is missing. =1, otherwise =0; Indicates the sampling interruption rate; This indicates the weight corresponding to the sampling interruption.

[0108] From the above formula, it can be seen that when =1 or When =1, the integrity score A value of 0 indicates that the data is unqualified due to missing modalities or missing key fields; only when the modality and fields are complete will the system deduct the integrity score based on the sampling interruption rate.

[0109] The higher the degree of anomaly, the lower the corresponding integrity score. In some embodiments, when the integrity score is lower than a preset threshold, the data can be directly determined as unusable.

[0110] Finally, this step outputs a data integrity score. and the integrity score This information is provided for subsequent comprehensive scoring steps to determine usability.

[0111] In this embodiment, by performing fine-grained fact-finding checks on multimodal perception data, state data, action data, and timestamps, and by combining the actual intervals between adjacent timestamps to deeply calculate the sampling interruption rate, the system can not only accurately intercept surface-level explicit errors such as missing files or blank fields, but also accurately capture temporal discontinuities caused by device lag. This overcomes the technical blind spot of conventional file size inspection, which cannot assess the continuity of the timeline, and improves the convergence efficiency and generalization stability of subsequent model training.

[0112] In one embodiment, the multimodal sensing data includes a red-green-blue video frame sequence and a depth video frame sequence; the video quality detection in the anomaly detection process includes:

[0113] Based on the pixel brightness information and image sharpness information of the red, green and blue video frame sequence, black frames, white frames and blurry frames are identified;

[0114] Based on the pixel differences between adjacent red, green and blue video frames and the changes in the corresponding depth video frames, frozen segments of the image are identified.

[0115] Based on the distribution of invalid depth values ​​in the depth video frame sequence, identify depth failure frames;

[0116] The abnormal proportions of the black frames, white frames, blurred frames, frozen frames, and depth failure frames are statistically analyzed, and the video quality detection results are generated based on the abnormal proportions.

[0117] Specifically, firstly, the system identifies black frames, white frames, and blurred frames based on the pixel brightness information and image sharpness information of the red-green-blue video frame sequence. In practice, the system traverses the red-green-blue video frame sequence frame by frame, calculating the global or local pixel mean value of each red-green-blue video frame to quantify the pixel brightness information. The system compares the calculated pixel mean value with a preset brightness threshold. If the pixel mean value is less than the set black frame threshold, the corresponding red-green-blue video frame is determined to be a black frame; if the pixel mean value is greater than the set white frame threshold, the corresponding red-green-blue video frame is determined to be a white frame. Simultaneously, the system uses edge detection algorithms such as the Laplacian operator to calculate the Laplacian response variance of each red-green-blue video frame, using this as a core indicator to measure the image sharpness information. When the Laplacian response variance is less than the set blur determination threshold, the corresponding red-green-blue video frame is determined to be out of focus or motion blurred, and it is marked as a blurred frame.

[0118] Next, the system identifies frozen segments based on the pixel differences between adjacent red-green-blue video frames and the changes in the corresponding depth video frames. In actual execution, the system calculates the average pixel difference between adjacent red-green-blue video frames on the time axis to capture dynamic changes in the image content; simultaneously, the system extracts the depth video frames that match the red-green-blue video frames and calculates the magnitude of change in the spatial depth dimension. When the average pixel difference of multiple consecutive red-green-blue video frames is consistently lower than a preset pixel difference threshold, and the magnitude of change in the corresponding depth video frames is also consistently lower than a preset depth change threshold, the system excludes interference from static environments and comprehensively determines that the multimodal perception data shows a frozen segment caused by hardware lag during this time period. Furthermore, the system identifies depth failure frames based on the distribution of invalid depth values ​​in the depth video frame sequence. The system scans the depth video frame sequence frame by frame, counts the percentage of invalid depth pixels in each depth video frame caused by sensor blind spots, material reflections, or exceeding the physical limits of ranging, and determines the corresponding depth video frame as a depth failure frame when the percentage of invalid depth pixels exceeds a preset depth failure threshold.

[0119] Finally, the system calculates the anomaly ratios of the black frames, white frames, blurred frames, frozen frames, and depth failure frames, and generates the video quality detection results based on these ratios. The system summarizes the total number of anomaly images and the cumulative duration of anomaly segments marked in each of the above detection steps, and calculates the percentage of each anomaly in the overall red-green-blue video frame sequence or the depth video frame sequence. The system can further combine preset weighting rules to perform weighted deduction calculations on the percentages of the above-mentioned anomalies, thereby outputting a quantitative score representing the video quality detection result and a specific anomaly distribution list, and transmitting the detection results to subsequent processing nodes.

[0120] Furthermore, in one embodiment, for brightness anomaly detection in RGB video, the system performs analysis on each frame of the RGB image. Calculate the pixel mean:

[0121]

[0122] in, and These represent the image height and width, respectively. Indicates the first Pixels in a frame image The pixel value.

[0123] when At that time, the judgment of the first The frame is a black frame; when At that time, the judgment of the first The frame is a white frame. Among them, The threshold for black frames. The threshold is the white frame threshold. This threshold can be preset based on the type of acquisition device, historical qualified sample statistics, or an empirical range.

[0124] For sharpness detection of RGB video, the system uses Laplacian response variance as the sharpness index to perform blur detection on each frame of the image:

[0125]

[0126] in, Indicates the first The Laplacian transform result of the frame image, This indicates variance calculation. When... At that time, the judgment of the first The frame is a blurred frame, where, This is the threshold for fuzzy judgment.

[0127] For the validity detection of depth video, the system performs a depth image analysis on each frame. The proportion of invalid depth values ​​is calculated. Invalid depth values ​​include zero, negative values, non-numerical values ​​(NaN), infinite values, or values ​​exceeding the sensor's measurement range. Then, the... Invalid proportions in the frame depth map can be represented as:

[0128]

[0129] in, Represents a set of invalid depth values. Indicates an indicator function. When > When this happens, the frame is determined to be a depth-failed frame. This is the threshold for deep failure.

[0130] For image freeze detection, the system calculates the average pixel difference between adjacent RGB frames to characterize the degree of change in image content over time:

[0131]

[0132] When continuous The frame satisfies the condition that the difference between adjacent RGB frames is less than a preset threshold. When the corresponding depth frame changes are also consistently small, the continuous interval is determined to have a frozen image. Using a joint determination method combining RGB video and depth video helps reduce misjudgments caused by static scenes, minimal local movement, or insignificant lighting changes.

[0133] After completing the above tests, the system statistically analyzes various abnormal results, obtaining the proportion of black frames, white frames, blurred frames, deep failure frames, and the percentage of frozen segment duration.

[0134] Among them, the black frame ratio This represents the proportion of RGB frames judged as black out of the total number of RGB video frames; the proportion of white frames. This represents the proportion of RGB frames judged as white frames out of the total number of RGB video frames; the proportion of blurred frames. This represents the proportion of RGB frames judged as blurry out of the total number of RGB video frames; the proportion of depth-failure frames. This indicates the percentage of depth frames deemed as depth failure frames out of the total number of depth video frames; and the percentage of frozen segment duration. This indicates the proportion of the number of frames corresponding to the frozen segment out of the total number of frames in the RGB video.

[0135] Then, the system calculates the video quality score based on the above-mentioned anomaly ratio. In one embodiment, the video quality score can be expressed as:

[0136]

[0137] Where α1, α2, α3, α4, and α5 represent the weights corresponding to black frames, white frames, blurred frames, depth failures, and frozen segments, respectively, and satisfy the following:

[0138]

[0139] The higher the anomaly rate, the lower the corresponding video quality score. Finally, this step outputs the video quality score. and the This information is provided for subsequent comprehensive scoring steps to determine usability.

[0140] In this embodiment, by deeply mining the brightness distribution, sharpness variance, and depth validity features at the bottom layer of image pixels, and combining the temporal differences between consecutive frames with cross-modal information, the system can not only eliminate black-and-white frames and blurry frames caused by sudden changes in ambient lighting or camera occlusion, but also use dual verification of red, green, and blue visual features and depth features to detect image freezes and crashes caused by single sensor failures or network congestion. This ensures that the multimodal visual data ultimately input into the downstream embodied intelligence training pipeline has extremely high physical realism, sharpness, and dynamic coherence, effectively avoiding the weakening and interference of poor-quality images on the feature extraction ability and generalization performance of the visual language action model.

[0141] In one embodiment, fusing the results of the anomaly detection processing to generate a comprehensive evaluation result includes:

[0142] Scores are calculated for the video quality detection, data integrity detection, temporal synchronization detection, and state action rationality detection, respectively.

[0143] The scores of each item are weighted and combined to obtain a comprehensive score that represents the overall evaluation result.

[0144] Specifically, based on the anomaly ratio and evaluation indicators output from each of the aforementioned detection stages, the system generates video quality scores after normalization. Completeness score Timing synchronization score Reasonableness score for state and action Next, the system performs a weighted fusion calculation on the scores of each item to obtain a comprehensive score representing the overall evaluation result. Specifically, this can be calculated using the following formula:

[0145]

[0146] in, Each of these corresponds to a weighting coefficient. The sum of all the weighting coefficients mentioned above is limited to one, and the specific values ​​are pre-configured by the system based on the sensitivity of the current embodied intelligent operation task to different physical dimensions.

[0147] Furthermore, in one embodiment, the weighting coefficients for the video quality score, the integrity score, the temporal synchronization score, and the state / action rationality score in the comprehensive scoring formula are as follows: These are dynamically adjustable parameters. The system is configured to adaptively adjust the distribution of the weight coefficients based on the specific embodied intelligent operation task type and target application scenario corresponding to the embodied data to be verified.

[0148] Specifically, because different embodied intelligence tasks have significantly different tolerance rates for various data dimensions, the system can achieve customized quality screening by dynamically adjusting the weight coefficients of each item. For example, in tasks such as precision assembly of tiny objects, threading a needle, or dexterous hand manipulation, the downstream embodied intelligence model has extremely high requirements for the edge sharpness of visual feedback and the microsecond-level timeliness of the response to underlying control commands. For such high-precision tasks, the system will significantly increase the weight coefficients representing video quality scores and temporal synchronization scores, thereby imposing more severe score reduction penalties for slight image defocusing or extremely small temporal lags. Furthermore, in tasks such as long-distance, large-scale autonomous navigation, warehouse logistics handling, or all-terrain patrol, the downstream embodied intelligence model has a high tolerance for occasional blurring of single frames, but is extremely sensitive to the long-term coherence of the overall spatial trajectory and the macroscopic logic of continuous transitions in physical states. For such long-cycle tasks, the system will increase the weight coefficients representing the integrity score and the rationality score of the state action to prevent sampling gaps or coordinate space jumps that do not conform to physical laws in the selected data.

[0149] In one embodiment, determining the usability determination result of the embodied data based on the comprehensive evaluation result and preset determination rules includes:

[0150] When the overall score is greater than or equal to the first threshold, the embodied data is determined to be the first type of data;

[0151] When the overall score is less than the second threshold, the embodied data is determined to be the second type of data;

[0152] When the overall score is between the first threshold and the second threshold, the embodied data is determined to be third type of data; wherein the first threshold is greater than the second threshold.

[0153] Specifically, the system determines the classification result of the embodied data based on the calculated comprehensive score and preset judgment rules. The system compares the comprehensive score with multiple preset hierarchical judgment thresholds: when the comprehensive score is greater than or equal to the first threshold, the embodied data is marked as first-category data and included in the target data set; when the comprehensive score is less than the second threshold, the embodied data is marked as second-category data and isolated or removed; when the comprehensive score is between the first and second thresholds, the embodied data is marked as third-category data and corresponding anomaly details are output for subsequent processing. The first threshold is numerically greater than the second threshold.

[0154] In some embodiments, to prevent the overall score from exceeding a predetermined range, the system may also truncate the scores obtained in each step and the overall score; when the score is less than 0, it is recorded as 0, and when the score is greater than 100, it is recorded as 100.

[0155] Finally, the system outputs a comprehensive score and judgment result for the data to be verified. If necessary, the system can also output the main anomaly type, corresponding anomaly index, and location of the anomaly segment that led to the judgment result, to facilitate subsequent manual review, data cleaning, and quality tracking.

[0156] In the above embodiments, by introducing a hierarchical judgment rule based on dual threshold intervals, the system divides the data into three categories: first, second, and third, and constructs corresponding data processing paths accordingly. This hierarchical mechanism ensures the automation of data processing while enabling differentiated processing of different data categories, thereby improving data filtering efficiency and enhancing the consistency and reliability of output data. Furthermore, the system performs truncation processing on each individual score and the overall score within a preset range to avoid the impact of abnormal fluctuations on the scoring results, thus improving the stability of the scoring results and the consistency of evaluation results across different batches of data.

[0157] Furthermore, while outputting the classification results, the system can also simultaneously output corresponding anomaly information, including anomaly type, anomaly index, and the location of the anomaly segment. Based on this anomaly information, the system can support the location and analysis of target data, eliminating the need to check all data item by item in subsequent processing. Instead, it can directly process the anomaly segments, thereby reducing data screening costs and improving processing efficiency.

[0158] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0159] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices.

[0160] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 2 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0161] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.

[0162] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0163] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0164] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.

[0165] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0166] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0167] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0168] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0169] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0170] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0171] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0172] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0173] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0174] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for anomaly detection and usability verification of embodied data, characterized in that, The method includes: Obtain the embodied data to be verified, which includes multimodal perception data as well as state data and action data associated with timestamps; Using the timestamp as a unified time reference, the multimodal perception data, the state data, and the action data are time-aligned, and the temporal correspondence between the data is extracted. Based on the temporal correspondence and the correlation between the changes in the action data and the response changes in the state data within a preset time range, anomaly detection processing is performed on the embodied data. The anomaly detection processing includes video quality detection based on image pixel features, data integrity detection based on data existence and continuity, temporal synchronization detection based on time deviation, and state-action rationality detection based on the relationship between state changes and action changes. The results of the anomaly detection processing are fused to generate a comprehensive evaluation result, and based on the comprehensive evaluation result and preset judgment rules, the usability judgment result of the embodied data is determined.

2. The method according to claim 1, characterized in that, The timing synchronization detection specifically includes: The overall alignment error is calculated based on the alignment deviations of the multimodal sensing data, the state data, and the action data relative to the timestamp. The state-action synchronization error is calculated based on the correspondence between the change in state data at adjacent time points and the action data at the previous time point. When the overall alignment error or the state action synchronization error is greater than a preset error threshold, it is determined that there is a synchronization anomaly.

3. The method according to claim 1, characterized in that, The rationality detection of the state action specifically includes: If the change in the state data between two adjacent time points is greater than the state transition threshold, it is determined that there is an abnormal state transition. When the motion data at adjacent time points exhibits repeated directional changes and the magnitude of these changes exceeds the motion jitter threshold, motion jitter is determined to exist. If the change in the action data at multiple consecutive time points is less than the action stillness threshold and the change in the state data is less than the state stillness threshold, it is determined that there is a stagnant segment.

4. The method according to claim 1, characterized in that, The data integrity check specifically includes: Detect whether the multimodal perception data, the state data, the action data, and the timestamp are missing; The sampling interruption rate is calculated based on the time interval between adjacent timestamps; When the sampling interruption rate exceeds the time interruption threshold, a sampling interruption anomaly is determined to exist.

5. The method according to claim 1, characterized in that, The multimodal sensing data includes red-green-blue video frame sequences and depth video frame sequences; the video quality detection specifically includes: Based on the pixel brightness information and image sharpness information of the red, green and blue video frame sequence, black frames, white frames and blurry frames are identified; Based on the pixel differences between adjacent red, green and blue video frames and the changes in the corresponding depth video frames, frozen segments of the image are identified. Based on the distribution of invalid depth values ​​in the depth video frame sequence, identify depth failure frames; The abnormal proportions of the black frames, white frames, blurred frames, frozen frames, and depth failure frames are statistically analyzed, and the video quality detection results are generated based on the abnormal proportions.

6. The method according to claim 1, characterized in that, The process of fusing the results of the anomaly detection processing to generate a comprehensive evaluation result includes: Scores are calculated for the video quality detection, data integrity detection, temporal synchronization detection, and state action rationality detection, respectively. The scores of each item are weighted and combined to obtain a comprehensive score that represents the overall evaluation result.

7. The method according to claim 6, characterized in that, The determination of the usability of the embodied data based on the comprehensive evaluation results and preset judgment rules includes: When the comprehensive score is greater than or equal to the first threshold, the embodied data is determined to be the first type of data, and the first type of data is included in the target data set; When the overall score is less than the second threshold, the embodied data is determined to be the second type of data, and the second type of data is isolated or removed. When the overall score is between the first threshold and the second threshold, the specific data is determined to be third type of data, and the corresponding abnormal details are output; wherein, the first threshold is greater than the second threshold.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.