Robot operation data segmentation and trusted labeling method and system
Patent Information
- Application Number
- CN202610982456.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2026-06-18
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-25
AI Technical Summary
(1)本发明以关节状态、末端执行器状态、夹持/接触状态和操作对象状态的四维物理变化量为核心依据进行动作边界检测,能够降低仅依赖视频帧变化、单一触觉突变或人工观察导致的边界漂移、边界遗漏和边界碎片化问题。
Smart Images

Figure CN122817902A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of embodied intelligent data processing, robot operation data governance, and artificial intelligence annotation technology. In particular, it relates to a method and system for performing motion segmentation, structured semantic annotation, credibility assessment, hierarchical processing, and consistency verification of preceding and subsequent states on multimodal data of continuous robot operation based on robot joint states, end effector states, contact / gripping states, operation object states, and task semantic information. Background Technology
[0002] The training, transfer, and evaluation of embodied intelligence models, robot operation strategy models, and visual-language-action models rely on a large amount of high-quality robot operation scenario data. During task execution, robots simultaneously generate multi-source heterogeneous data, including video frames, image frames, voice or text task commands, joint angles, joint velocities, joint torques, end effector poses, gripper opening and closing degrees, contact forces, torques, tactile or pressure signals, object poses, object states, and task results. Unlike ordinary image or video annotation, robot operation data not only needs to indicate "what's in the scene," but also when the robot starts executing a certain action, when it completes the action, what the object being manipulated is, what state the object or environment should be in before execution, what state the object or environment actually becomes after execution, and whether the action satisfies a continuous state transition relationship with the preceding and following actions.
[0003] Current methods for labeling robot task data typically rely on manual viewing of long-duration videos or playback of sensor recordings to manually determine the start and end times of actions, action categories, and objects being manipulated. This approach suffers from low labeling efficiency, high labor costs, strong subjectivity in action boundaries, poor consistency among labelers, and difficulty in expressing complex task state relationships. When a robot task consists of multiple consecutive actions such as "approaching an object, grasping an object, moving an object, placing an object, and releasing an object," relying solely on visual frame changes or subjective human judgment can easily lead to premature, delayed, or fragmented action boundary segmentation, thus affecting the quality of subsequent training data.
[0004] Existing technology CN119538161A discloses a method for constructing a human-machine collaboration dataset and detecting anomalies by integrating a multimodal large model. This scheme uses human-machine action tags to control a virtual digital human and a virtual robotic arm to perform virtual human-machine collaboration. It acquires human collision points, multi-angle human-machine scene images, and joint posture data. After time-stamp alignment of the above data, a multimodal dataset is formed. A multimodal large model is then used to generate a visual-language instruction dataset and to achieve anomaly detection in human-machine collaboration. While this scheme illustrates the general approach to multimodal data synchronization, dataset construction, and anomaly detection, its technical objective and processing object are concentrated on anomaly detection in virtual human-machine collaboration. It does not propose a closed-loop scheme for automatic action boundary segmentation, physical signal boundary scoring, atomic skill pre- and post-state verification, and label credibility grading for continuous robot operation data.
[0005] Existing technology CN113894779A discloses a multimodal data processing method for robot interaction. This scheme acquires target visual and tactile information data, performs fusion processing based on a multimodal data fusion model to obtain fused command information data, and outputs it to the robot's motion components. The focus of this scheme is on fusing visual and tactile information to improve the accuracy of interactive command recognition. It does not address the automatic segmentation and structured annotation of long-term robot task datasets, nor does it disclose the use of four-dimensional physical changes—joint state, end effector state, gripping / contact state, and object state—for motion boundary scoring, nor does it disclose the use of consistency between the pre-motion and post-motion states to determine the reliability of the annotation.
[0006] Existing technology CN119312066A discloses a method and electronic device for calculating the quality of user annotations for large models. This scheme evaluates the annotation quality of crowdsourcing users through weak model triggering, annotation score, instruction diversity score, and quality ranking score. The evaluation object of this scheme is "whether the annotation user is reliable," rather than "whether the annotation of a certain action segment in the robot's continuous operation data conforms to the physical state and task state transition rules." It does not handle the relationships between joints, end effectors, contact forces, object states, and action segment state transitions in robot operations, nor does it provide a structured labeling system and automatic quality verification mechanism for robot tasks.
[0007] The paper "An Atomic Skill Library Construction Method for Data-Efficient Embedded Manipulation" discloses the idea of decomposing end-to-end tasks into subtasks through a vision-language planning module, then abstracting them into atomic skill definitions, and constructing an atomic skill library through data acquisition and fine-tuning of a vision-language action model. This paper focuses on "how to construct, expand, and call the atomic skill library," but does not disclose using atomic skill templates as the basis for verifying the pre-state, post-state, and state transition rules of robot operation annotation results. It also does not disclose using a four-dimensional weighted average of joint states, end effector states, gripping / contact states, and object states for action boundary scoring, nor does it disclose a closed-loop system for reliable robot operation data annotation that combines five-dimensional reliability weighting, hierarchical processing, and consistency verification of pre- and post-states.
[0008] Furthermore, while existing motion segmentation research has disclosed the basic ideas for online boundary detection or motion segmentation of robot operations, it usually focuses on visual sequence segmentation, motion boundary models, or motion category recognition itself. It does not provide a complete technical solution that combines joint motion changes, end effector motion changes, gripping / contact state changes, and object state changes to form a boundary score, and further couples it with atomic skill templates, consistency verification of preceding and following states, and five-dimensional credibility grading.
[0009] Therefore, existing technologies still lack a method and system that can automatically segment action segments based on physical signals in real or simulated robot work scenarios, associate action segments with atomic skill templates, object state evolution, and task execution results, and perform hierarchical processing, correction, and backtracking of automatic annotation results through five-dimensional credibility and state consistency rules. Summary of the Invention
[0010] The technical problems to be solved by this invention are: how to accurately detect action boundaries by making full use of physical signals such as joint state, end effector state, gripping / contact state and operation object state in continuous robot operation scenarios; how to automatically map the segmented action fragments into structured annotations with action category, operation object, previous state, subsequent state and execution result; and how to improve the reliability, traceability and reusability of automatic annotation results through five-dimensional confidence weighting, hierarchical processing and consistency verification of previous and subsequent states.
[0011] To address the aforementioned technical problems, this invention provides the following technical solution: a method and system for robot operation data segmentation and reliable annotation, comprising the following steps:
[0012] S1: Multimodal data access and unified timeline construction. Acquire video frames or image frames, task instructions, joint states, end effector states, gripping / contact states, force / torque or tactile states, manipulated object states, and task result states during robot operations; construct a unified timeline at preset time intervals and synchronize continuous state data and discrete visual data in time.
[0013] S2: Construction of Robot Operation State Vector and Physical Changes. A robot state vector is formed on a unified time axis. Joint motion changes, end effector motion changes, gripping / contact state changes, and object state changes are extracted to form four-dimensional physical changes for action boundary detection.
[0014] S3: Automatic motion boundary detection driven by physical signals. Normalizes and jointly weights the four-dimensional physical changes to obtain the motion boundary score at each moment; based on threshold judgment, local extremum screening, shortest duration constraint and adjacent boundary merging rules, candidate motion boundaries are determined, and the robot's complete operation process is divided into a set of motion segments.
[0015] S4: Semantic matching of action fragments based on atomic skill templates. Establish an atomic skill template library containing action categories, operation object types, tool types, preceding states, following states, task constraint rules, and state transition rules; extract the semantic vector and physical dynamic features of each action fragment and match them with the atomic skill templates to determine candidate atomic skill types.
[0016] S5: Structured annotation generation. Based on the visual content of the action segment, joint state, end effector state, gripping / contact state, object state, task instructions, and candidate atomic skill templates, structured annotation results are generated; the structured annotation results include at least the action category, operation object, tool type, action start time, action end time, previous state, subsequent state, state change, execution result, and evidence source.
[0017] S6: Five-dimensional credibility weighting and consistency verification of preceding and following states. The credibility of visual language reasoning, skill template matching, object recognition, physical boundary, and state consistency are calculated separately and then weighted and fused. Simultaneously, based on the preceding and following states and task constraint rules in the atomic skill template, consistency verification is performed on the state transitions within a single action segment and the connection between preceding and following states between adjacent action segments.
[0018] S7: Graded processing and output of annotation results. Based on the comprehensive credibility and state consistency verification results, the annotation results are divided into three levels: automatic pass, manual review, and re-analysis; for action segments that do not meet the state consistency requirements, candidate correction results are generated or re-segmentation and re-inference are triggered; finally, a structured annotation dataset with a timeline is output.
[0019] It should be noted that the unified time axis construction, linear interpolation, and nearest neighbor mapping in this invention are only basic preprocessing steps for multi-source robot operation data to enter the unified state space. The core improvement of this invention is that: action boundary scores are formed in the unified state space using four-dimensional physical change quantities, action segments are matched with atomic skill templates containing pre-state and post-state, and the annotation results are graded and closed-loop corrected by combining five-dimensional credibility and state consistency verification.
[0020] Preferably, in step S1, the unified time axis is set as follows:
[0021] For continuous state data such as joint status, end effector status, force / torque, and tactile feedback, linear interpolation is used for time synchronization.
[0022] in, Indicates the first Continuous modes at a unified time The state value, and The adjacent sampling times for this mode.
[0023] For discrete visual data such as image frames or video frames, the nearest neighbor mapping method is used to obtain the visual frames corresponding to a unified time:
[0024] After completing time synchronization, construct a unified robot state vector:
[0025] in, Indicates visual state. , , These represent joint angle, joint velocity, and joint torque, respectively. and These represent the position and orientation of the end effector, respectively. Indicates the opening and closing degree of the grippers. Indicates the contact force or force / torque state. Indicates the state of the object being operated on. Indicates the semantic state of the task.
[0026] Preferably, in steps S2 and S3, the change in joint motion is:
[0027] The change in motion of the end effector is:
[0028] in, The distance representing the attitude change can be expressed using a rotation matrix or the angular distance corresponding to a quaternion.
[0029] The change in clamping / contact state is:
[0030] in, Indicates the contact state or clamping state. This is an indicator function.
[0031] The change in the state of the operation object is:
[0032] in, and These represent the position and orientation of the manipulated object, respectively. Indicates the state of an object, such as "on the desktop", "held", "lifted up", or "located in the target area".
[0033] Robust normalization of the four-dimensional changes:
[0034] Action boundary score:
[0035] in, ,and .
[0036] At a certain moment A candidate action boundary is determined when the following conditions are met:
[0037] in, For boundary thresholds, Let be the radius of the local extremum window. For the candidate boundary set... Further execution of the shortest action duration constraint and adjacent boundary merging yields the boundary sequence. And form a collection of action segments:
[0038] Preferably, in step S4, each atomic skill template is defined as:
[0039] in, Indicates the action category, Indicates the object type, Indicates the tool or end effector type. Indicates the previous state. Indicates the subsequent state. This represents the task constraint rules. This indicates the dynamic feature description corresponding to the template.
[0040] Action clips With template The matching score is:
[0041] in, For the semantic vector of the action segment, For template semantic vectors, For dynamic feature consistency score, The score is determined by matching the previous state. The score is determined by matching the subsequent state, and all weights are non-negative and sum to 1. The candidate atomic skill template is as follows:
[0042] Preferably, in step S5, the structured annotation result is represented as follows:
[0043] in, Indicates the action category, Indicates the object being operated on. Indicates the tool or end effector type. and Indicates the start and end times. and These represent the observed preceding and following states, respectively. Indicates a change in state. Indicates the execution result. Indicate the source of the evidence.
[0044] Preferably, in step S6, the state consistency loss is:
[0045] in, The state distance function, This indicates whether there are conflicts in the order rules, object relationships, or state transition rules between adjacent action segments. The state consistency reliability is:
[0046] The five-dimensional credibility includes the credibility of visual language reasoning. Skill template matching credibility Reliability of object recognition Physical boundary credibility State consistency reliability The overall credibility is:
[0047] in, ,and .
[0048] Preferably, the physical boundary confidence level can be calculated from the boundary scores near the start and end boundaries of the segment:
[0049] in, For the Sigmoid function, This is the smoothing coefficient.
[0050] Preferably, in step S7, the hierarchical processing rule is as follows:
[0051] For action segments that enter the re-analysis level, the system selects from the candidate label set. Select the correction annotation:
[0052] Therefore, this invention forms a closed loop encompassing action boundary detection, atomic skill semantic annotation, five-dimensional credibility assessment, and consistency verification between preceding and subsequent states, rather than simply performing multimodal temporal alignment, ordinary visual / tactile fusion, user annotation quality scoring, or the construction of a general atomic skill library. These steps work in tandem: physical boundary scores define the time range of action segments, atomic skill templates define the state transition relationships that action segments should satisfy, and five-dimensional credibility and hierarchical processing determine whether the annotation results are automatically approved, subject to manual review, or trigger re-analysis, thereby improving the reliability and traceability of robot operation data annotation.
[0053] This invention also provides a method and system for robot operation data segmentation and reliable annotation, comprising the following modules: Multimodal data access module: used to access images / videos, task instructions, joint status, end effector status, gripping / contact status, force / torque status, tactile status, object status, and task result status during robot operation.
[0054] Time synchronization and state vector construction module: used to construct a unified time axis, interpolate continuous data, perform nearest neighbor mapping on discrete visual frames, and generate a unified robot state vector.
[0055] Physical change calculation module: used to calculate joint motion changes, end effector motion changes, clamping / contact state changes, and object state changes.
[0056] Action boundary detection module: used to generate action boundary scores based on four-dimensional physical changes, determine action boundaries based on thresholds, local extrema and duration constraints, and output a set of action segments.
[0057] Atomic Skill Template Matching Module: Used to maintain the atomic skill template library and match action fragments with atomic skill templates that include both preceding and following states.
[0058] Semantic reasoning and annotation module: used to call visual language models or rule-based reasoning models, and combine action fragments, physical features, task instructions and candidate templates to generate structured annotation results.
[0059] Credibility Assessment Module: This module calculates the credibility of visual language reasoning, skill template matching, object recognition, physical boundary, and state consistency, and outputs the overall credibility.
[0060] State consistency verification module: Used to check whether the preceding and following states within an action segment meet the template rules, and to check whether the state transitions between adjacent action segments are consistent.
[0061] The hierarchical processing and dataset output module is used to classify the annotation results into automatic pass, manual review or re-analysis based on the comprehensive credibility and state consistency verification results, and output a structured dataset with a time axis.
[0062] Compared with the prior art, the present invention has at least the following beneficial effects: (1) The present invention uses the four-dimensional physical changes of joint state, end effector state, clamping / contact state and operation object state as the core basis for motion boundary detection, which can reduce the boundary drift, boundary omission and boundary fragmentation caused by relying solely on video frame changes, single tactile mutations or manual observation.
[0063] (2) The present invention matches action segments with atomic skill templates that include action category, operation object, tool type, pre-state, post-state and task constraint rules, so that the annotation results not only include action name, but also express object, state change, execution result and evidence source.
[0064] (3) This invention incorporates state transition relationships such as “the hand should be empty and the gripper should be open before grabbing”, “the object should be clamped or lifted after grabbing”, and “the object should be located in the target area and the gripper should be released after placement” into automatic verification through consistency verification of the preceding and following states. This can detect logical errors that are difficult to detect by simple large-scale model semantic reasoning.
[0065] (4) The present invention adopts a five-dimensional confidence weighting, which incorporates visual language reasoning, template matching, object recognition, physical boundary and state consistency into the evaluation, so as to avoid the direct acceptance of erroneous labels with high confidence of a single model but inconsistent physical state.
[0066] (5) This invention combines automatic approval, manual review and re-analysis, which can concentrate manual review resources on low-confidence or state conflict samples, thereby improving the annotation efficiency of large-scale robot datasets.
[0067] (6) The structured timeline labels output by this invention can be used for robot operation strategy training, embodied intelligence model fine-tuning, action behavior analysis, task failure attribution and cross-robot platform data reuse, and each label can be traced back to the corresponding video frame, physical curve and object state evidence. Attached Figure Description
[0068] Figure 1 This is an overall flowchart of the method in an embodiment of the present invention; Figure 2 This is a flowchart of physical signal-driven action boundary detection in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the consistency verification of atomic skill template matching and pre- and post-state states in an embodiment of the present invention. Figure 4 This is a schematic diagram of the action segment timeline of the task of grasping the wooden block in an embodiment of the present invention; Figure 5 This is a block diagram of the system module structure in an embodiment of the present invention.
[0069] Figure labeling: 100 - Multimodal data access module; 200 - Time synchronization and state vector construction module; 300 - Physical change calculation module; 400 - Action boundary detection module; 500 - Atomic skill template matching module; 600 - Semantic reasoning annotation module; 700 - Credibility assessment module; 800 - State consistency verification module; 900 - Hierarchical processing and dataset output module. Detailed Implementation
[0070] The technical solution of the present invention will be clearly and completely described below with reference to embodiments. The described embodiments are only for illustrating the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can combine, substitute, or modify the various embodiments without departing from the concept of the present invention.
[0071] Example 1: Intelligent Labeling Method for Continuous Robot Operation Data like Figure 1 As shown, this embodiment provides an intelligent annotation method for continuous robot operation data. When a robot performs a task, the system first receives data such as video frames, joint angles, joint velocities, joint torques, end effector poses, gripper opening and closing degrees, contact forces, object poses, and task commands. The system constructs a unified time axis at fixed time intervals, performs linear interpolation on continuous modalities, and performs nearest neighbor mapping on discrete image frames to form a unified robot state vector. This time synchronization step ensures that each modality can be compared and combined at the same time, and the subsequent action boundaries are determined by the joint scoring of four-dimensional physical changes.
[0072] Subsequently, the system calculates four-dimensional physical changes at each unified instant. Joint motion changes reflect the velocity, acceleration, and torque changes of each joint of the robotic arm; end effector motion changes reflect changes in the end effector's position and orientation; gripping / contact state changes reflect abrupt changes in gripper opening, contact force, and contact state; and object state changes reflect changes in the object's position, orientation, and state category. After normalizing these changes, the system calculates motion boundary scores and determines the motion boundaries based on local extrema and threshold values.
[0073] For example, when a robot transitions from approaching an object to grasping it, the end effector speed may decrease, the gripper opening angle may change rapidly, and the contact force may increase sharply. When the robot completes the grasp and begins to move the object, the object's state changes from "on the table" to "gripped / lifted." These physical changes typically reflect the transition between motion stages more stably than simple changes in image frames. Therefore, in this embodiment, motion boundary detection focuses on four-dimensional physical changes: joints, end effector, gripping / contact, and object state, while visual frames are primarily used for object recognition and object state estimation.
[0074] After an action fragment is generated, the system establishes an atomic skill template library. Each template includes not only the action category but also the pre-action states that must be met before the action is executed and the post-action states that must be achieved after the action is executed. For example, the pre-action states of the "grab" template may include the end being close to the target object, the gripper opening, and the target object not being gripped; the post-action states may include the gripper closing, the contact force being within the effective range, and the target object being gripped or moving with the end. The pre-action states of the "place" template may include the target object being gripped and the end being above the target area; the post-action states may include the target object being in the target area, the gripper releasing, and the object remaining stationary.
[0075] The system extracts semantic vectors and physical dynamic features from each action segment and calculates its matching score with each atomic skill template. If the end of an action segment approaches the target object, the gripper changes from open to closed, the contact force increases, and the object subsequently moves with the end, then its matching score with the "grab" template is the highest. Based on this, the system determines candidate skill types and generates structured annotation results by combining a visual language model or a rule-based reasoning model. Unlike simply building a skill library, this embodiment uses the template's pre-state, post-state, and task constraint rules as conditions for verifying the credibility of the annotation.
[0076] After generating structured annotations, the system does not directly accept all model outputs. Instead, it calculates five dimensions of credibility: visual language reasoning credibility evaluates the stability of the model's output semantics; skill template matching credibility evaluates the degree of matching between action segments and candidate templates; object recognition credibility evaluates whether the object being operated on is reliably identified; physical boundary credibility evaluates whether there are stable physical change peaks near the start and end times of the action; and state consistency credibility evaluates whether the connection between the preceding state, the following state, and adjacent action segments conforms to the rules. These five dimensions of credibility collectively determine the processing level of the annotation results.
[0077] If the overall credibility is high and the state consistency credibility meets the requirements, the system automatically approves the annotation; if the overall credibility is in the middle range or there is a slight state conflict, the system pushes the annotation for manual review; if the overall credibility is low or there is a serious state conflict, the system triggers re-segmentation, re-template matching, or re-semantic reasoning. Thus, the system can avoid erroneous annotations that "appear semantically correct but are not supported by the physical state" from entering the training dataset.
[0078] Example 2: Complete annotation process for the task of grabbing wood blocks like Figure 4 As shown, this embodiment uses "grabbing a wooden block from the table and placing it on a tray" as an example to illustrate the complete process of the present invention. The robot's task instruction is: "Grab the wooden block on the table and place it on the tray." During the execution process, the robot generates the following data: RGB video frames, six-axis joint angles of the robotic arm. Joint velocity End effector position and posture Gripper opening and closing degree End contact force Center position of the wooden block Wooden block posture and the state of the wooden block .
[0079] First, the system unifies the above data to On the time axis, a state vector is formed. The system calculates the four-dimensional physical changes and obtains the action boundary score. In this embodiment, the boundary score exhibits a significant peak in the following stages: (1) After the robot approaches the wooden block, it begins to decelerate and adjust its end-effector posture to form the boundary between "approaching" and "grabbing preparation"; (2) The gripper changes from open to closed, and the contact force rises rapidly, forming the boundary of the "grabbing" action; (3) The state of the wooden block changes from "on the table" to "being held / lifted", forming the boundary between the "grabbing completed" and "transporting" actions; (4) After the end reaches the top of the pallet, the speed decreases and the position of the wooden block tends to the target area, forming the boundary between the "carrying" and "placing" actions; (5) The grippers open, the contact force decreases, and the state of the wooden block changes to "located in the tray and stationary", forming the "placement completed" boundary.
[0080] Based on this, the system obtains four main action segments:
[0081] For action segments The system extracted the following from the fragment: the end is near the wooden block, the gripper opening changes from large to small, the contact force changes from low to high, and the wooden block subsequently moves with the end. This fragment matches the "grab" atomic skill template and generates the following structured annotation: the action category is "grab", the object is "wooden block", the tool type is "gripper", the preceding state is "wooden block on the table, gripper open, end near the wooden block", the following state is "gripper closed, contact force effective, wooden block is gripped and can move with the end", and the execution result is "success".
[0082] For action segments The system verifies whether its previous state is consistent with... The subsequent states are consistent. If The subsequent state is "the wooden block is clamped", and If the preceding state also observes "the wooden block is clamped and moves with the end", then the two states are consistent; if the visual language model mistakenly interprets... If the label is marked as "moving cup", the credibility of object recognition and the credibility of state consistency will decrease, and the system will transfer the label to manual review or re-analysis.
[0083] In this embodiment, it is assumed that The five-dimensional credibility is , , , , And all five weights are The overall credibility is:
[0084] If the threshold is automatically passed Set as State consistency threshold Set as If the overall credibility of another segment is [value missing], then the annotation will be automatically approved. However, the reliability of state consistency is only For example, if the subsequent state of the "placement" segment still shows that the wooden block has not been released, the system will not automatically pass the test, but will trigger a re-analysis or manual review. This approach can prevent the high-confidence output of a single semantic model from masking conflicts in the actual physical state.
[0085] Ultimately, the timeline annotation data output by the system can be represented as:
[0086] The above data can be directly used for robot operation strategy training, action failure analysis, task replay retrieval, and cross-platform skill data governance.
[0087] Example 3: System Implementation like Figure 5 As shown, this embodiment provides an intelligent annotation system for multimodal data of robot operations. The multimodal data access module 100 accesses various modal data through a robot operating system, data logger, or simulation platform; the time synchronization and state vector construction module 200 synchronizes multi-frequency data; the physical change calculation module 300 outputs four-dimensional changes in joint, end effector, gripping / contact, and object states; the motion boundary detection module 400 outputs a set of motion segments; the atomic skill template matching module 500 maintains a template library and performs template matching; the semantic reasoning annotation module 600 calls a visual language model, object detection model, or rule reasoning module to generate structured annotations; the credibility assessment module 700 calculates five-dimensional credibility; the state consistency verification module 800 checks whether state transitions satisfy pre- and post-rules; and the hierarchical processing and dataset output module 900 outputs the final annotated dataset.
[0088] In one implementation, the system also includes a template update unit. This unit is not intended to build a general atomic skill library, but rather to maintain the skill template specifications required for annotation. Once a corrected annotation of a low-confidence segment is manually reviewed and confirmed, the system can supplement the template library with the preceding, following, or constraint rules from the corrected result to improve the consistency of subsequent annotations for similar robot operation data.
[0089] In another implementation, the system also includes an evidence backtracking unit. This unit associates each structured tag with its corresponding video frame, joint curve, end effector trajectory, gripper opening / closing curve, contact force curve, and object state sequence. When a human reviewer verifies a tag, they can simultaneously view the physical evidence corresponding to that tag, thereby improving review efficiency and consistency.
[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. For different robot platforms, the state vector dimension, change weight, template rules, and classification threshold can be adjusted according to the robot arm's degrees of freedom, end effector type, sensor type, and task type. However, as long as a combination of motion segmentation based on physical signals such as joints / end effects, consistency verification of atomic skill states before and after, and five-dimensional credibility classification processing is adopted, it falls within the scope of the present invention.
Claims
1. A method for robot task data segmentation and reliable annotation, characterized in that, Includes the following steps: S1: Acquire video frames or image frames, task instructions, joint states, end effector states, gripping / contact states, force / torque or tactile states, manipulated object states, and task result states during robot operation, and complete the time synchronization of multimodal data on a unified time axis; S2: Construct a robot state vector on the unified time axis, and calculate the joint motion changes, end effector motion changes, gripping / contact state changes, and object state changes respectively to form four-dimensional physical changes; S3: Normalize and weight the four-dimensional physical change quantities to calculate the action boundary score at each moment; when the action boundary score at a certain moment meets the boundary threshold condition and is the local maximum value within a preset window, the moment is determined as a candidate action boundary, and then the action segment set is obtained. S4: Establish an atomic skill template library, wherein each atomic skill template in the atomic skill template library includes an action category, an operation object type, a tool or end effector type, a preceding state, a following state, a task constraint rule, and a state transition rule; match each action fragment in the action fragment set with the atomic skill templates to determine a candidate atomic skill template for each action fragment; S5: Based on the physical dynamic features of the action fragment, the corresponding task instructions, and the candidate atomic skill templates, a structured annotation result is generated through a visual language model. The structured annotation result includes the action time period, action category, operation object, tool type, observation pre-state, observation post-state, and execution result. S6: Calculate the credibility of visual language reasoning, skill template matching, object recognition, physical boundary, and state consistency to form a five-dimensional credibility assessment; and perform two-layer consistency verification on the structured annotation results: the first layer verifies whether the observation pre-state and observation post-state of a single action segment meet the constraints of the candidate atomic skill template, and the second layer verifies whether the state transition between the post-state and the pre-state of adjacent action segments is consistent. S7: Based on the comprehensive credibility and the consistency verification results of the two-layer state, the structured annotation results are classified into automatic pass, manual review or re-analysis, and a structured credible annotation dataset with time axis is output. The action boundary is jointly triggered by the four-dimensional physical change quantity, and the consistency verification of the two-layer state is performed according to the pre-state, post-state and adjacent action segment state connection rules in the atomic skill template.
2. The method according to claim 1, characterized in that, The unified timeline is Linear interpolation is used for time synchronization of continuous state data such as joint state, end effector state, force / torque, or tactile state; nearest neighbor mapping is used for time synchronization of discrete visual data such as video frames or image frames.
3. The method according to claim 1, characterized in that, The robot state vector includes at least the visual state, joint angles, joint velocities, joint torques, end effector position, end effector posture, gripper opening degree, contact force, object state, and task semantic state.
4. The method according to claim 1, characterized in that, Action boundary score Calculate using the following formula: in, , , and These represent the normalized changes in joint motion, end effector motion, gripping / contact state, and manipulated object state, respectively. .
5. The method according to claim 1, characterized in that, Determining candidate action boundaries includes: when the action boundary score at a certain moment is greater than or equal to a preset boundary threshold, and the action boundary score at that moment is the local maximum value within a preset window, that moment is determined as a candidate action boundary; and the candidate action boundaries are corrected by the shortest action duration constraint and the adjacent boundary merging rule.
6. The method according to claim 1, characterized in that, Each atomic skill template includes an action category, an object type, a tool or end effector type, a preceding state, a following state, task constraint rules, and a dynamic feature description; the preceding state is used to define the state that the object, robot, or environment should meet before the action is executed, and the following state is used to define the state that the object, robot, or environment should reach after the action is executed.
7. The method according to claim 1, characterized in that, The structured annotation results include at least the action category, the object of operation, the tool type, the action start time, the action end time, the pre-observation state, the post-observation state, the state change, the execution result, and the source of inference evidence; the source of inference evidence records the visual features, physical signal fragments, and matching template identifiers on which the annotation results are generated.
8. The method according to claim 1, characterized in that, The state consistency reliability is calculated based on the state consistency loss, which includes: the difference between the observed pre-state and the candidate atomic skill template pre-state, the difference between the observed post-state and the candidate atomic skill template post-state, the difference between the current action segment post-state and the next action segment pre-state, and whether there are conflicts in sequence rules, object relationships, or state transition rules between adjacent action segments.
9. The method according to claim 1, characterized in that, The overall credibility Calculate using the following formula: in, Indicating the credibility of visual language reasoning, Indicates the credibility of skill template matching. Indicates the credibility of object recognition. Indicates the credibility of physical boundaries. Indicates the reliability of state consistency. .
10. The method according to claim 1, characterized in that, The hierarchical processing includes: when the overall credibility is greater than or equal to the automatic pass threshold and the state consistency credibility is greater than or equal to the state consistency threshold, the structured annotation result is automatically passed; when the overall credibility is within the manual review range or there is a slight state conflict, the structured annotation result is sent for manual review; when the overall credibility is lower than the re-analysis threshold or there is a serious state conflict, the corresponding action segment is re-segmented, re-template matched, or re-semantic reasoning is performed.
11. A robot operation data segmentation and reliable annotation system, characterized in that, include: The system comprises a multimodal data access module, a time synchronization and state vector construction module, a physical change calculation module, an action boundary detection module, an atomic skill template matching module, a semantic reasoning annotation module, a credibility assessment module, a state consistency verification module, and a hierarchical processing and dataset output module; the system is used to execute the method described in any one of claims 1 to 10.
12. The system according to claim 11, characterized in that, The motion boundary detection module is used to generate motion boundary scores based on joint motion changes, end effector motion changes, clamping / contact state changes, and object state changes, and output a set of motion segments according to boundary thresholds, local extrema, and shortest motion duration constraints.
13. The system according to claim 11, characterized in that, The state consistency verification module is used to verify whether the observation pre-state and observation post-state of a single action segment meet the pre-state and post-state rules in the candidate atomic skill template, and to verify whether the state transitions between adjacent action segments are consistent.
Citation Information
Patent Citations
Multi-modal data processing method applied to robot interaction
CN113894779A
Large model user annotation quality calculation method and electronic equipment
CN119312066A
Man-machine cooperation data set construction and anomaly detection method fusing multi-modal large model
CN119538161A