Metadata set driven robot fine control heterogeneous data real machine simulation evaluation method, medium and equipment
By constructing a unified metadata dataset and task graph, and combining real machine and simulation collaborative verification, the evaluation problem of robot precision manipulation tasks is solved, realizing cross-platform unified evaluation and efficient and reliable evaluation results, and improving the automation and interpretability of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-17
AI Technical Summary
In existing technologies, the evaluation of robot fine manipulation tasks suffers from data heterogeneity, lack of task graph representation, complexity of the evaluation process, and anti-cheating challenges, resulting in inconsistent evaluation results, insufficient accuracy and credibility, and a lack of a systematic and efficient evaluation framework.
We adopt a meta-dataset-driven approach, construct a unified meta-dataset, use task graphs for modeling and evaluation, combine real machine and simulation for collaborative verification, and introduce an anti-cheating mechanism to ensure the integrity and credibility of the evaluation data.
It achieves unified evaluation across platforms and tasks, improves evaluation efficiency and interpretability, ensures the fairness and accuracy of evaluation results, and has good scalability and automation.
Smart Images

Figure CN121680280A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to robot precision control technology, and more particularly to a method, medium, and equipment for unified evaluation and collaborative verification of heterogeneous data driven by meta-datasets for robot precision control, as well as a method for real-machine simulation. Background Technology
[0002] With the rapid development of robotics technology, improving the precision manipulation capabilities of robots has become one of the core challenges in the field of robot applications. Precision manipulation tasks typically involve grasping, manipulating, and moving complex objects, requiring high-precision and robust control strategies. Currently, precision manipulation technology is widely used in various fields such as industrial automation, service robots, and medical robots, and requires rigorous evaluation and verification in both simulation and real-world environments.
[0003] In evaluating robots' precision manipulation tasks, especially in environments with multiple data sources and heterogeneous platforms, several pressing challenges need to be addressed:
[0004] 1. Data Heterogeneity: Current robot evaluation typically relies on datasets from different sources, including simulation datasets and real-machine experimental datasets. Because these datasets differ significantly in task definition, annotation methods, sensor configurations, and control parameters, it is difficult to compare evaluation results across platforms or tasks, and a unified standard and methodology are lacking.
[0005] 2. Lack of Task Graph Representation: Fine-grained manipulation tasks can usually be broken down into multiple sub-tasks (atomic actions), and these sub-tasks have clear execution order, constraints, and dependencies. In existing technologies, it is difficult to effectively model and represent the execution process of these sub-tasks within a unified framework, making it difficult for task evaluation to comprehensively and systematically examine the integrity, continuity, and robustness of the operation.
[0006] 3. Complexity of the evaluation process: Task evaluation in real-world environments often consumes significant time and resources, and it is difficult to comprehensively cover all possible operational scenarios and environmental disturbances. Although simulation environments can provide richer scenario simulations and efficient evaluation mechanisms, the gap between simulation and real machines leads to deviations in the mapping of simulation results to the real environment, affecting the accuracy and reliability of the evaluation.
[0007] 4. Anti-fraud and Verifiability: Ensuring the authenticity, integrity, and immutability of evaluation data during multi-party evaluation and data sharing has become a critical issue. Currently, many evaluation processes suffer from data tampering and falsified results, casting doubt on the fairness and transparency of the evaluation process.
[0008] To address the above issues, existing technologies primarily rely on manual annotation and experience-based evaluation methods, lacking a systematic, standardized, and efficient evaluation framework. This limits the advancement of precision robot manipulation technology, particularly posing significant challenges to the comparability, interpretability, and scalability of evaluation and verification across multiple platforms and data sources.
[0009] Therefore, there is an urgent need for a new method and system that can transform and evaluate heterogeneous datasets based on a unified standard and framework, and provide a more accurate, reliable and efficient robot fine manipulation evaluation solution through real machine and simulation co-verification. Summary of the Invention
[0010] To address the aforementioned issues, this invention provides a method, medium, and equipment for real-machine simulation evaluation of heterogeneous data for precise robot manipulation driven by metadata datasets. By combining real-machine and simulation environments, it evaluates and optimizes precise robot manipulation tasks through collaborative verification. This invention is applicable to fields such as imitation learning, visual-language-action models, and closed-loop control systems, enabling unified evaluation across platforms and tasks, and improving the efficiency and interpretability of the evaluation process.
[0011] The specific plan is as follows:
[0012] Firstly, a method, medium, and device for real-machine simulation evaluation of heterogeneous data for precise robot manipulation driven by meta-datasets are provided, including the following steps:
[0013] Step 1: Construct a unified meta-dataset. Through a predefined stable action library and attribute library, perform unified modeling of atomic actions, object attributes, and states involved in the fine manipulation of the robot. Establish object adapters and robot adapters respectively, and abstract the geometric and topological information of task-related objects, robot coordinate system relationships, and sensor configurations into a unified descriptive form. Based on the stable action library, attribute library, object adapter, and robot adapter, preconstruct a directed acyclic task graph template for describing the fine manipulation task, thereby forming the schema foundation of the unified meta-dataset.
[0014] Step 2: Based on the unified metadata dataset, the heterogeneous dataset is transformed. For simulation datasets and real machine logs from different sources, the task annotations, object information, robot execution trajectories, and sensor observation data are parsed. The object adapter and robot adapter are used to map the original data into unified object descriptions and unified robot descriptions. Based on the stable action library and attribute library, the original control sequences are summarized into atomic action sequences. Then, the atomic action sequences are matched and instantiated with the pre-built task graph template to generate instantiated task graphs and corresponding task cards with annotations of preconditions, effect conditions, and constraints. This transforms the heterogeneous dataset into fine-grained manipulation task samples in the unified metadata dataset format.
[0015] Step 3: Design evaluation criteria based on the task graph. Based on the instantiated task graph, design an evaluation index system bound to the task graph structure, including a Process Score (PS) model to characterize the quality of the execution process, and an Area Under Curve (AUC) metric to characterize the success rate as a function of task difficulty. The Process Score model, using task graph nodes as granularity, quantifies and aggregates the effectiveness of each atomic action, the legality of preconditions, the rationality of the execution order, and constraint violations into a task-level PS. The AUC is calculated based on the perturbation axis in the simulation scenario and on the covariate quantile in the real machine scenario. Furthermore, define evidence level and coverage metrics according to the evaluation configuration to characterize the credibility and coverage of the evaluation results, and provide rules for combining PS and AUC into a unified final score.
[0016] Step 4: Implement collaborative evaluation using real and simulation. In the real-device scenario, deploy an anti-cheating client to collect the strategy execution process. Through strategy fingerprint proof, action log integrity proof, and audio / video and clock alignment mechanisms, bind the model, control code, and calibration parameters to the specific evaluation session and generate real-device execution evidence. Based on the real-device execution data, align the actual trajectory to the corresponding instantiated task graph, calculate the real-device success rate and the corresponding process score PS, and use the Theta Encoder to structure the scene configuration, geometric relationships, contact attributes, and sensor configuration of a single real-device execution into an environmental parameter vector θ. In the simulation environment, reconstruct the corresponding scene based on the environmental parameter vector θ, perform multiple conditional replays for the same θ, and use the Posterior Predictive Inference (PPI) method to obtain the large-sample conditional success rate and its confidence interval. Simultaneously, in the pure simulation scenario, based on the preset disturbance axis batch operation strategy, calculate the success rate and process score PS under each disturbance level and obtain the corresponding AUC. Finally, generate a unified evaluation result and report according to the scoring rules, evidence level, and coverage design in Step 3.
[0017] In a second aspect, a computer storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0018] In a second aspect, an electronic device is provided, including a memory and one or more processors, the memory being used to store one or more programs; when the one or more programs are executed by the one or more processors, they implement the method described above.
[0019] Compared with the prior art, the significant advantages of this invention are:
[0020] (1) Unified metadata and cross-platform evaluation capabilities
[0021] This invention addresses the issues of inconsistent evaluations and difficulty in comparison caused by heterogeneous data sources and platform differences in existing technologies by introducing a unified meta-dataset framework. Traditional robot evaluation typically relies on a single data source or a specific platform, making it difficult to handle data from different simulation systems and real-world experiments. This invention, however, transforms datasets from different sources into a unified evaluation format based on predefined stable action and attribute libraries, as well as object and robot adapter definitions. This not only enhances the comparability of evaluation results but also provides a standardized foundation for seamless integration between different platforms and tasks, significantly improving the efficiency and accuracy of cross-platform and cross-task evaluations.
[0022] (2) Modeling and evaluation of fine-grained manipulation tasks based on task graphs
[0023] This invention employs a task graph as a modeling tool for fine-grained manipulation tasks. It decomposes tasks into multiple independently evaluable sub-tasks (atomic actions) and represents the sequence, constraints, and execution relationships of these tasks using a directed acyclic graph (DAG). This method effectively solves the problems of unclear task execution logic and inflexible task flow modeling in existing technologies. Each task node can be closely associated with specific operational goals, environmental states, and constraints, and the dependencies between nodes can be intuitively expressed through the graph structure. This provides a systematic and standardized framework for evaluating fine-grained manipulation capabilities, further improving the interpretability and reproducibility of task execution.
[0024] (3) Innovation of anti-cheating and data verification mechanisms
[0025] This invention provides an anti-fraud client that combines Policy Fingerprint Proof (PFA), Action Merkel Log (AML), Verifiable Execution Experiment (VEE), dual audio and video watermarking, and multi-clock alignment mechanisms to record and verify the entire evaluation process. These technologies effectively prevent data tampering and result fraud, ensuring the integrity and credibility of the evaluation data. This innovative anti-fraud mechanism has significant advantages in robot evaluation, especially in cross-platform evaluation scenarios with multiple participants, ensuring the authenticity of the data and the fairness of the evaluation, and avoiding the problems of susceptibility to human interference in traditional methods.
[0026] (4) Highly efficient real machine and simulation co-verification
[0027] This invention combines real-machine evaluation with simulation evaluation, calibrating simulation evaluation results using the posterior predictive inference (PPI) method. Specifically, it utilizes the simulation environment to repeatedly replay real execution conditions, combining a large number of simulation evaluation samples to narrow the confidence interval and improve the accuracy of the evaluation results. In this process, simulation evaluation not only provides a large amount of evaluation data but also reveals potential deviations between simulation and real execution under condition replay, further enhancing the reliability of the evaluation. Compared with traditional methods that rely solely on simulation or real-machine evaluation, this invention's real-machine and simulation collaborative verification scheme has significant advantages, enabling a more comprehensive evaluation of robot strategy performance in different environments and allowing for effective comparison and optimization across multiple tasks and platforms.
[0028] (5) Scalability and versatility
[0029] The method of this invention has good scalability and can be adapted and transformed to different datasets as needed. By adapting to different robot control systems and task types, this invention can flexibly handle a variety of different robot fine manipulation tasks. This method is not only applicable to traditional industrial robots, but also to modern robot control strategies such as Visual Language Action Modeling (VLA) and imitation learning. With the introduction of more datasets, the system can automatically adjust and expand the existing evaluation framework according to new data patterns, thereby maintaining efficiency and adaptability in the rapidly evolving field of robotics.
[0030] (6) A systematic and automated evaluation framework
[0031] This invention automates and standardizes the evaluation of robot fine manipulation strategies by constructing a systematic task graph and evaluation framework. Traditional robot evaluation methods often rely on manual annotation and debugging, which is time-consuming and prone to bias. This invention, however, significantly improves evaluation efficiency and reduces the possibility of human error by automating task generation, data conversion, and evaluation report generation. Furthermore, evaluation results can be quantified using a unified scoring system (such as AUC and process score PS), ensuring the comparability and transparency of the evaluation results.
[0032] (7) Enhanced explainability and decision support capabilities
[0033] The fine-grained task graph and scoring model provided by this invention allow users to clearly understand the robot's behavior in different task executions and to conduct detailed analysis of the effects, execution conditions, and constraint violations of each action step. This interpretability not only helps evaluators identify potential problems in robot operation but also provides data support for further optimization of control strategies and a scientific basis for strategy improvement. Attached Figure Description
[0034] Figure 1 This is a flowchart of the method of the present invention.
[0035] Figure 2 This diagram illustrates the construction process of the meta-dataset in this invention. It details the transformation process from the original heterogeneous dataset (including simulation data and real machine logs) to a unified meta-dataset.
[0036] Figure 3 The result is used to generate a transformation task graph for heterogeneous datasets.
[0037] Figure 4 This is a schematic diagram of step-by-step evaluation based on the task graph. Detailed Implementation
[0038] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0039] Combination Figure 1 A method for unified evaluation and collaborative verification of heterogeneous data driven by meta-datasets for precise robot manipulation, specifically including the following steps:
[0040] Step 1 constructs a unified meta-dataset. Using a predefined stable action library and attribute library, it uniformly models the atomic actions, object attributes, and states involved in the fine manipulation of the robot, and establishes object adapters and robot adapters respectively. It abstracts task-related object geometry and topology information, robot coordinate system relationships, and sensor configurations into a unified descriptive form. Based on the stable action library, attribute library, object adapter, and robot adapter, it pre-constructs a directed acyclic task graph template to describe the fine manipulation task, thus forming the schema foundation of the unified meta-dataset. Specifically, in existing technologies, various datasets often use their own independent task naming rules and action encoding methods, making it difficult to compare and analyze the evaluation results of the same strategy on different datasets under a unified standard. Therefore, this invention first provides a unified task description schema, uniformly mapping various simulation datasets and real machine logs to the same semantic and structural space, thereby forming a unified meta-dataset that can be shared among different data sources.
[0041] To achieve the aforementioned unification, step one first defines a stable action library, which is used to abstract the basic operations in the process of fine robot manipulation at the semantic level. This stable action library is denoted as the action set. Each element represents an atomic action type. Action set This includes at least reaching, contacting, grasping, aligning, pulling, pushing, rotating, inserting, placing, pressing, switching, and waiting actions. A stable action library allows for the segmentation and categorization of underlying control trajectories from different datasets into unified atomic action types, achieving unified semantic modeling of the fine manipulation process. Building upon this, step one further constructs an attribute library to uniformly describe the geometric and semantic features of the manipulated object and its environment. The attribute library is denoted as attribute set P, where each attribute characterizes a semantic slot in the fine manipulation task. Attribute set P includes at least object identification attributes, component identification attributes, facing attributes, spatial orientation attributes, color attributes, state attributes, and target pose attributes. Object identification attributes are used to distinguish different manipulated objects; component identification attributes are used to identify local components of an object; orientation attributes describe the orientation of each face of an object; spatial direction attributes describe the direction of operation in a specified reference frame; color attributes identify the color of an object or component; state attributes identify the discrete opening / closing, insertion, or locking states of an object or mechanism; target pose attributes describe the desired position and orientation of the object or end effector in space during the task. The attribute library is linked to a tolerance table, which sets uniform allowable deviation ranges for geometric position, orientation, and state determinations, so that the action effects can be judged according to a unified standard in subsequent evaluations.
[0042] To map object and robot information from datasets of different sources to the unified semantic space described above, step one establishes object adapters and robot adapters. The object adapter converts the object identifiers and geometric information recorded in each dataset into a unified object description format. For each object, the object adapter provides a description including at least a bounding box, joint axes, slots or mating planes, surface normals, and optional color and status labels, allowing the geometric and semantic information of the same type of object from different datasets to be represented using the same fields in the unified metadata set. The robot adapter abstracts the geometric structure and control characteristics of different robot platforms into a unified robot description format. The robot adapter provides the transformation relationships between the world coordinate system, robot base coordinate system, end effector coordinate system, and various camera coordinate systems, records the installation configuration of static cameras, wrist cameras, depth cameras, tactile sensors, and force / torque sensors, and describes characteristics such as control command update frequency and control latency. Through the object and robot adapters, different datasets can use a unified data structure to represent task-related object and robot information from the perspective of the unified metadata set.
[0043] Based on the aforementioned unified semantics and unified adaptation, this invention models the execution process of fine manipulation tasks using a graph structure. Since fine manipulation tasks of robots can typically be broken down into several sub-steps with clear semantics, each sub-step corresponds to an atomic action with a clear goal and boundary, such as reaching, contacting, grasping, aligning, pulling, and inserting. There are strict sequential dependencies, selection relationships, or optional relationships between different sub-steps. Recording these steps only in linear sequence or plain text form makes it difficult to simultaneously express complex logical structures such as branching, parallelism, optional paths, and error recovery paths, and also makes it difficult to analyze and evaluate the process quality of individual steps during evaluation. To solve these problems, this invention uses a graph to model fine manipulation tasks in the specific implementation of the unified metadata dataset. Each sub-step is represented as a node in the graph, and the sequence constraints and logical relationships between sub-steps are represented as directed edges in the graph. The unified representation simultaneously encodes "what operations are involved," "how operations are connected," and "under what conditions a transition is possible," providing a clear and computable structural foundation for subsequent process scoring calculations and collaborative verification between real machines and simulations.
[0044] Specifically, in this invention, the task graph used to describe a single fine manipulation task is represented in the form of a directed acyclic graph, and its mathematical representation is as follows:
[0045]
[0046] Where V is the set of nodes in the task graph; E is the set of directed edges in the task graph; Σ is the set of semantic attributes actually used in this task instance, and Σ is a subset of the attribute set P; Θ is the set of constraint parameters related to this task; Π is the set of predicates used to define preconditions, effect conditions, and violation conditions; and A is the global setting information related to this task. At the node level, each node v in set V represents an atomic action instance, and node v is represented as:
[0047] ,
[0048] Here, `type` is the action type identifier, indicating the type of atomic action corresponding to this node, and its value comes from the action set `𝒜`; `slots` is the semantic slot binding result, used to instantiate the attributes in the attribute set `Σ` to the current atomic action. `slots` represents a set of "attribute name – attribute value" pairs, such as "object = drawer", "part = drawer handle", "direction = +x direction in the object coordinate system", etc.; `pre` is the precondition, used to describe the conditions that the environment should meet before executing the atomic action corresponding to the current node. `pre` consists of one or more predicates in the predicate set `Π` and their logical combinations; `effect` is the effect condition, used to describe the state that the environment should reach after the atomic action corresponding to the current node is successfully completed. `effect` also consists of predicates in the predicate set `Π` and their logical combinations; `constraints` are the constraint conditions, used to limit the physical or safety restrictions during the execution of the current atomic action. `constraints` references one or more parameters in the constraint parameter set `Θ`, used to represent the joint angle range, sliding stroke range, contact gap, tolerance threshold, and upper limits of speed and force, etc.; `timeout` The execution timeout parameter limits the maximum time allowed for the current atomic action to execute. If the effect condition is not met within the timeout period, the node is judged as a timeout failure or a preset exception handling logic is triggered. At the edge level, each directed edge in set E... Directed edges represent the execution dependencies between task steps. Represented as:
[0049]
[0050] in, The source node represents the starting node of the current directed edge; `target` represents the terminal node of the current directed edge; `gate` is a logical gateway type parameter used to represent the target node's dependencies on its multiple predecessor nodes. The value of `gate` can include AND, OR, or OPTIONAL. When the `gate` type of multiple incoming edges is AND, it means that the target node is only allowed to enter the executable state if all corresponding predecessor nodes are in a successful state and their effect conditions are met. When the `gate` type of multiple incoming edges is OR, it means that the target node can enter the executable state as long as any one of its predecessor nodes is in a successful state. When an incoming edge is marked as OPTIONAL, it means that the path is an optional execution path, and the success or failure of this predecessor node does not affect the execution of the target node when other predecessor conditions are met.
[0051] Set Σ represents the set of semantic attributes actually used in a specific task instance within the task graph T. Σ selects object identifiers, component identifiers, orientation information, spatial orientation, color attributes, state attributes, and target pose attributes related to the current task from attribute set P. These are used for specific binding in the node's `slots` field, explicitly recording the attribute range upon which the task depends. Set Θ represents the set of constraint parameters used in the task graph T, summarizing the constraints that the task must adhere to under the robot platform and physical environment. These include joint angle limits, sliding stroke limits, maximum permissible speed and force, contact gap thresholds, and safety boundary parameters related to collisions and boundary violations. The node's `constraints` field references parameters from set Θ to provide a unified constraint description for the execution process of individual atomic actions. Set Π represents the set of predicates used in the task graph T, used to construct the node's preconditions (`pre`), effect conditions (`effect`), and logical conditions related to constraint violations. Each predicate is defined on object attributes, robot states, and sensor observations. Symbol A represents the global settings information of the task graph T, including at least the reference coordinate system settings for interpreting spatial orientation and pose, the tolerance table version identifier used for effect judgment, and environmental parameters for recording the evaluation environment configuration.
[0052] After defining the action set 𝒜, attribute set P, object adapter, robot adapter, and task graph form T = ⟨V, E, Σ, Θ, Π, A>, this invention can pre-define a set of task graph templates for common fine manipulation tasks, such as opening a drawer, unscrewing a bottle cap, picking up a specified pen by color, and picking up a knife handle. Each task graph template provides node, edge, attribute, constraint, and predicate configurations under the aforementioned unified pattern. The task graph templates, together with the action library, attribute library, and adapter, constitute the pattern foundation of the unified metadata dataset, enabling subsequent task samples from different simulation platforms and real machine systems to be uniformly mapped to task graph instances conforming to this pattern, thereby completing the construction of the unified metadata dataset in step one.
[0053] Step 2 transforms the heterogeneous dataset based on the unified metadata dataset, specifically including: obtaining raw multimodal data from the target simulation dataset and / or real machine logs, including scene image frames, robot joint states synchronized with the images, end effector poses, and task-related metadata; inputting the scene image frames and their corresponding robot state information into a pre-trained Vision-Language Model (VLM), and combining the unified object description and unified robot description obtained in Step 1 through the object adapter and robot adapter, providing object categories, component candidates, and robotic arm types as prompts to the Vision-Language Model, which then performs retrieval and discrimination within the unified semantic space determined by the stable action library and attribute library, and outputs a set of fine manipulation candidate tasks supported in the current data and the natural language task instructions corresponding to each candidate task;
[0054] Based on candidate tasks and corresponding natural language task instructions output by the visual language model, the system calls upon the set of atomic action types defined in the stable action library and the attributes such as objects, parts, directions, and states defined in the attribute set to discretize and normalize the continuous action descriptions output by the visual language model. This transforms the descriptions into atomic action sequences composed of atomic actions and their semantic slots. Furthermore, candidate tasks that do not conform to the unified metadata dataset pattern constraints are filtered and corrected. Figure 2 As shown;
[0055] Based on the normalized atomic action sequence and its semantic slot binding results, a task graph template matching the task instruction is retrieved from the pre-built task graph template library, and the task graph format given in step 1 is used. For the set of nodes in the task graph The set of edges E, the set of semantic attributes Σ, and the set of constraint parameters Θ are instantiated to generate an instantiated task graph labeled with preconditions, effect conditions, and constraints, such as... Figure 3 As shown.
[0056] After obtaining the instantiated task graph, a fine-grained manipulation task data generation file suitable for the dataset is generated by combining the recording format of the target dataset itself and the scene configuration file. The fine-grained manipulation task data generation file includes at least the task identifier, natural language task instructions, corresponding instantiated task graph identifier, the binding relationship between the objects and components involved in the unified attribute set Σ, the mapping relationship between the object and the original trajectory file or scene configuration file, and the parameter configuration for sampling task instances under different difficulty levels and different covariate conditions in the dataset. This enables the unified conversion of heterogeneous simulation datasets and real machine logs and the generation of fine-grained manipulation tasks based on the unified meta-dataset model.
[0057] Step 3 is a task graph definition strategy evaluation method based on a unified metadata set, specifically including: the task graph representation obtained in step 1.
[0058]
[0059] Based on this, the evaluation of the strategy execution process is divided into two parts: process score and success rate on the difficulty axis. Here, V is the set of nodes in the task graph, where each node represents an atomic action instance; E is the set of directed edges; Σ is the set of semantic attributes used by the task instance; Θ is the set of constraint parameters; Π is the set of predicates; and A is global setting information. First, a process score function is constructed based on the task graph node set V, calculating the single-step process score for the atomic action corresponding to each node, and then performing weighted aggregation at the task level. For the j-th node in the task graph... Define the single-step process score of this node as The specific calculation formula is as follows:
[0060]
[0061] in, This is a performance indicator used to measure the achievement of a node. The effect condition is the degree to which the relevant predicates in the set of predicates Π satisfy the condition. This is a precondition validity indicator, used to measure the validity of nodes during effect determination. Whether the precondition pre is satisfied; This is an execution order rationality index, used to indicate the degree of consistency between the actual execution order and the topological order given by E in the task graph; A normalized metric used to constrain the number of violations or the severity of violations, which comprehensively reflects whether the set of constraint parameters Θ referenced during node execution is violated; This is a truncation function used to restrict the input value x to the interval [0,1].
[0062] For the entire task graph T, the task-level process score (PS) is defined as the weighted average of the individual process scores of each node. The specific calculation formula is as follows:
[0063]
[0064] in, For nodes The weighting parameters are used to highlight the contribution of key steps to the overall process quality. This is useful when a node cannot be calculated due to a lack of necessary observations or unavailable predicates. At that time, the corresponding weights can be... The value is set to zero to gray out undecidable steps. Through the above definition, the process score PS can quantify the rationality and standardization of the strategy execution process within a unified task graph structure.
[0065] Secondly, a success rate-difficulty curve is constructed on the difficulty axis based on task samples in a unified metadata set, and an area under the curve (AUC) metric is defined. In the simulation evaluation scenario, one or more perturbation parameters are selected as the difficulty axis, such as initial pose offset, contact friction coefficient, sensor noise amplitude, etc. The perturbation axis is discretized into L difficulty levels, denoted as the set of difficulty levels.
[0066]
[0067] Regarding the difficulty level Let the total number of trials performed under this difficulty condition be . The number of successful trials was The success rate at that difficulty level is... Defined as:
[0068]
[0069] Under discrete difficulty axis conditions, the success rate versus area under the difficulty curve of the simulation scenario It can be approximated as:
[0070]
[0071] in, For the corresponding difficulty level The weighting coefficients, The settings can be configured according to the range of disturbance parameters or the representativeness of each level in the experimental design to meet the requirements.
[0072]
[0073] In real-device evaluation scenarios, one or more covariates are selected as the natural difficulty axis, such as occlusion ratio, target distance, motion blur level, video frame rate, or control latency. Each covariate is then quantized and binned, dividing its values into Q quantile intervals, denoted as the quantile set.
[0074]
[0075] For quantile intervals Let the total number of trials for performing the task within this interval be . The number of successful trials was The success rate within that quantile interval Defined as:
[0076]
[0077] The corresponding success rate in real-device scenarios – area under the quantile curve It can be defined as:
[0078]
[0079] in, These are the weighting coefficients for each quantile interval. It can be set to equal weight, or based on the proportion of each quantile interval in the actual distribution. (This satisfies...)
[0080]
[0081] After obtaining the process score PS and the area under the success rate-difficulty curve in simulation or real-machine scenarios, this invention defines a unified comprehensive evaluation score to simultaneously reflect the final success rate and the quality of the execution process. The comprehensive score F can be defined as:
[0082]
[0083] AUC is selected based on the evaluation scenario. or , This is the process scoring weighting coefficient, used to control the degree of influence of process scoring on the final score. When only the final success rate is available and there is no complete process score, it can be... Setting it to zero degenerates the overall score F into a simple AUC metric. Based on the above definitions, step 3 presents an evaluation method for fine-grained manipulation strategies, building upon the task graph in the unified metadata set. This allows for quantitative comparisons of strategy performance across different datasets, platforms, and perturbation conditions using the same scoring system.
[0084] Step 4, the real-device-simulation joint evaluation scheme, specifically involves: During the real-device evaluation phase, an anti-cheating client is deployed to record and verify the entire process of each real-device evaluation session. At the start of the session, the anti-cheating client acquires the session identifier and task configuration, calculates the fingerprints of the current strategy model, control code, and calibration parameters, and generates a strategy fingerprint proof to bind the session to a specific strategy implementation. During task execution, the client records the observation data, control actions, and timestamps at each moment and constructs an action Merkle log, organizing the records at each moment into a Merkle tree structure and periodically reporting to the Merkle root to ensure the integrity and immutability of the execution log. Simultaneously, based on minor perturbation instructions issued by the server, the client performs controllable perturbations on sensor signals or control commands and detects the strategy's response to the perturbations, thus forming a verifiable execution experiment to determine whether the strategy is being executed online in real time rather than offline playback. Watermarks associated with the session identifier and time information are embedded in the video and audio streams, and multiple clock alignments and deviation records are performed on the local time, robot control time, and server time to constitute evidence of real-device execution.
[0085] After obtaining the execution data from the actual machine, based on the task graph in the unified metadata set, the actual execution trajectory is aligned to the node set V of the corresponding task graph. The process score for this execution is calculated using the process scoring formula defined in step 3, and the process score for a single execution on the actual machine is denoted as . Where n is the real-machine test number. Let be the total number of tests performed on the same task under a given configuration in the real-machine evaluation. The number of successful trials was The corresponding success rate on a real device Defined as:
[0086]
[0087] Simultaneously, the process scores of all real-machine tests can be averaged or weighted to obtain the task-level real-machine process score. ,For example:
[0088]
[0089] After the actual device execution is complete, through The Theta Encoder structures the environment and execution conditions of a single real-machine test into an environment parameter vector. For the nth real-machine test, the environmental parameter vector is denoted as... It includes at least the following components: scene lighting and background configuration parameters, robot end effector initial pose, key object initial pose, occluder geometry and attitude parameters, contact and friction-related proxy parameters, and sensor and communication parameters such as camera external participation frame rate and control link latency, i.e.:
[0090]
[0091] in, This indicates scene configurations such as lighting, background, and frame rate. This indicates the initial pose of the end effector and the key object. Represents the geometry and orientation of the occluding object. Indicates contact and friction proxy parameters, This indicates parameters such as sensor extrinsic parameters, sampling frequency, and delay.
[0092] During the simulation evaluation phase, based on the environmental parameter vector Reconstruct the initial scenario configuration in the simulation environment to match the nth real-machine test, and, while keeping the policy model and its parameters unchanged, perform a specific test on each... Perform conditional replay. For a given... The simulation experiment is repeated K times under different random seed conditions. Let K be the total number of simulation experiments under these conditions, and let the number of successful experiments be _____. Then under the condition Simulation success rate estimate
[0093]
[0094] Here, the subscript "ppi" indicates that the success rate is a simulation estimate obtained through posterior predictive inference. By analyzing each... By analyzing the binomial distribution or its approximate distribution of the K replay trials, further calculations can be performed. The confidence interval, for example, at a given confidence level. The next confidence level is obtained With upper confidence boundary .
[0095] For multiple real machine condition samples This invention can obtain an overall PPI success rate estimate through a weighted average. ,For example:
[0096]
[0097] And through all The replay experiments were statistically analyzed together to obtain the overall confidence interval. During the simulation replay, the corresponding simulation process score can also be calculated based on the unified task graph. The process score of the k-th simulation experiment in the n-th conditional replay is denoted as... The scores are then aggregated to obtain the PPI simulation process score. ,For example:
[0098]
[0099] In obtaining a high success rate with real devices Real device process scoring PPI success rate And PPI simulation process scoring Subsequently, this invention defines a real-machine comprehensive score and a PPI comprehensive score to reflect the overall performance under real-world conditions and under conditional simulation prediction, respectively. Real-machine comprehensive score It can be defined as:
[0100]
[0101] PPI Composite Score It can be defined as:
[0102]
[0103] Through the aforementioned joint evaluation scheme of real machine and simulation, this invention, while ensuring the authenticity of real machine evaluation and anti-cheating capabilities, utilizes simulation condition replay based on environmental parameters to obtain large-sample statistical results, thereby achieving real machine-simulation collaborative verification of robot fine manipulation strategies within a unified meta-dataset framework. Figure 4 As shown.
[0104] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.
Claims
1. A method for evaluating a robot fine manipulation heterogeneous data real machine simulation driven by metadata sets, characterized in that, Comprising the following steps: Step 1, constructing a unified metadata set, through a pre-defined stable action library and attribute library, atomically modeling the atomic actions, object attributes and states involved in fine manipulation of robots, and respectively establishing object adapters and robot adapters, abstracting task-related object geometry and topological information, robot coordinate system relationship and sensor configuration into a unified description form; Based on the stable action library, attribute library, object adapter and robot adapter, a directed acyclic task graph template for describing fine manipulation tasks is pre-constructed, thereby forming the mode basis of the unified metadata set Meta-Dataset; Step 2, based on the unified metadata set, converting heterogeneous data sets, for simulation data sets and real machine logs of different sources, parsing the task annotations, object information, robot execution trajectories and sensor observation data therein, using the object adapter and robot adapter to map the original data to unified object description and unified robot description, and based on the stable action library and attribute library, the original control sequence is induced into an atomic action sequence; Then match and instantiate the atomic action sequence with the pre-constructed task graph template to generate an instantiated task graph with precondition, effect condition and constraint condition annotations and the corresponding task card, thereby converting the heterogeneous data set into a fine manipulation task sample in the format of the unified metadata set; Step 3, based on the task graph, design the evaluation, on the basis of the instantiated task graph, design an evaluation index system bound to the task graph structure, including a process score PS model for describing the quality of the execution process, and an area under the success rate-difficulty curve AUC index for describing the success rate as the task difficulty changes; The process score PS model takes the task graph node as the granularity, quantifies and aggregates the effect achievement degree, legality of preconditions, rationality of execution order and constraint violation of each atomic action into task-level PS, and the area under the success rate-difficulty curve AUC index is calculated based on the disturbance axis setting in the simulation scene and the covariate quantile setting in the real scene; And according to the evaluation configuration, define the evidence level and coverage rate index to describe the credibility and coverage range of the evaluation results, and give the rule of combining PS and AUC into a unified final score; In the real machine scene, the anti-cheating client collects the policy execution process, and through the policy fingerprint proof, action log integrity proof, and audio / video and clock alignment mechanism, the model, control code and calibration parameters are bound to the specific evaluation session and real machine execution data is generated; based on the real machine execution data, the actual trajectory is aligned to the corresponding instantiated task graph, the real machine success rate and the corresponding process score PS are calculated, and the scene configuration, geometric relationship, contact attribute and sensor configuration of a single real machine execution are structured into an environment parameter vector θ through the θ encoder; in the simulation environment, the corresponding scene is reconstructed based on the environment parameter vector θ, multiple conditional replays are performed for the same θ, and a large sample conditional success rate and its confidence interval are obtained by using the posterior predictive inference PPI method; at the same time, in the pure simulation scene, the strategy is batched based on the preset perturbation axis, the success rate and process score PS at each perturbation level are calculated, and the corresponding AUC is obtained, and finally the unified evaluation result and report are generated according to the scoring rules, evidence level and coverage rate designed in step 3.
2. The metadata set driven robot fine manipulation heterogeneous data real machine simulation evaluation method according to claim 1, characterized in that, In step 1, a stable action library is first defined for abstracting the basic operations in the process of fine manipulation of the robot at the semantic level; the stable action library is denoted as an action set where each element represents an atomic action type; Through the stable action library, the underlying control trajectories in different data sets are segmented and summarized into unified atomic action types, realizing unified action semantic modeling of fine manipulation processes; The attribute library is denoted as an attribute set P, wherein each attribute is used to describe a semantic slot in the fine manipulation task; the attribute library is associated with a tolerance table, and the tolerance table sets a unified allowed deviation range for geometric position, attitude and state judgment; The object adapter is established to convert the object identification and geometric information recorded in each data set into a unified object description form; The object adapter gives a description for each object, so that the geometric and semantic information of the same type of object in different data sets is represented in the same field in the unified metadata set; The robot adapter is used to abstract the geometric structure and control characteristics of different robot platforms into a unified robot description form; through the object adapter and the robot adapter, the object and robot information related to the task in different data sets are represented in a unified data structure under the perspective of the unified metadata set; In the specific implementation of the unified metadata set, the fine manipulation task is modeled in the form of a graph, each sub-step is represented as a node in the graph, and the sequence constraints and logical relationships between sub-steps are represented as directed edges in the graph. In the unified representation, "what operations are performed", "how to link between operations", and "under what conditions can the transition be made" are encoded at the same time. Specifically, the task graph used to describe a single fine manipulation task is represented in the form of a directed acyclic graph, and its mathematical representation is: ; Wherein, V is a node set in the task graph; E is a directed edge set in the task graph; Σ is a semantic attribute set actually used in the task instance, Σ is a subset of the attribute set P; Θ is a constraint parameter set related to the task; Π is a predicate set used to define preconditions, effect conditions and exception conditions; A is global setting information related to the task; at the node level, each node v in the set V represents an atomic action instance, and the node v is represented as: ; Wherein, type is an action type identifier, used to indicate the atomic action type corresponding to the node, taking values from the action set A; slots is a semantic slot binding result, used to instantiate the attributes in the attribute set∑ to the current atomic action; pre is a precondition, used to describe the conditions that the environment should satisfy before executing the atomic action corresponding to the current node, pre is composed of one or more predicates and their logical combinations in the predicate set∑; effect is an effect condition, used to describe the state that the environment should reach after the successful completion of the atomic action corresponding to the current node, effect is also composed of predicates and their logical combinations in the predicate set∑; constraints is a constraint condition, used to limit the physical or safety restrictions in the execution process of the current atomic action, constraints reference one or more parameters in the constraint parameter set∑; timeout is an execution timeout parameter, used to limit the maximum time allowed for the execution of the current atomic action, when the effect condition is not met within the timeout limit, the node will be judged as timeout failure or trigger the preset exception handling logic; at the level of edges, each directed edge in the set E represents the execution dependency between task steps, the directed edge represents the execution dependency between task steps, the directed edge is represented as: ; wherein, is a source node, representing the starting node of the current directed edge; is a target node, representing the terminating node of the current directed edge; gate is a logic gateway type parameter, used to represent the dependency relationship of the target node to its multiple predecessor nodes; when the gate type of multiple incoming edges is AND, it means that only when all corresponding predecessor nodes are in a successful state and meet their effect conditions, the target node is allowed to enter an executable state; when the gate type of multiple incoming edges is OR, it means that as long as any one of the predecessor nodes is in a successful state, the target node can enter an executable state; when a certain incoming edge is marked as OPTIONAL, it means that this path is an optional execution path, and the success or failure of this predecessor node does not affect the execution of the target node when other predecessor conditions are met.
3. The metadata set driven robot fine manipulation heterogeneous data real machine simulation evaluation method according to claim 1, characterized in that, In step 2, the heterogeneous data sets are converted based on the unified metadata set, specifically including: obtaining original multi-modal data from the target simulation data set and / or the real machine log, including scene image frames, robot joint states synchronized with the image, end effector poses, and task-related meta information; inputting the scene image frames and the corresponding robot state information into a pre-trained visual language model VLM, and combining the unified object description and unified robot description obtained through the object adapter and the robot adapter in step 1, providing the object category, component candidate and robot type as prompt information to the visual language model, and performing retrieval and discrimination in the unified semantic space determined by the stable action library and the attribute library, outputting a set of fine manipulation candidate tasks that can be supported in the current data and a natural language task instruction corresponding to each candidate task; Based on the candidate tasks and the corresponding natural language task instructions output by the visual language model, the atomic action type set defined in the stable action library and the object, component, direction, state attribute defined in the attribute set are called to discretize and standardize the continuous action description output by the visual language model, convert it into an atomic action sequence composed of atomic actions and semantic slots, and filter and correct candidate tasks that do not conform to the mode constraints of the unified metadata set; According to the normalized atomic action sequence and the semantic slot binding result thereof, a task graph template matching the task instruction is searched in a pre-constructed task graph template library, and the task graph form given in step 1 is utilized , the node set in the task graph , the edge set E, the semantic attribute set Σ, and the constraint parameter set Θ are instantiated to generate an instantiated task graph with precondition, effect condition and constraint condition annotations; wherein Π is a predicate set; A is global setting information; After obtaining the instantiated task graph, a fine manipulation task data generation file suitable for the data set is generated in combination with the recording format and scene configuration file of the target data set, which at least includes task identification, natural language task instruction, corresponding instantiated task graph identification, binding relationship of objects and components in the unified attribute set Σ, mapping relationship with the original trajectory file or scene configuration file, and parameter configuration for sampling task instances under different difficulty levels and different covariant conditions in the data set, thereby realizing unified conversion and fine-grained fine operation task generation of heterogeneous simulation data sets and real machine logs based on the unified metadata set mode.
4. The metadata set driven robotic fine-motor manipulation heterogeneous data real machine simulation evaluation method of claim 1, wherein, Step 3 specifically includes: the task graph representation obtained in step 1. Based on this, the evaluation of the strategy execution process is divided into two parts: process score and success rate on the difficulty axis. Here, V is the set of nodes in the task graph, where each node represents an atomic action instance; E is the set of directed edges; Σ is the set of semantic attributes used by the task instance; Θ is the set of constraint parameters; Π is the set of predicates; and A is global setting information. First, a process scoring function is constructed based on the task graph node set V to calculate the single-step process score for the atomic action corresponding to each node, and then performs weighted aggregation at the task level. For the j-th node in the task graph... Define the single-step process score of this node as The specific calculation formula is as follows: ; wherein, is an effect achievement indicator, used to measure the degree to which the effect condition effect of the node is satisfied by the relevant predicates in the predicate set Π; is a precondition legality indicator, used to measure whether the precondition pre of the node is satisfied at the time of effect determination; is an execution order rationality indicator, used to represent the degree of consistency between the actual execution order and the topological order given by E in the task graph; is a normalized indicator of constraint violation frequency or violation severity, used to comprehensively reflect whether the constraint parameter set Θ referenced in the execution process of the node is violated; is a truncation function, used to limit the input value x within the interval [0, 1]; For the entire task graph T, the task-level process score PS is defined as the weighted average of the node single-step process scores, and the specific calculation formula is: ; in, For nodes The weighting parameters are used to highlight the contribution of key steps to the overall process quality. This is useful when a node cannot be calculated due to a lack of necessary observations or unavailable predicates. At that time, the corresponding weights will be... Set to zero to gray out undecidable steps; through the above definition, the process score PS can quantify the rationality and standardization of the strategy execution process under a unified task graph structure; Secondly, based on the task samples in the unified metadata set, a success rate-difficulty curve on the difficulty axis is constructed, and a success rate-difficulty curve area index AUC is defined based on the area under the curve; in the simulation evaluation scene, one or more disturbance parameters are selected as the difficulty axis, and the disturbance axis is discretized into L difficulty levels, denoted as ; For a difficulty level , the total number of trials to perform a task under this difficulty condition is , where the number of successful trials is , then the success rate under this difficulty level is defined as: ; Area under the success rate - difficulty curve of the simulated scenario under the discrete difficulty axis condition is approximately expressed as: ; wherein, is a weight coefficient corresponding to the difficulty level is a weight coefficient corresponding to the difficulty level is set according to the value range of the disturbance parameter or the representative of each level in the experimental design, and meets ; In the real machine evaluation scene, one or more covariates are selected as natural difficulty axes, each covariate is divided into Q quantile intervals, and the quantile set is denoted as ; For quantile interval , let the total number of trials in which the task is performed be , and the number of successful trials be , then the success rate in this quantile interval is defined as: ; Corresponding true scene success rate - area under the quantile curve is defined as: ; wherein, is a weight coefficient for each quantile interval, is set as equal weight, or is set according to the proportion of each quantile interval in the real distribution; satisfies ; After obtaining the process score PS and the area under the success rate-difficulty curve in the simulation or real scene, a unified comprehensive evaluation score is defined to reflect the terminal state success rate and the execution process quality; the comprehensive score F is defined as: ; AUC is selected based on the evaluation scenario. or , This is the process scoring weighting coefficient, used to control the degree of influence of process scoring on the final score; when only the final success rate is available and there is no complete process score, it will be... Setting it to zero degenerates the overall score F into a simple AUC metric. Based on the above definition, step 3 provides an evaluation method for fine-grained manipulation strategies on the basis of the task graph in the unified metadata set, enabling quantitative comparison of strategy performance under different datasets, platforms, and perturbation conditions within the same scoring system.
5. The metadata set driven robotic fine-motor manipulation heterogeneous data real machine simulation evaluation method of claim 1, wherein, Step 4 specifically comprises: in the real machine evaluation stage, by deploying an anti-cheating client, the whole process of each real machine evaluation session is recorded and proved; the anti-cheating client obtains the session identifier and task configuration at the beginning of the session, calculates the fingerprint of the current strategy model, control code and calibration parameters, and generates a strategy fingerprint proof for binding the current session with a specific strategy implementation; during the task execution process, the observation data, control action and timestamp of each time point are recorded, and an action Merkle log is constructed, the records of each time point are organized into a Merkle tree structure, and the Merkle root is reported periodically to ensure the integrity and non-tamperability of the execution log; at the same time, according to the micro disturbance instruction issued by the server, controllable disturbance is made to the sensor signal or control command, and the response of the strategy to the disturbance is detected, so as to form a verifiable execution experiment for determining whether the strategy is online real execution or offline playback; the watermark associated with the session identifier and time information is embedded in the video and audio stream, and the local time, robot control time and server time are multi-clock aligned and deviation recorded to form real machine execution evidence; After obtaining the real machine execution data, based on the task graph in the unified metadata set, the real execution trajectory is aligned to the node set V of the corresponding task graph, and the process score of this execution is calculated by using the process score PS defined in step 3, and the process score of single real machine execution is recorded as , wherein n is the real machine test number; and the total test number of the same task under the given configuration in the real machine evaluation is recorded as , wherein the number of successful tests is , the corresponding real machine success rate is , and is defined as: ; At the same time, all the true machine test process score is averaged or weighted aggregation, get the task level true machine process score : ; After the end of the real machine execution, the environment and execution conditions of the single real machine test are structured into an environment parameter vector by the encoder ; for the nth real machine test, the environment parameter vector is denoted as which includes at least the following components: scene lighting and background configuration parameters, robot end effector initial pose, key object initial pose, occluder geometry and pose parameters, contact and friction related proxy parameters, and camera extrinsic and frame rate, control link delay sensor and communication parameters, namely: ; wherein, represents lighting, background and frame rate scene configuration, represents initial poses of end effector and key objects, represents geometry and pose of occluder, represents contact and friction agent parameters, represents sensor extrinsic, sampling frequency and latency parameters; In the simulation evaluation phase, according to the environment parameter vector In the simulation environment, the initial scene configuration matching the nth real machine test is reconstructed, and each is conditionally replayed; for a given , the simulation test is repeatedly executed K times under different random seed conditions, and the total number of simulation tests under the condition is K, wherein the number of successful tests is , then the simulation success rate estimation value under the condition ; ; where the subscript "ppi" indicates that the success rate is a simulation estimate obtained by the posteriori prediction inference method; by analyzing the binomial distribution or its approximate distribution of K playback tests under each , the confidence interval of is further calculated; the lower confidence limit and the upper confidence limit are obtained at a given confidence level ; For multiple real-condition samples The overall PPI success rate estimate is obtained by weighted average : ; And through the joint statistics of all replay tests, the overall confidence interval is obtained ; in the simulation replay process, the corresponding simulation process score is also calculated based on the unified task graph, and the process score of the kth simulation test in the nth conditional replay is recorded as , and it is aggregated to obtain the PPI simulation process score : ; After obtaining the real machine success rate , the real machine process score , the PPI success rate , and the PPI simulation process score , the real machine comprehensive score and the PPI comprehensive score are defined respectively to reflect the comprehensive performance under the real environment and the conditional simulation prediction respectively; the real machine comprehensive score is defined as: ; PPI composite score defined as: ; Through the above real machine and simulation joint evaluation scheme, while ensuring the authenticity and anti-cheating ability of real machine evaluation, large sample statistical results are obtained by using simulation condition replay based on environmental parameters, and real machine-simulation collaborative verification of fine control strategy of robot in the unified metadata set framework is realized.
6. A computer storage medium, characterized in that A computer program is stored thereon, and the computer program is executed by a processor to realize the steps of the method of any one of claims 1 to 5.
7. An electronic device, comprising: A memory is included, and one or more programs are stored in the memory; when the one or more programs are executed by one or more processors, the method of any one of claims 1 to 5 is realized.
Citation Information
Cited By
An unmanned intelligent agent autonomous flight decision and verification method and system
CN122239495A