Human preference labeling method of robot operation track and corresponding robot learning system

By dividing the data sets of the robot's operation trajectory attributes according to safety, efficiency and performance, and clustering the data sets using DTW distances to form trajectory selection for labeling personnel, the problems of inconsistent data labeling and low quality in existing tools are solved, and more efficient and consistent data labeling is achieved.

CN120011835APending Publication Date: 2025-05-16HONG KONG IND ARTIFICIAL INTELLIGENCE & ROBOTICS RES & DEV CENT LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510039163.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing robot operating trajectory preference data annotation tools are difficult to ensure the consistency and quality of data, and human evaluators are prone to inattention and fatigue during long-term labeling, which makes it difficult to accurately collect data preferences.

Method used

By dividing the data set of the robot operation trajectory attributes according to safety, efficiency and performance, and clustering the data sets using dynamic time wrapping (DTW) distances, a trajectory is formed to select the preferred trajectory for labelers. This method combines database clustering based on attributes and principles, back-end active query model, and front-end hybrid active user interface, improving the efficiency and consistency of data annotation.

Benefits of technology

This method effectively improves the quality and consistency of the robot's operating trajectory preference data, reduces attention fatigue in human evaluators during the annotation process, and improves the performance and efficiency of the robot learning algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011835A_ABST
    Figure CN120011835A_ABST
Patent Text Reader

Abstract

The invention provides a human preference labeling method for a robot operation track and a corresponding robot learning system. The human preference labeling method comprises a database clustering method based on attributes and principles, a back-end active query model and a front-end hybrid active user interface. The robot trajectory data set is clustered through three main principles: safety, efficiency and performance; for each principle, quantifiable attributes are collected, the quantifiable attributes are connected into a main vector in each time step, and a main vector time sequence of the whole trajectory is formed. Then, utilizing DTW to cluster the trajectory data set according to each main vector sequence; when the preferences are collected according to each principle, the back-end model can calculate weighted average values of several indexes designed according to user requirements in the preference extraction process, for example, the difficulty of comparing track pairs, the divergence among human evaluators and the skewness of marked and unmarked track pairs; the front end presents track video pairs for human evaluators to compare and specify their preferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data labeling method for collecting human preferences for robot operation trajectories and a system for using the labeled data for robot learning and training. Background Art

[0002] Autonomous robots have emerged in various industrial applications, such as pre-operative rehabilitation, companionship, and assembly tasks. Manipulation robots with various artificial intelligence (AI) algorithms typically learn from a large amount of data from human teachers to perform tasks.

[0003] Human preferences are one of the most common types of human data in robot learning algorithms. Existing human data collection tools only provide trajectory videos for human evaluators to annotate their preferences, but a robot needs a large amount of human data to learn a task, and the performance of the robot is highly dependent on the quality of human data. The shortcomings of existing basic tools include: 1) human evaluators may mix various principles when specifying preferences, resulting in inconsistent preference data and bringing additional complexity to robot task learning algorithms; 2) human evaluators may not notice the detailed attributes in trajectory videos; 3) during the preference annotation process of long trajectories, human evaluators may lose focus and easily fatigue.

[0004] Currently, state-of-the-art robot learning algorithms often rely on leveraging skilled human demonstrations, which are expensive to collect and filter. These algorithms often have difficulty maintaining a balance between multiple factors that humans consider critical in robot operation, such as safety and efficiency. In addition, existing methods for learning from human preferences often focus on distinguishing between optimal and suboptimal datasets without explicitly manipulating human preferences for robot learning.

[0005] Existing data labeling tools have been widely used in text, speech, image, and video annotation tasks. However, robot operation trajectories are non-stationary and high-dimensional, which brings higher cognitive load and unique challenges to the data annotation task. Therefore, existing data labeling or annotation tools may not be able to solve the challenge of specifying human preferences for robot operation trajectories. Summary of the invention The object of the present invention is to provide a novel robot trajectory preference labeling method for human evaluators, and a system for using the labeled data for robot learning training.

[0006] The present invention is achieved through the following technical solutions: The method for labeling the human preference of the robot operation trajectory includes dividing the collected data set of the robot operation trajectory according to the following principle attributes: Safety, including collision, distance, and contact force features: efficiency, including speed, trajectory length, time and cost, and; performance, including smoothness, stability, orientation, and posture; For each feature of the principle attribute, a corresponding principal vector related to the principle attribute is established, and the DTW distance between the eigenvalue of the feature corresponding to the attribute in the data set and the corresponding principal vector is calculated, and the data set is clustered accordingly; the clustered data sets are paired to form trajectory pairs, and the trajectory pairs are sent to the annotator for judgment and selection of the optimal trajectory; the optimal trajectory is formed into a training data set for the artificial intelligence model of robot operation control.

[0007] Preferably, the data set includes feature values ​​and video data in the robot operation trajectory; the labeling personnel label the trajectory pairs by observing the video data and select the preferred trajectory.

[0008] Preferably, the pairing method of trajectory pairs includes pairing data sets with larger and smaller DTW distances, and pairing data sets with similar DTW distances; and measuring the work indicators of the labeling personnel to assign trajectory pairs with different attributes or different pairing methods for labeling, and the above work indicators include coverage, familiarity, similarity, deviation and fatigue.

[0009] Preferably, the video data also includes key video frames corresponding to the feature values, and the annotator selects the preferred trajectory by observing the key video frames and annotating them in combination with the video data.

[0010] Preferably, it also includes a method for evaluating the consistency of the annotation personnel's annotations on the corresponding principle attribute clustering, which is to perform statistics and analysis based on the annotation results of the annotation personnel, and judge the consistency of the annotation personnel's evaluation of the principle attribute clustering annotations. If it is judged that the annotation results deviate from the statistical results, the corresponding key frames in the feature are displayed to allow the annotation personnel to make further annotations.

[0011] A robot learning system based on a human preference labeling method for robot operation trajectories, comprising: The robot operation trajectory information collection module is used to collect the robot's actions, states, explicit ratings and implicit feedback in the operation trajectory; The eigenvalue calculation module performs normalization calculation on the eigenvalue of the corresponding feature; A feature algebra module for performing feature manipulation through feature algebra based on features annotated by human preferences based on robot operation trajectories; The diffusion model generation module controls the actions of the generated robot by combining the feature algebra module.

[0012] Preferably, the implicit feedback includes eye tracking and facial expression analysis of the annotator.

[0013] Compared with the prior art, the invention has the beneficial effects of providing a novel robot trajectory preference labeling tool for human evaluators. The tool includes a database clustering method based on attributes and principles, a back-end active query model, and a front-end hybrid active user interface.

[0014] The robot trajectory dataset is clustered by three main principles: safety, efficiency, and performance. For each principle, we collect quantifiable properties, concatenate them into a principal vector at each time step, and form a principal vector time series for the entire trajectory. Then, we cluster the trajectory dataset according to each principal vector sequence using the Dynamic Temporal Wrapping (DTW) distance.

[0015] When collecting preferences according to each principle, the backend model of our tool computes the weighted average of several metrics that we designed based on user needs during preference elicitation, such as the difficulty of comparing trajectory pairs, the disagreement between human raters, and the skewness of labeled and unlabeled trajectory pairs.

[0016] The front end presents trajectory-video pairs for human raters to compare and specify their preferences. The hybrid-active user interface consists of buttons showing keyframes of attributes for each trajectory video, attribute statistics and graphs that are adaptively displayed but also adaptable to the user, and a progress bar at the top of the user interface. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, further description is given below in conjunction with the accompanying drawings.

[0018] Figure 1 Describes the components and framework diagram of the present invention; Figure 2 It is a flow chart of active learning system; Figure 3 It is an active query process that prompts the annotator about the trajectory pairs. DETAILED DESCRIPTION

[0019] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0020] The inventors conducted a formative study to explore how general users compare trajectories and elicit their challenges and needs when using a basic trajectory preference annotation system. The study recruited 12 participants with an average age of 22 years (SD = 2.6). The subjects reported low to moderate familiarity with data labeling and robotic systems, with an average of 1.9 (SD = 0.9) on a 5-point Likert scale. A basic trajectory preference annotation system was developed. The system utilized a pick-and-place dataset from the Robomimic task, which contained six cans and four boxes. Robot trajectories consisted of state-action pairs, with states including the positions and velocities of joints and objects, simulated and recorded through the Robosuite framework. Trajectory videos were captured at 20 fps. K-means clustering was used to analyze states, with 9 representative samples generated for each cluster. The system randomly presented 36 pairs of trajectories for annotation during the study. Participants labeled trajectory pairs, which took 6-22 minutes (M = 15.4, SD = 4.0). This was followed by a retrospective think-aloud study and semi-structured interviews to explore challenges, comparison criteria, potential issues with larger datasets, and system requirements. The sessions were recorded, transcribed, and analyzed thematically. Analysis of the transcripts revealed three main themes: trajectory comparison principles, preference elicitation challenges, and labeling system requirements.

[0021] According to the interviews with the participants, there are three main trajectory comparison principles: safety, efficiency, and performance. The safety principle is determined by the attributes collision, distance, and contact force. The efficiency principle is determined by speed, trajectory length, time, and cost. The performance principle is determined by smoothness, stability, direction, and posture.

[0022] Interviews revealed three major challenges for trajectory analysis: poorly defined standards, neglect of key trajectory details, and difficulty focusing. Detailed issues include imprecise feature coverage, insufficient robot comprehension, limited video observability, mental fatigue, and insufficient feedback. These challenges highlight the inherent complexity of robotics and video comparison tasks.

[0023] To address these issues, the design framework proposed in the patent application outlines three core requirements: adaptive marker sorting, emphasis on important trajectory elements, and continuous attention tracking and feedback. The framework introduces five key features: feature-based trajectory clustering, selective keyframe extraction, dynamic pairing of trajectories, consistency verification, and flexible visualization of key data points. This comprehensive approach aims to improve standard establishment, detail identification, and continuous engagement throughout the trajectory comparison process.

[0024] like Figure 1As shown, the tool consists of four main components: a clustering method for organizing databases based on attributes and principles, an active learning model based on the database, a back-end active query model, and a front-end hybrid active user interface.

[0025] For basic robotic manipulation tasks, this patent defines mathematical formulas for three principles with properties summarized from formative research. This patent application collects the preferences of trajectory pairs for each principle. For each principle, this application's clustering method calculates the DTW distance between the main vectors of the trajectory pair. Then, the method clusters the trajectory dataset based on the DTW distance.

[0026] With a number of initially labeled data, we train an active learning model, actively add trajectories with uncertain labels to the label pool, and then update the model according to the labels updated by the front end through the query process. Figure 2 shown. Figure 2 A comprehensive active learning framework is outlined, which can be described by the following pseudocode: Algorithm PreferencePrediction Input: Model, UnlabeledData, HumanEvaluator Output: LabeledData 1. Initialize: LabeledData = {} 2. For each data point x = (t_i, t_j) in UnlabeledData do: a. Predict the preference probability P using the Model b. If uncertainty in P is high then i. Query HumanEvaluator for the preference label pref_ij ii. If L is provided then A. Add (x, pref_ij) to LabeledData B. Update Model with (x, pref_ij) 3. Return LabeledData We can choose different algorithms for the model for pairwise preference prediction. Some common ranking algorithms include RankNet, RankBoost, and RankSVM. The following pseudo code shows these three ranking algorithms: Algorithm RankNet Input: PairwisePreferenceData Output: TrainedModel 1. Initialize neural network Model with parameters θ 2. For each training epoch do: a. For each pair (x_i, x_j) in PairwisePreferenceData do: i. Compute scores s_i and s_j for x_i and x_j using the Model ii. Compute preference probability P_ij = sigmoid(s_i - s_j) iii. Compute cross-entropy loss: L = - [L_ij * log(P_ij) + (1 - L_ij) * log(1 - P_ij)] iv. Backpropagate the loss and update θ 3. Return TrainedModel Algorithm RankBoost Input: PairwisePreferenceData, WeakLearners Output: CombinedModel 1. Initialize weights w_i,j = 1 / |PairwisePreferenceData| 2. Initialize CombinedModel 3. For each iteration t do: a. 使用加权成对偏好数据训练弱学习器WeakLearner_t b. 对于每一对(x_i, x_j),执行以下操作: i. 使用弱学习器WeakLearner_t预测偏好分数 ii. 根据预测准确性更新权重w_i,j c. 将弱学习器WeakLearner_t添加到具有权重α_t的组合模型中 4. 返回组合模型 排序支持向量机算法 输入:成对偏好数据 输出:训练好的模型 1. 使用参数θ初始化支持向量机模型 2. 对于成对偏好数据中的每一对(x_i, x_j),执行以下操作: a. 计算差异向量x_ij = x_i - x_j b. 如果x_i比x_j更受偏好,则将y_ij赋值为1,否则y_ij赋值为 -1 3. 在(x_ij, y_ij)对上训练支持向量机模型 4. 返回训练好的模型 During the preference collection process, according to Figure 3, the backend model of our tool dynamically arranges the query flow based on metrics such as coverage, familiarity, similarity, divergence, and skewness. These metrics are dynamically converted into ranking-based scores and then averaged to determine the final ranking score. The higher the score, the higher the priority of presenting the track pair to the user. Initially, when the user has not yet viewed all clusters, the server selects track pairs from uncovered clusters and ranks them for presentation to the user based on the average of other metrics. This initial stage is designed to give the user an overview of features from different clusters and simplify the process of establishing criteria. Once the user has viewed all clusters, the server selects the track pair with the lowest label skewness score and ranks it for presentation based on other metrics. These metrics balance labeling difficulty, user familiarity, and divergence.

[0027] The present invention addresses the limitations of current robotic learning methods by introducing a novel system that effectively combines various techniques.

[0028] A module for collecting robot operation trajectory information, including explicit ratings and implicit feedback of eye tracking and facial expression analysis.

[0029] The eigenvalues ​​are calculated based on the human preference labeling method of the robot operation trajectory mentioned above.

[0030] Characteristic algebra is used to control the generation of diffusion models.

[0031] The system allows for more flexible control of robot learning, making it possible to incorporate features that are preferred by humans and to facilitate adjustments to learning based on specific features associated with the robot's trajectory.

[0032] Collection of robot operation trajectory information: The system collects robot demonstrations in the form of action-state pairs , , represents the robot's actions in tshike de, and Represents the state at time t. In addition to these basic elements, the system also collects: 1. Explicit rating for each operation trace ( ) 2. Explicit ratings and implicit feedback from eye tracking and facial expression analysis These data are used to provide a more fine-grained rating for each operation trace, denoted as ,and represents the weight of each feature i. Track structure: The initial trajectory structure is expressed as , and all The values ​​are normalized.

[0033] Eigenvalue calculation: For each feature i, the system calculates the feature value at time t and normalizes it, expressed as , using the human preference labeling method based on the robot operation trajectory. The extended trajectory structure is obtained feature: Applications of Characteristic Algebra: The system uses the calculated eigenvalues ​​as a subspace to control the generation of the diffusion model. In this method, the action vector is generated based on the state vector. By extending the conceptual algebra of score-based generative modeling, its mathematical framework is as follows Human preference characteristics and potential information — the action vector in the time window, — the state vector in the time window (including observations); As a latent variable, capturing all the information about the robot's actions related to the robot's state observed by the human, we get: (1) —Preference characteristic variables ( measurable), Related to Yes, the sample space .

[0034] A set of preference feature pairs { } satisfies X if ,in For example, the preference feature "safety" would be sufficient for the class of trajectories with a sequence of "no collision" states. Trajectories themselves induce distributions of many preference features (e.g. "efficiency"), but these other preference features are independent of the given feature "safety" = the class of trajectories with a sequence of "no collision" states. Then, (2) Distribution of human preference traits Following equation (2), we can look at each state vector Specify a distribution Latent preference characteristics This observation allows us to precisely determine the relationship between state and preference characteristics.

[0035] Feature distribution Represents the distribution of human characteristics. Each state vector Specify the feature distribution as That is, we move from viewing the state vector as expressing a specific preference for features (i.e., safety preference feature score = 1) to expressing a probability distribution over human preference features ( . is more general in terms of probability - features expressed deterministically can be represented as degenerate distributions. This additional generality is necessary: ​​for example, a certain class of trajectory states leads to a non-degenerate distribution over safety features.

[0036] The generative model of feature distribution state control introduces the state vector And produces random output Such a model is a mapping from the state vector Probability density space expression about As is sufficient to specify the output distribution.

[0037] —Density is defined as Assume that the model learns the true data distribution of ; (3) Feature Representation is a function that maps the distribution of features Corresponding expression , is a vector space. The representation of the state vector is an expression of the distribution of related concepts, The second reason is that it allows us to reason about representations that do not correspond to any state vector. Every state vector defines a feature distribution, but not every feature distribution can be defined by a state vector. This is important because we ultimately want to reason about the meaning of features of representation vectors created by algebraic manipulation of state representations. Such vectors need not correspond to any state vector.

[0038] Arithmetic Composition We now have the tools to define what it means for a representation to be arithmetically composable. We define composability for a pair of properties and In subsequent developments, our goal will be to manipulate Deviating from fixed .

[0039] A statement is arithmetically composable with respect to the features , If there is a vector space and So that for all characteristic distributions of the form , (4) Said as well as In general: We limit product distribution to meet specific and can freely operate with each other (typically and One or both of them are degenerate, when all mass is attributed to a single point mass). This definition requires that there exist fixed subspaces for each feature, in the sense that, for example, only changes in Only causes change.

[0040] Fractional Notation Fractional representation Feature distribution Defined as: The middle score represents is defined as Here, It is a function itself and , indicating space is a vector space of functions.

[0041] Causal separability fractions do not have an arithmetically composable structure for every pair of features. The point is that concepts are reflected in representations based on their effects on things. If they affect the way Fundamentally depends on some interaction between the two features, indicating that we cannot hope to disentangle them. Therefore, we must rule out this scenario.

[0042] We say Cause and effect are separable. If there is a unique - Measurable variables ,as well as ,as follows For some reversible and verifiable functions ,as well as if is relative to and Cause and effect are separable, then the middle score represents relative to and can be combined by arithmetic operations.

[0043] Characteristic Algebra To modify specific features We only want to modify the representation on the subspace correspond For example, consider changing the safety feature "number of collisions" to "0 collisions". Intuitively, we need an operation of the following form: Said is the projection onto the subspace corresponding to the “number of collisions” feature.

[0044] Feature manipulation via projection (6) Here —— The subspace of (8) as well as --exist The projection of Project to If we can calculate , we can just edit The expression of each point. That is, we transform the score function of each point: (9) We then draw samples from the stochastic differential equation defined by .

[0045] let yes Conceptually any fixed distribution, yes Any reference distribution on . Then, assuming The causal separability of (10) We use equation (10) to define The goal is to find the basis of the subspace using the state vector This derives a distribution of the form For example, to identify the “number of conflicts” feature, we use the conflict-free state vector , collides with 1 , ... until Collision, based on this idea (11) have the same marginal distribution We then define the estimated subspace using the state vector as (12) In other words, we designed State Vector , so each in have different distributions, but in Then, the estimation space is calculated according to equation (12).

[0046] Characteristic Algebra In summary, this algebraic approach to concepts is: 1. Find the state vector , so that each one gets Different distribution, but The same distribution. That is, For each 2. Constructing an estimated representation space Following equation (12), we define the project As a projection of this space.

[0047] 3. Sample of discretized SDE defined by manipulation score representation (13) The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for labeling human preference of robot operation trajectory, characterized by: This includes dividing the collected robot operation trajectory data set into data according to the following principle attributes: Safety, including collision, distance, and contact force features: efficiency, including speed, trajectory length, time and cost, and; performance, including smoothness, stability, orientation, and posture; For each feature of the principle attribute, a corresponding principal vector related to the principle attribute is established, and the DTW distance between the eigenvalue of the feature corresponding to the attribute in the data set and the corresponding principal vector is calculated, and the data set is clustered accordingly; the clustered data sets are paired to form trajectory pairs, and the trajectory pairs are sent to the annotator for judgment and selection of the optimal trajectory; the optimal trajectory is formed into a training data set for the artificial intelligence model of robot operation control.

2. The method for labeling the human preference of the robot operation trajectory according to claim 1, characterized in that: The data set includes the feature values ​​and video data in the robot operation trajectory; the labeling personnel select the preferred trajectory by labeling the trajectory pairs by observing the video data.

3. The method for labeling the human preference of the robot operation trajectory according to claim 1, characterized in that: The pairing method of trajectory pairs includes pairing data sets with larger and smaller DTW distances, and pairing data sets with similar DTW distances; and measuring the work indicators of the labelers to assign trajectory pairs with different attributes or different pairing methods for labeling. The above work indicators include coverage, familiarity, similarity, deviation and fatigue.

4. The method for labeling the human preference of the robot operation trajectory according to claim 2, characterized in that: The video data also includes key video frames corresponding to the feature values. When making judgments, the annotators select the preferred trajectory by observing the key video frames and annotating them in combination with the video data.

5. The method for labeling the human preference of the robot operation trajectory according to claim 4, characterized in that: It also includes a method for evaluating the consistency of the annotation personnel's annotations on the corresponding principle attribute clustering, which is to perform statistics and analysis based on the annotation results of the annotation personnel, and judge the consistency of the annotation personnel's evaluation of the principle attribute clustering annotation. If the annotation result is judged to deviate from the statistical result, the corresponding key frame in the feature is displayed to allow the annotation personnel to make further annotations.

6. A robot learning system based on the method for labeling the human preference of the robot operation trajectory according to any one of claims 1 to 2, characterized in that: include The robot operation trajectory information collection module is used to collect the robot's actions, states, explicit ratings and implicit feedback in the operation trajectory; The eigenvalue calculation module performs normalization calculation on the eigenvalue of the corresponding feature; A feature algebra module for performing feature manipulation through feature algebra based on features annotated by human preferences based on robot operation trajectories; The diffusion model generation module controls the actions of the generated robot by combining the feature algebra module.

7. A robot learning system according to claim 5, characterized in that: The implicit feedback includes eye tracking and facial expression analysis of the annotator.