Zero sample learning-based motion data labeling method, equipment and medium

By using a zero-shot learning-based method, positive and negative sample pairs are constructed using label features and motion features. A contrastive learning neural network model is then used to annotate human 3D motion data, solving the problem of low efficiency in existing technologies and achieving efficient data annotation.

CN120873667APending Publication Date: 2025-10-31XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510865075.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods for annotating human 3D motion data suffer from low efficiency and high cost due to manual annotation, while neural network-based methods require readjusting the training model to adapt to new motion types, resulting in low efficiency.

Method used

We employ a zero-shot learning approach, which constructs positive and negative sample pairs by collecting label features and motion features from labeled data. We then train a contrastive learning neural network model and use a pre-trained language model and a spatiotemporal graph convolutional network for feature extraction and prediction, thereby achieving the labeling of unlabeled data.

Benefits of technology

It improves the model's generalization ability, reduces the cost of labeling new categories of data, avoids the need to re-collect and label large amounts of data, and improves labeling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873667A_ABST
    Figure CN120873667A_ABST
Patent Text Reader

Abstract

The invention relates to a motion data labeling method and device based on zero sample learning and a medium, and the method comprises the steps: collecting human motion data with a label, and extracting a label feature and a motion feature; constructing a training set based on the positive sample pair and the negative sample pair corresponding to each label; constructing a neural network model based on comparative learning, setting the input of the model as a reference feature set composed of all label features and the motion features of the motion data, and setting the output of the model as the label category to which the motion data belongs; training the model through the training set; when the motion data needs to be labeled, extracting motion features of to-be-labeled motion data, then jointly inputting the motion features into the trained model in combination with the reference feature set, and taking the output of the model as a label of the to-be-labeled motion data; and when a newly added tag type is received, extracting tag features of the newly added tag type and adding the tag features to the reference feature set. According to the method, the generalization ability of the model is improved, and the data annotation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a method, device and medium for motion data annotation based on zero-shot learning. Background Technology

[0002] Human motion data annotation refers to the precise annotation of human motion posture in three-dimensional space at each moment. High-quality annotated data is often crucial for data analysis, model training and application development, and provides important technical support for fields such as action recognition, human-computer interaction, virtual reality and augmented reality.

[0003] In existing technologies, human 3D motion data can be labeled using manual annotation or neural network-based automatic annotation methods. Manual annotation requires a significant amount of manual work, which, while ensuring accuracy, is inefficient and incurs substantial manpower and time costs. Neural network-based methods are currently the mainstream automatic annotation approach. These methods train suitable neural network models using a large amount of labeled data, thus achieving automatic annotation. A common network structure is the spatiotemporal graph convolutional neural network (SPCNN), which can handle data with temporal and spatial dependencies and is suitable for labeling 3D motion data. It captures spatial relationships in motion data through a graph structure and uses temporal information to predict motion trajectories. However, in practical applications, many types of motion exist that are not present in the training data. This neural network-based approach struggles to handle data outside the training data. When new motion types require annotation, the data needs to be re-labeled, and the model updated to adapt to the new data.

[0004] In summary, the current technologies used for human three-dimensional motion data annotation have the following drawbacks: (1) Although manual annotation methods are accurate, they are inefficient and costly; (2) Neural network-based annotation methods can significantly improve annotation efficiency, but require readjustment of the training model when faced with new types of motion. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a motion data annotation method, device, and medium based on zero-shot learning.

[0006] The specific plan is as follows:

[0007] A motion data annotation method based on zero-shot learning includes the following steps:

[0008] S1: Collect labeled human motion data and extract the label features of each labeled data and the motion features of each motion data.

[0009] S2: Based on label features and motion features, construct positive and negative sample pairs for each label, and then construct a training set based on the positive and negative sample pairs for each label.

[0010] S3: Construct a neural network model based on contrastive learning. The input of the model is a reference feature set consisting of all labeled features and the motion features of the motion data. The output of the model is the label category to which the motion data belongs. Train the model using the training set.

[0011] S4: When it is necessary to label motion data, extract the motion features of the motion data to be labeled, and input them into the trained model in combination with the reference feature set. Use the output of the model as the label of the motion data to be labeled.

[0012] S5: When a new label type is received, extract the label features of the new label data, add it to the reference feature set, and return to S4.

[0013] Furthermore, the process of extracting label features is as follows: after obtaining the motion posture descriptions of each human body part corresponding to the label data based on the large language model, the motion posture descriptions are encoded by the pre-trained language model BERT.

[0014] Furthermore, before extracting motion features from the motion data, the motion data is preprocessed; the preprocessing includes outlier removal, missing value imputation, and data alignment to a human skeletal model.

[0015] Furthermore, motion features are extracted using a spatiotemporal graph convolutional network.

[0016] Furthermore, the motion feature extraction process employs a contrastive learning approach. The method for constructing positive and negative sample pairs used in contrastive learning is as follows: Samples s = {s1,...,s...} are constructed based on motion data of a fixed number of consecutive frames. t}, s t This represents the motion data of frame t, and constructs its corresponding positive sample s based on sample s. + ={s 1+Δt ,...,s t+Δt} or s + ={s 1-Δt ,...,s t-Δt} and negative samples s - ={s t ,...,s1}, where Δt is a frame number much smaller than t, then the positive and negative sample pairs are represented as <s,s + ,s - >

[0017] Furthermore, the neural network model employs a fully connected network.

[0018] Furthermore, the loss function L of the model is:

[0019]

[0020] Where q′ represents the model's subsequent motion features, t i Let represent the i-th label feature in the reference feature set, K represent the total number of labels, t represent the true label of the motion data corresponding to the motion feature input to the model, τ represent the temperature hyperparameter, and sim represent the similarity calculation.

[0021] A motion data annotation terminal device based on zero-shot learning includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the embodiments of the present invention.

[0022] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above in the embodiments of the present invention.

[0023] The present invention adopts the above technical solution, which utilizes existing labeled data and semantic information to identify unlabeled new categories, greatly improving the generalization ability of the model, thereby avoiding the need to collect and label a large amount of data for each new category and greatly reducing the cost of data labeling. Attached Figure Description

[0024] Figure 1 The diagram shown is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation

[0025] To further illustrate the various embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, primarily used to illustrate the embodiments, and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementations and the advantages of the present invention.

[0026] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.

[0027] Example 1:

[0028] This invention provides a motion data annotation method based on zero-shot learning, such as... Figure 1 As shown, the method includes the following steps:

[0029] S1: Collect labeled human motion data and extract the label features of each labeled data and the motion features of each motion data.

[0030] In this embodiment, the human motion data collected is frame-by-frame motion data collected over a continuous time period. This data corresponds to different motion labels (such as running, walking, sitting, etc.) at different time periods. The collected human motion data has two modalities: one is text-formatted label data, and the other is three-dimensional motion data (usually frame-by-frame collected joint motion data, represented as spatial three-axis coordinate data). The two types of data need to be feature extracted separately.

[0031] To extract the label features from the label data, this embodiment first performs data augmentation on the label data. This involves leveraging the powerful generative capabilities of a large language model by providing a prompt: "Please describe the movement posture of the human body [part1,...,partn] when performing [action]." This yields a detailed description of the movement posture of various parts of the body (such as hands, legs, hips, and head) during the movement. The "action" corresponds to the label text, and "part1,...,partn" represents the movement posture of the human body. n Corresponding to n parts of the human body.

[0032] After obtaining the motion posture descriptions of each human body part corresponding to the labeled data, it is necessary to encode each motion posture description. In this embodiment, the pre-trained language model BERT is preferably used for encoding, and the encoding results are used as the label features of the labeled data.

[0033] This embodiment includes preprocessing the motion data before extracting its motion features. Preprocessing includes outlier removal, missing value imputation, and data alignment to a human skeletal model. Outlier removal can be performed using statistical methods (such as Z-scores) or methods based on physical constraints. Considering the high continuity of the motion data, interpolation or mean-based methods are preferred for imputing missing values. The human skeletal model can use existing models such as SMPL and NTU-RGB+D, which have fixed joints and can achieve a fixed topological structure.

[0034] In this embodiment, considering the good spatial and temporal properties of motion data, a spatiotemporal graph convolutional network trained with a halo is preferred for motion feature extraction. It is worth noting that due to the presence of convolutional blocks, the spatiotemporal graph convolutional network encodes a frame along with its surrounding frames (e.g., a feature is actually the feature encoding of multiple surrounding motion frames), which to some extent strengthens the feature representation. After discarding the full-image average pooling layer and fully connected layers, the encoded data format is (N, C, T, V), where N represents the batch size, C represents the number of features per node (e.g., the coordinate information of joints, usually three, i.e., X, Y, Z coordinates), T represents the length of the time series (i.e., the number of frames), and V represents the number of joints in the human skeleton, depending on the human model used.

[0035] Furthermore, in this embodiment, the enhancement of motion features during the extraction process employs a contrastive learning approach. The method for constructing positive and negative sample pairs used in contrastive learning is as follows: Samples s = {s1,...,s...} are constructed based on motion data of a fixed number of consecutive frames. t}, s t Represents the motion data of frame t, and constructs its corresponding positive sample s based on sample s. + ={s 1+Δt ,...,s t+Δt} or s + ={s 1-Δt ,...,s t-Δt} and negative samples s - ={s t ,...,s1}, where Δt is a frame number much smaller than t, then the positive and negative sample pairs are represented as <s,s + ,s - >

[0036] The motion feature encoding uses a pre-trained spatiotemporal graph convolutional network feature extraction model, which is fine-tuned using the positive and negative sample pairs constructed above to enhance the motion features encoded by the model. A contrastive loss is used, and the specific loss function is as follows:

[0037]

[0038] Among them, q, q + q - These represent the motion features of the samples, positive samples, and negative samples output by the feature extraction model, respectively.

[0039] Through the comparative learning method described above, the features output by the feature extraction model can be enhanced, that is, the boundary between the end and the beginning of an action can be more clearly defined.

[0040] S2: Based on label features and motion features, construct positive and negative sample pairs for each label, and then construct a training set based on the positive and negative sample pairs for each label.

[0041] A positive sample pair contains the motion features of a sample and its true label; a negative sample pair contains the motion features of a sample and other labels besides its true label.

[0042] In this embodiment, each sample is composed of motion data of a fixed number of consecutive frames. These consecutive frames of motion data all correspond to the same label and form blocks.

[0043] S3: Construct a label prediction model based on contrastive learning. The input of the model is a reference feature set consisting of all label features and the motion features of the motion data. The output of the model is the label category to which the motion data belongs. The model is trained using a training set.

[0044] The label prediction model includes a feature alignment module, a similarity pair calculation module, and a label prediction module. In this embodiment, the feature alignment module uses a fully connected network. It outputs aligned motion and label features based on the input motion and label features. The similarity calculation module then calculates the similarity between the aligned motion and label features. Finally, the label prediction module predicts the label corresponding to the motion data based on the similarity score. Since there are multiple labels, this embodiment is a multi-classification model. The model outputs the probability that the motion data belongs to each label, and the label with the highest probability is used as the predicted label for the motion data.

[0045] Since the model training process requires calculating the contrast loss between the two features output by the alignment network, this embodiment sets the loss function L used in the training process as follows:

[0046]

[0047] Where q′ represents the model's subsequent motion features, t i Let represent the i-th label feature in the reference feature set, K represent the total number of labels, t represent the true label of the motion data corresponding to the motion feature input to the neural network model, τ represent the temperature hyperparameter, and sim represent the similarity calculation. In this embodiment, cosine similarity is used.

[0048] This embodiment continuously optimizes the model's network parameters through comparative learning, maximizing the correlation between features of two modalities within the same class. This ensures that features of the two modalities are aligned to the same feature space, resolving the potential issue of dimensionality inconsistencies between features. By continuously optimizing the model's network parameters through comparative learning, the correlation between features of two modalities within the same class is maximized.

[0049] S4: When it is necessary to label motion data, extract the motion features of the motion data to be labeled, and input them into the trained model in combination with the reference feature set. The output of the model is used as the label of the motion data to be labeled.

[0050] S5: When a new label type is received, extract the label features of the new label data, add it to the reference feature set, and return to S4.

[0051] This embodiment uses zero-shot learning to identify unlabeled new categories by utilizing existing labeled data and semantic information, which greatly improves the model's generalization ability and avoids the need to collect and label a large amount of data for each new category, thus significantly reducing data labeling costs.

[0052] Example 2:

[0053] The present invention also provides a motion data annotation terminal device based on zero-shot learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method embodiment described above in Embodiment 1 of the present invention.

[0054] Furthermore, as an executable solution, the zero-shot learning-based motion data annotation terminal device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The zero-shot learning-based motion data annotation terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above-described structure of the zero-shot learning-based motion data annotation terminal device is merely an example and does not constitute a limitation on the zero-shot learning-based motion data annotation terminal device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the zero-shot learning-based motion data annotation terminal device may also include input / output devices, network access devices, buses, etc., and this embodiment of the invention does not limit this.

[0055] Furthermore, as an executable solution, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices. The general-purpose processor can be a microprocessor or any conventional processor. This processor is the control center of the zero-shot learning-based motion data annotation terminal device, connecting all parts of the device via various interfaces and lines.

[0056] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the zero-shot learning-based motion data annotation terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0057] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described in the embodiments of the present invention.

[0058] If the modules / units integrated in the zero-shot learning-based motion data annotation terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc.

[0059] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A motion data annotation method based on zero-shot learning, characterized in that, Includes the following steps: S1: Collect labeled human motion data and extract the label features of each labeled data and the motion features of each motion data. S2: Based on label features and motion features, construct positive and negative sample pairs for each label, and then construct a training set based on the positive and negative sample pairs for each label; S3: Construct a label prediction model based on contrastive learning. The input of the model is a reference feature set consisting of all label features and the motion features of the motion data. The output of the model is the label category to which the motion data belongs. The model is trained using a training set. S4: When it is necessary to label motion data, extract the motion features of the motion data to be labeled, and input them into the trained model in combination with the reference feature set. Use the output of the model as the label of the motion data to be labeled. S5: When a new label type is received, extract the label features of the new label data, add it to the reference feature set, and return to S4.

2. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: The process of extracting label features is as follows: after obtaining the motion posture descriptions of each human body part corresponding to the label data based on the large language model, the motion posture descriptions are encoded by the pre-trained language model BERT.

3. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: Before extracting motion features from motion data, the process also includes preprocessing the motion data; preprocessing includes outlier removal, missing value imputation, and data alignment to a human skeletal model.

4. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: Motion features are extracted using a spatiotemporal graph convolutional network.

5. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: The motion feature extraction process employs a contrastive learning approach. The method for constructing positive and negative sample pairs used in contrastive learning is as follows: Samples s = {s1,...,s...} are constructed based on motion data of a fixed number of consecutive frames. t }, s t Represents the motion data of frame t, and constructs its corresponding positive sample s based on sample s. + ={s 1+Δt ,...,s t+Δt } or s + ={s 1-Δt ,...,s t-Δt } and negative samples s - ={s t ,...,s1}, where Δt is a frame number much smaller than t, then the positive and negative sample pairs are represented as <s,s + ,s - > 6. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: It includes a feature alignment module, a similarity pair calculation module, and a label prediction module, wherein the feature alignment module adopts a fully connected network.

7. The motion data annotation method based on zero-shot learning according to claim 1, characterized in that: The loss function L of the model is: Where q′ represents the model's subsequent motion features, t i Let represent the i-th label feature in the reference feature set, K represent the total number of labels, t represent the true label of the motion data corresponding to the motion feature input to the model, τ represent the temperature hyperparameter, and sim represent the similarity calculation.

8. A motion data annotation terminal device based on zero-shot learning, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.