A method for selecting training data and related devices

By automatically selecting training data from the data set, the problem that training data selection in the existing technology depends on expert experience is solved, and the automatic selection of high-quality training data is realized, and the model performance and robustness are improved.

CN115130598BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210789247.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-05-30
Estimated Expiration
2042-07-06

AI Technical Summary

Technical Problem

In the prior art, the training data selection of machine learning models is heavily dependent on expert experience, resulting in poor model performance, poor robustness and generalization.

Method used

A method of automatically selecting training data from the data set is proposed. By performing feature encoding processing on candidate data in the target data set, its representative parameters are determined, and training data is selected based on these parameters until the data selection end condition is met.

Benefits of technology

Automatic selection of high-quality training data is realized, which improves the performance of the model, enhances the robustness and generalization of the model, and reduces the cost and efficiency of training data selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130598B_ABST
    Figure CN115130598B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method for selecting training data and related devices in the field of artificial intelligence. The method includes: performing feature encoding processing on each candidate data in the target data set to obtain the encoded features of each candidate data; selecting candidate data from the target data set as the initial training data, and migrating the training data from the target data set to the training sample set; repeatedly performing the training data selection operation until the data selection end condition is met; the training data selection operation includes: determining the representative parameters corresponding to each candidate data in the target data set according to the encoded features of each candidate data in the target data set and the encoded features of each training data in the training sample set; selecting candidate data from the target data set as the training data according to the representative parameters corresponding to each candidate data, and migrating the training data from the target data set to the training sample set. The model trained based on the training data selected by this method has better performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for selecting training data and related devices. Background Art

[0002] Machine learning technology is used to study how to enable a computer to simulate or implement human learning behaviors, so as to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. It is the core of artificial intelligence technology and the fundamental way to make a computer achieve intelligence.

[0003] The implementation of machine learning technology usually requires a large amount of labeled data, and the acquisition cost of high-quality labeled data is high and the acquisition difficulty is relatively high. At present, a common way to obtain labeled data is that experts select several data from existing relevant data sets as training data for machine learning, and then experts label the selected training data.

[0004] However, the selection of training data from relevant data sets by experts will seriously depend on the experience and judgment of experts, and the effects of models trained based on different training data often vary. In many cases, the performance of the model trained based on the training data selected by experts is not ideal, and the robustness and generalization of the model are poor. Summary of the Invention

[0005] Embodiments of this application provide a method for selecting training data and related devices, which can automatically select training data for training a model from a data set and ensure that the model trained based on the selected training data has better performance.

[0006] In view of this, in the first aspect of this application, a method for selecting training data is provided, and the method includes:

[0007] Performing feature encoding processing on each candidate data in the target data set to obtain the encoded features of each candidate data in the target data set;

[0008] Selecting at least one candidate data from the target data set as the initial training data, and migrating the training data from the target data set to the training sample set; the training data included in the training sample set is the data to be labeled for training the target model;

[0009] Repeatedly performing the training data selection operation until the data selection end condition is met;

[0010] Among them, the training data selection operation includes: determining, according to the encoding features of each candidate data in the target data set and the encoding features of each training data in the training sample set, the representative parameter corresponding to each candidate data in the target data set, where the representative parameter is used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoding feature space; and selecting at least one candidate data from the target data set as training data according to the representative parameter corresponding to each candidate data in the target data set, and migrating the training data from the target data set to the training sample set.

[0011] The second aspect of the present application provides a training data selection device, and the device includes:

[0012] A feature encoding module, configured to perform feature encoding processing on each candidate data in the target data set to obtain the encoding feature of each candidate data in the target data set;

[0013] An initial data selection module, configured to select at least one candidate data from the target data set as initial training data, and migrate the training data from the target data set to the training sample set; the training data included in the training sample set is unlabeled data for training the target model;

[0014] A data selection module, configured to repeatedly execute the training data selection operation until the data selection end condition is met; among them, the training data selection operation includes: determining, according to the encoding features of each candidate data in the target data set and the encoding features of each training data in the training sample set, the representative parameter corresponding to each candidate data in the target data set, where the representative parameter is used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoding feature space; and selecting at least one candidate data from the target data set as training data according to the representative parameter corresponding to each candidate data in the target data set, and migrating the training data from the target data set to the training sample set.

[0015] The third aspect of the present application provides a computer device, and the device includes a processor and a memory:

[0016] The memory is used to store a computer program;

[0017] The processor is configured to execute the steps of the training data selection method as described in the first aspect above according to the computer program.

[0018] A fourth aspect of the present application provides a computer-readable storage medium for storing a computer program for executing the steps of the training data selection method described in the first aspect above.

[0019] A fifth aspect of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the training data selection method described in the first aspect above.

[0020] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0021] The embodiment of the present application provides a training data selection method, which innovatively proposes a mechanism for automatically selecting training data from a data set. Specifically, in this method, each candidate data in the target data set is first subjected to feature encoding processing to obtain the encoded features of each candidate data; then, at least one candidate data is selected from the target data set as the initial training data, and the training data is migrated from the target data set to the training data set; furthermore, the training data selection operation is repeatedly executed until the data selection end condition is met; the training data selection operation here includes: determining the respective representative parameters (used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoded feature space) of each candidate data according to the encoded features of each candidate data in the target data set and the encoded features of each training data in the training sample set, and then selecting at least one candidate data from the target data set as the training data according to the respective representative parameters of each candidate data, and migrating the training data from the target data set to the training sample set. When selecting training data from the target data set by the above method, the training data is selected from the target data set according to the distance between the candidate data in the target data set and each training data already selected into the training sample set in the encoded feature space. The training data selected in this way is usually the candidate data with more representativeness in the encoded feature space; the so-called representativeness can be understood as that the selected training data can represent the candidate data group in the target data set with similar features to it, and the candidate data groups represented by each selected training data are different, that is, the selected training data is evenly and dispersedly distributed in the encoded feature space. The target model trained based on such training data usually has better model performance, better robustness and generalization ability; and compared with the solution of manually selecting training data by experts, the solution of automatically selecting training data in the embodiment of the present application has higher training data selection efficiency and lower implementation cost. Description of the Drawings

[0022] Figure 1 Schematic diagram of the application scenario of the training data selection method provided by the embodiment of the present application;

[0023] Figure 2 Schematic diagram of the flow of the training data selection method provided by the embodiment of the present application;

[0024] Figure 3 Schematic diagram of the principle of selecting training data provided by the embodiment of the present application;

[0025] Figure 4 Schematic diagram of the working principle of the discriminator network provided by the embodiment of the present application;

[0026] Figure 5 Schematic diagram of the data annotation interface for the annotation object provided by the embodiment of the present application;

[0027] Figure 6 Schematic diagram of the flow of the training method of the target feature encoder provided by the embodiment of the present application;

[0028] Figure 7 Schematic diagram of the implementation principle of the double-layer contrast learning mechanism provided by the embodiment of the present application;

[0029] Figure 8 Schematic diagram of the implementation process of the training data selection method provided by the embodiment of the present application;

[0030] Figure 9 Schematic diagram of the experimental comparison results provided by the embodiment of the present application;

[0031] Figure 10 Schematic diagram of the structure of the training data selection device provided by the embodiment of the present application;

[0032] Figure 11 Schematic diagram of the structure of the terminal device provided by the embodiment of the present application;

[0033] Figure 12 Schematic diagram of the structure of the server provided by the embodiment of the present application. Detailed implementation manners

[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0035] In the description and claims of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0036] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.

[0037] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0038] The solution provided by the embodiments of this application relates to the machine learning technology in artificial intelligence technology, and is specifically illustrated through the following embodiments:

[0039] In the related art, training data is usually selected from existing relevant datasets by experts based on their own experience and judgment. This way of selecting training data highly depends on the experience and judgment of experts. If the selected training data is not ideal (for example, the features of the selected training data are not rich enough and too concentrated), it will seriously affect the performance of the model trained based on these training data. For example, it may cause the trained model to only be able to accurately process specific types of data and cannot accurately process other types of data in a generalized manner; in addition, this way of selecting training data is inefficient and costly.

[0040] To solve the above technical problems, the embodiments of the present application provide a method for selecting training data. This method proposes a mechanism for automatically selecting training data from a dataset and can ensure that the selected training data is representative. Based on the selected training data, a model with better performance can be trained.

[0041] Specifically, in this method for selecting training data, first, feature encoding processing is performed on each candidate data in the target dataset to obtain the encoded features of each candidate data in the target dataset. Then, at least one candidate data is selected from the target dataset as the initial training data, and this training data is migrated from the target dataset to the training sample set. The training data included in the training sample set here is the data to be labeled for training the target model. Furthermore, the training data selection operation is repeatedly executed until the data selection end condition is met; the training data selection operation here includes: determining the respective representative parameters of each candidate data in the target dataset according to the encoded features of each candidate data in the target dataset and the encoded features of each training data in the training sample set. The representative parameter here is used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoded feature space; and according to the respective representative parameters of each candidate data in the target dataset, at least one candidate data is selected from the target dataset as the training data, and this training data is migrated from the target dataset to the training sample set.

[0042] When selecting training data from the target dataset using the above training data selection method, the training data is selected from the target dataset based on the distances between the candidate data in the target dataset and each piece of training data already selected into the training sample set in the encoded feature space. The training data selected in this way is usually the candidate data that is more representative in the encoded feature space. The so-called representativeness can be understood as the selected training data can represent the candidate data group in the target dataset with similar features, and the candidate data groups represented by each selected piece of training data are different, that is, the selected pieces of training data are evenly and dispersedly distributed in the encoded feature space. The target model trained based on such training data usually has better model performance. The target model can process various types of data more accurately and has better robustness and generalization ability. Moreover, compared with the solution of manually selecting training data by experts, the automatic training data selection solution provided in the embodiments of the present application has higher training data selection efficiency and lower implementation cost.

[0043] It should be understood that the training data selection method provided in the embodiments of the present application can be executed by a computer device with data processing capabilities, and the computer device can be a terminal device or a server. Among them, the terminal device includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server. In addition, the relevant data involved in the embodiments of the present application can be stored in a blockchain network.

[0044] To facilitate the understanding of the training data selection method provided in the embodiments of the present application, the application scenario of the training data selection method will be exemplarily introduced below with the execution subject of the training data selection method being a server.

[0045] See Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the training data selection method provided in the embodiments of the present application. As Figure 1 shown, this application scenario includes a server 110, a database 120, and a terminal device 130. The server 110 can access the database 120 through a network, or the database 120 can also be integrated in the server 110. The server 110 communicates with the terminal device 130 through a network. Among them, the server 110 is used to execute the training data selection method provided in the embodiments of the present application to select the training data for training the model from the target dataset. The database 120 is used to store the target dataset. The terminal device 130 is used to obtain the training data selected by the server 110 and label the training data.

[0046] In practical applications, the server 110 can retrieve a target data set that matches the model training requirements from the database 120 according to the model training requirements. Exemplarily, if the server 110 needs to train a motion recognition model for recognizing action types based on motion sequence data, the server 110 can retrieve a target data set including a large amount of motion sequence data from the database 120; it should be understood that the data included in the target data set are all unlabeled data.

[0047] After the server 110 obtains the target data set, it can use the feature encoder to perform feature encoding processing on each candidate data in the target data set respectively, so as to obtain the encoded features of each candidate data. Then, the server 110 can randomly select at least one candidate data from the target data set as training data, and migrate the selected training data from the target data set to the training sample set, that is, add the training data to the training sample set, and at the same time delete the training data from the target data set.

[0048] Furthermore, the server 110 can perform a training data selection operation until the data selection end condition is met. The training data selection operation here includes: determining the respective representative parameters corresponding to each candidate data in the target data set according to the encoded features of each candidate data in the target data set and the encoded features of each training data in the training sample set, and the representative parameter is used to characterize the distance between the corresponding candidate data and each training data in the training sample set in the encoded feature space; and selecting at least one candidate data from the target data set as training data according to the respective representative parameters corresponding to each candidate data in the target data set, and migrating the training data from the target data set to the training sample set. It should be understood that each time a training data selection operation is completed, the target data set and the training sample set will be updated accordingly. The candidate data selected as training data will be missing in the target data set, and the training data selected from the target data set will be added to the training sample set; correspondingly, when performing the next training data selection operation, it will be based on the target data set and the training sample set updated after this training data selection operation. The data selection end condition here can be set according to actual needs. For example, it can be that the training data in the training sample set reaches a preset data volume.

[0049] After the server 110 determines that the data selection end condition is met, it can send the training data in the training sample set to the terminal device 130 to display the selected training data through the terminal device 130. Then, the terminal device 130 can receive the annotation result configured by its user for the displayed training data, and return the annotation result to the server 110, so that the server 110 can store the annotation result corresponding to the training data accordingly for subsequent training of the target model based on the training data and its corresponding annotation result.

[0050] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the training data selection method provided in the embodiments of the present application can also be applied to other scenarios, such as various scenarios like games, cloud technologies, intelligent transportation, assisted driving, etc. No limitations are imposed on the application scenarios applicable to the training data selection method provided in the embodiments of the present application herein.

[0051] The training data selection method provided in the present application will be introduced in detail below through method embodiments.

[0052] See Figure 2 , Figure 2 which is a schematic flowchart of the training data selection method provided in the embodiments of the present application. For ease of description, the following embodiments will still be described by taking the execution subject of this training data selection method as a server as an example. As Figure 2 shown, the training data selection method includes the following steps:

[0053] Step 201: Perform feature encoding processing on each candidate data in the target data set to obtain the encoded features of each candidate data in the target data set.

[0054] In the embodiments of the present application, when the server needs to select training data for training a model from the target data set to obtain the annotation results corresponding to the training data and use the training data and its corresponding annotation results to train the model, the server needs to first perform feature encoding processing on each candidate data in the target data set to obtain the encoded features of each candidate data in the target data set. The feature encoding processing here is a processing method that converts candidate data into digital data that can be recognized by a machine. This digital data is the encoded feature of the candidate data, which can to a certain extent reflect the characteristics of the candidate data.

[0055] It should be noted that the above target data set is a data set including a large amount of unlabeled data, and the unlabeled data included in the target data set matches the model training requirements; in the embodiments of the present application, the unlabeled data in the target data set is regarded as candidate data.

[0056] Exemplarily, assume that the target model to be trained is a motion recognition model for identifying action types based on motion sequence data. Then the target data set used should include a large amount of unlabeled motion sequence data; the motion sequence data here includes a series of sequentially arranged action data for characterizing action postures. This motion sequence data can be generated, for example, by capturing the actions of a specific object in a feature scene over a period of time, or can also be artificially simulated and created. The target data set in the above scenario can be represented as D=(x 1 ,x 2 ,……,xN ), where x 1 , x 2 , ……, x N is the motion sequence data in the target dataset. The motion sequence data x i can be represented, for example, as x i = (s i,1 , s i,2 , ……, s i,T ), where s i,1 , s i,2 , ……, s i,T are the action data in the motion sequence data x i . s i,j ∈ R J×3 (j = 1, 2, ……, T) represents the three-dimensional coordinates of J body joints in the action posture represented by the j-th action data. The embodiments of the present application aim to select a part of the motion sequence data from the target dataset D to form a training sample set D train .

[0057] It should be understood that in practical applications, the types of candidate data included in the target dataset change with the change of model training requirements, that is, the candidate data included in the target dataset can also be other types of data. The present application does not make any limitation on the types of candidate data included in the target dataset.

[0058] In a possible implementation manner, the server can perform feature encoding processing on each candidate data in the target dataset through a target feature encoder, so as to obtain the encoded features of each candidate data in the target dataset. The target feature encoder is trained by the unsupervised training method provided by the embodiments of the present application. The training method of the target feature encoder will be introduced in detail below.

[0059] Step 202: Select at least one candidate data from the target dataset as the initial training data, and migrate the training data from the target dataset to the training sample set; the training data included in the training sample set is the data to be labeled for training the target model.

[0060] In the embodiments of the present application, when the server selects the training data for training the model from the target dataset, it can first randomly select at least one candidate data from the target dataset as the initial training data, and then migrate the selected training data from the target dataset to the training sample set, that is, add the selected training data to the training sample set and delete the training data in the target dataset at the same time.

[0061] It should be noted that the above training sample set is a data set used to carry the training data selected from the target data set, and the training data included therein are all unlabeled data for training the target model; after the selection of the training data is completed, each training data in the training sample set can be labeled, and then the target model can be trained based on each training data and its corresponding labeling result.

[0062] Exemplarily, refer to Figure 3 , Figure 3 which is a schematic diagram of the principle of selecting training data provided by an embodiment of the present application. Figure 3 In (a), (b), (c), ……, (n), the inside of each corresponding circle represents an encoded feature space, which is determined according to the encoded features of each candidate data in the target data set, and each circle in the encoded feature space corresponds to an encoded feature. As Figure 3 shown in (a), the server can first randomly select an encoded feature in the encoded feature space, and the candidate data corresponding to the encoded feature is selected as the initial training data. The encoded feature selected as the training data can be correspondingly represented as a black solid circle, that is Figure 3 the solid circle 1 in (a).

[0063] It should be understood that in practical applications, when the server selects the initial training data from the target data set, it can select one training data or multiple training data with encoded features far apart in the encoded feature space. The present application does not make any limitation on the number of the selected initial training data.

[0064] Step 203: Repeat the training data selection operation until the data selection end condition is met; wherein, the training data selection operation includes: determining the respective representative parameters of each candidate data in the target data set according to the encoded features of each candidate data in the target data set and the encoded features of each training data in the training sample set, where the representative parameters are used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoded feature space; and selecting at least one candidate data from the target data set as training data according to the respective representative parameters of each candidate data in the target data set, and migrating the training data from the target data set to the training sample set.

[0065] After the server selects the initial training data and migrates it from the target data set to the training sample set, it can perform a round of training data selection operation based on the current target data set and training sample set. The training data selection operation is as follows: According to the encoding features of each candidate data in the target data set and the encoding features of each training data in the training sample set, determine the respective representative parameters corresponding to each candidate data in the target data set (used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoding feature space). Furthermore, according to the respective representative parameters corresponding to each candidate data in the target data set, select at least one candidate data from the target data set as training data again, and migrate the selected training data from the target data set to the training sample set, that is, add the selected training data to the training sample set, and at the same time delete the training data from the target data set.

[0066] After performing a round of training data selection operation, both the target data set and the training sample set will be updated. Specifically, the target data set will be missing the training data selected in this round of training data selection operation, while the training sample set will increase the training data selected in this round of training data selection operation. Furthermore, the server can further perform the next round of training data selection operation based on the updated target data set and training sample set, and repeat the execution of several rounds of training data selection operations until the data selection end condition is met.

[0067] It should be noted that the data selection end condition here is a condition used to measure whether to stop performing the training data selection operation, and it can be set according to actual needs. Exemplarily, the data selection end condition can be that the number of training data included in the training sample set reaches a preset number, the data selection condition can also be that the number of rounds of training data selection operations performed reaches a preset number of rounds, and the data selection condition can also be that the ratio of the number of training data in the training sample set to the number of candidate data in the target data set reaches a preset ratio. This application does not make any limitations on the data selection end condition here.

[0068] In a possible implementation manner, when the server performs each round of training data selection operation, it can determine the respective representative parameters corresponding to each candidate data in the target data set through a discriminator network. Specifically, the server can regard each candidate data in the target data set and each training data in the training sample set as data to be processed. Then, through the discriminator network, according to the encoding features of each data to be processed and the data type corresponding to each data to be processed, determine the respective representative parameters corresponding to each data to be processed belonging to the target data set. The data type here is used to characterize whether the corresponding data to be processed belongs to the target data set or the training sample set.

[0069] Exemplarily, a three-layer Multilayer Perception (MLP) can be pre-designed as the discriminator network. The working principle of this discriminator network can be as Figure 4 shown. This discriminator network processes the input feature f through three layers of Rectified Linear Activation Function (ReLU), and then obtains the representative parameters corresponding to the data to be processed. Specifically, the server can use the data to be processed as the input feature f, or the server can also preprocess the data to be processed based on a preset data preprocessing rule and use the preprocessed result as the input feature f. Then, the server can input the input feature f into the discriminator network composed of three layers of MLP. Each layer of MLP in this discriminator network analyzes and processes the input feature f in turn, and finally outputs the representative parameters corresponding to the data to be processed. The above discriminator network uses a lightweight three-layer MLP design. Compared with the traditional complex discriminator network structure, this lightweight discriminator network can be trained quickly and can predict representative parameters quickly.

[0070] In the embodiments of the present application, when performing each round of training data selection operation, each candidate data in the target data set and each training data in the training sample set can be used as the data to be processed, and the data type of each data to be processed is marked, that is, it is marked whether each data to be processed belongs to the target data set or the training sample set. Then, the server can input the encoded feature and data type of each data to be processed into the discriminator network. This discriminator network analyzes and processes the input data and correspondingly outputs the representative parameters corresponding to each data to be processed belonging to the target data set.

[0071] It should be noted that the representative parameters corresponding to the candidate data in the target data set can characterize the distance between the candidate data and each training data in the encoded feature space. It should be understood that the distance between the candidate data and the training data in the encoded feature space can reflect the feature difference between the candidate data and the training data; the farther the distance between the candidate data and the training data in the encoded feature space, the greater the feature difference between the candidate data and the training data, and the closer the distance between the candidate data and the training data in the encoded feature space, the smaller the feature difference between the candidate data and the training data. It should be understood that when there are multiple training data in the training sample set, the distance characterized by the representative parameters corresponding to the candidate data is determined by comprehensively considering the distance between the candidate data and these multiple training data in the encoded feature space, that is, the representative parameters can comprehensively reflect the feature difference between the candidate data and each training data.

[0072] In this way, through the above discriminator network, the representative parameters corresponding to each candidate data in the target dataset can be determined, and the representative parameters corresponding to each candidate data in the target dataset can be efficiently determined, and the accuracy and reliability of the determined representative parameters can be ensured.

[0073] It should be understood that in practical applications, the server can also determine the representative parameters corresponding to each candidate data in the target dataset in other ways; for example, the server can calculate the distance between the encoded feature of each candidate data and the encoded feature of each training data. Here, the encoded features of the candidate data and the training data are both obtained through step 201. When specifically determining the distance between the encoded features of the candidate data and the training data, the server can calculate the Euclidean distance or cosine distance between the encoded feature of the candidate data and the encoded feature of the training data, so as to use the calculated distance to represent the distance between the candidate data and the training data in the encoded feature space; furthermore, calculate the average value of the distances between the encoded feature of the candidate data and the encoded features of each training data as the representative parameter corresponding to the candidate data. This application does not make any limitation on the method for determining the representative parameters corresponding to the candidate data in the target dataset.

[0074] In a possible implementation manner, when the server performs each round of training data selection operation, it can select the candidate data with the largest corresponding representative parameter from the target dataset according to the representative parameters corresponding to each candidate data in the target dataset. Here, the larger the representative parameter, the farther the corresponding candidate data is from each training data in the training sample set in the encoded feature space, and the smaller the representative parameter, the closer the corresponding candidate data is to each training data in the training sample set in the encoded feature space.

[0075] As Figure 3 shown in (b), (c), ……, (n) in Figure 3 As shown in (b) in Figure 3As shown in (c), the candidate data corresponding to the solid circle 3 can be selected as the training data in this round of training data selection operation. When selecting this training data, it is necessary to comprehensively consider the encoding features and the distances between the solid circle 1 and the solid circle 2 in the encoding feature space, and then determine the candidate data corresponding to the encoding feature with the farthest comprehensive distance as the training data; and so on, perform each round of training data selection operation until the data selection end condition is met, and the training data corresponding to each of the solid circles 1 to 8 can be selected in the encoding feature space.

[0076] It should be understood that in practical applications, when the server performs each round of training data selection operation, it can select only one candidate data as the training data, or select multiple candidate data as the training data, such as selecting multiple candidate data with relatively high corresponding representative parameters as the training data. This application does not make any limitation on the number of training data selected in each round of training data selection operation.

[0077] After the server completes the training data selection operation through the above operations, it can push the selected training data to the terminal device facing the annotation object, and the terminal device will display the training data accordingly; furthermore, the annotation object can configure the corresponding annotation result for the training data for subsequent training of the target model based on the training data and its corresponding annotation result.

[0078] Taking the training data selected by the server as motion sequence data as an example, Figure 5 FIG. is an exemplary data annotation interface schematic diagram provided for the embodiment of the present application. As Figure 5 shown, the annotation object can trigger the acquisition of training data in the training sample set by clicking the data recommendation control in the interface, and the acquired training data will be displayed in the data display area on the left side of the interface; correspondingly, the annotation object can configure the corresponding annotation result for the training data displayed in the data display area through the annotation result configuration control on the right side of the interface. When the running sequence data is x i =(s i,1 , s i,2 , ……, s i,T ), where s i,1 , s i,2 , ……, s i,T are the action data in the motion sequence data x i , and s i,j ∈R J×3 (j = 1, 2, ……, T) represents the three-dimensional coordinates of J body joints in the action posture represented by the jth action data, configuring the corresponding annotation result for the motion sequence data x i is essentially to configure the corresponding annotation result for each action data s i in the running sequence data x i,jConfigure a binary label vector c i,j = {0, 1} m , c i,j,k = 1 (k = 1, 2, ……, m) indicates that the action posture represented by the action data s i,j belongs to the k-th action type.

[0079] When the above training data selection method selects training data from the target dataset, it will select training data from the target dataset according to the distances between the candidate data in the target dataset and each training data that has been selected into the training sample set in the encoded feature space. The training data selected in this way is usually the candidate data that is more representative in the encoded feature space; the so-called representativeness can be understood as the selected training data can represent the candidate data group in the target dataset with similar features to it, and the candidate data groups represented by each selected training data are different, that is, the selected training data is evenly and dispersedly distributed in the encoded feature space. The target model trained based on such training data usually has better model performance. The target model can process various types of data more accurately and has better robustness and generalization; and compared with the scheme of manually selecting training data by experts, the scheme of automatically selecting training data provided by the embodiments of the present application has higher training data selection efficiency and lower implementation cost.

[0080] As mentioned in the embodiments above Figure 2 When the server performs feature encoding processing on each candidate data in the target dataset, it can use the target feature encoder to perform feature encoding processing on each candidate data in the target dataset, so as to obtain the encoded features of each candidate data. The training method of the target feature encoder will be introduced in detail below through method embodiments.

[0081] See Figure 6 , Figure 6 which is a schematic flowchart of the training method of the target feature encoder provided by the embodiments of the present application. For the convenience of description, the following embodiments still take the server as the execution subject of the training method for introduction. As Figure 6 shown, the training method includes the following steps:

[0082] Step 601: Construct the first positive sample, the second positive sample and the negative sample corresponding to the reference data.

[0083] In the embodiments of the present application, before the server trains the target feature encoder, it can first construct the corresponding first positive sample, second positive sample and negative sample for the reference data.

[0084] It should be noted that the above-mentioned reference data is the training data used when training the target feature encoder, and the data type of the reference data is the data type of the data that the target feature encoder needs to process; for example, assuming that the target feature encoder is used to perform feature encoding processing on motion sequence data, then the reference data should be motion sequence data; in one possible implementation method, the reference data can be candidate data in the target data set.

[0085] The first positive sample and the second positive sample corresponding to the reference data are two positive samples constructed based on the reference data, and the semantics represented by the first positive sample and the second positive sample are the same or similar to the semantics represented by the reference data; exemplarily, the first positive sample and the second positive sample can be obtained by performing data enhancement processing on the reference data based on different data enhancement processing parameters using data enhancement processing operations. The negative sample corresponding to the reference data is a negative sample constructed based on the reference data, and the semantics represented by the negative sample is usually very different from the semantics represented by the reference data; exemplarily, the negative sample can be obtained by performing data enhancement processing on the reference data using reverse data enhancement processing operations.

[0086] The reason for constructing two corresponding positive samples and one negative sample based on the reference data is to match the contrastive learning idea in the embodiment of the present application. The contrastive learning idea in the embodiment of the present application is that the predicted coding features of each positive sample constructed based on the same reference data should be close, while the predicted coding features of the positive sample and the negative sample constructed based on the same reference data should be quite different. The present application aims to construct a loss function from the above two dimensions to train the target feature encoder. Training the target feature encoder from these two dimensions is conducive to better constraining the position of the contrastive learning representation in the feature space, thereby ensuring that the trained target feature encoder has better feature coding performance.

[0087] In one possible implementation, when the target feature encoder to be trained is used to encode motion sequence data, the reference data should be motion sequence data, including multiple frames of motion data arranged in sequence to reflect motion postures. Typically, motion sequence data is captured at a high sampling rate (such as 120FPS), and the difference between the motion data of adjacent frames is small, which can lead to ambiguity, even if the target feature encoder is confused when processing the motion data of adjacent frames; in order to eliminate this ambiguity, the embodiment of the present application proposes a data enhancement method for motion sequence data.

[0088] Specifically, for each frame of action data in the reference data, the server can determine the associated action data for each frame corresponding to the action data in the reference data according to the arrangement position of the action data in the reference data and the reference time window; and determine the enhanced action data corresponding to the action data according to the difference between the action data and its corresponding associated action data for each frame and the action data itself; furthermore, according to the enhanced action data corresponding to each frame of action data in the reference data, determine the enhanced reference data corresponding to the reference data.

[0089] Exemplarily, for the j-th frame of action data s in the reference data i,j , the server can, according to the arrangement position of the j-th frame of action data s i,j in the reference data and the length of the preset reference time window, determine the associated action data for each frame corresponding to the j-th frame of action data s i,j in the reference data; for example, assuming that the length of the reference time window is t, and the sampling rate and sampling time interval corresponding to the reference data are r and l respectively, the server can determine the associated action data sampling parameter n = t * r / l, and then can determine that in the reference data, from the (j - nl)-th frame of action data s i,j-nl to the (j + nl)-th frame of action data s i,j+nl (excluding the j-th frame of action data itself), are all the associated action data corresponding to the j-th frame of action data s i,j .

[0090] Furthermore, the server can strengthen the features expressed by the action data itself according to the difference between the action data and its corresponding associated action data for each frame and the action data itself, to obtain the enhanced action data corresponding to the action data.

[0091] As an example, the server can construct the enhanced action data corresponding to the action data in the following way: for each frame of associated action data arranged before the action data in the reference data, subtract the associated action data from the action data to obtain the enhanced value of the associated action data; for each frame of associated action data arranged after the action data in the reference data, subtract the action data from the associated action data to obtain the enhanced value of the associated action data. Then, arrange the enhanced values of each frame of associated action data and the action data itself according to the arrangement position of each frame of associated action data and the action data in the reference data, to obtain the enhanced action data corresponding to the action data.

[0092] Exemplarily, for the j-th frame of action data s in the reference data i,j , its corresponding associated action data includes the (j - nl)-th frame of action data s i,j-nl to the (j + nl)-th frame of action data s i,j+nl(Excluding the j-th frame action data itself). For each associated action data arranged before the j-th frame action data s in the reference data, i.e., the action data s of the (j - nl)-th frame i,j to the action data s of the (j - l)-th frame i,j-nl , the server can subtract the associated action data from the j-th frame action data s i,j-l itself to obtain the enhancement value of each frame of associated action data, i.e., s i,j -s i,j 、……、s i,j-nl -s i,j 、……、s i,j-l . For each associated action data arranged after the j-th frame action data s in the reference data, i.e., the action data s of the (j + l)-th frame i,j to the action data s of the (j + nl)-th frame i,j+l , the server can subtract the j-th frame action data s i,j+nl itself from the associated action data to obtain the enhancement value of each frame of associated action data, i.e., s i,j -s i,j+l 、……、s i,j -s i,j+nl 、……、s i,j .

[0093] Furthermore, the server can arrange the enhancement values of each frame of associated action data and the j-th frame action data s i,j-nl according to the arrangement positions of each frame of associated action data (i.e., the action data s of the (j - nl)-th frame i,j+nl to the action data s of the (j + nl)-th frame i,j ) and the j-th frame action data s i,j in the reference data, so as to obtain the enhanced action data s i,j ' corresponding to the j-th frame action data s i,j . The enhanced action data s i,j ' is specifically shown as follows:

[0094] s i,j ' = (s i,j -s i,j-nl ,…,s i,j -s i,j-l ,s i,j ,s i,j+l -s i,j ,…,s i,j+nl -s i,j )

[0095] After the server obtains the enhanced action data corresponding to each frame of action data in the reference data, it can arrange the enhanced action data corresponding to each frame of action data accordingly according to the arrangement positions of each frame of action data in the reference data, so as to obtain the enhanced reference data corresponding to the reference data. For example, for the reference data xi , and its corresponding enhanced reference data x i ’ = (s i,1 ’, s i,2 ’, …, s i,T ’).

[0096] Correspondingly, when the server constructs the first positive sample, the second positive sample, and the negative sample corresponding to the reference data, it can construct the first positive sample, the second positive sample, and the negative sample based on the enhanced reference data corresponding to the reference data; that is, perform data augmentation processing based on the enhanced reference data corresponding to the reference data to obtain the first positive sample, the second positive sample, and the negative sample.

[0097] In this way, when the reference data used to train the target feature encoder is motion sequence data, by performing the above-mentioned enhancement processing on the motion sequence data, the ambiguity between the action data in the motion sequence data can be clarified, that is, it is ensured that there are obvious differences between the enhanced action data of each frame in the enhanced reference data obtained after the enhancement processing, which is beneficial to enabling the trained target feature encoder to learn richer knowledge.

[0098] In a possible implementation manner, when the target feature encoder to be trained is used to encode motion sequence data, the above-mentioned reference data should be motion sequence data, which includes multiple frames of action data arranged in sequence and used to reflect action postures. At this time, the first positive sample, the second positive sample, and the negative sample corresponding to the reference data can be constructed in the following manner:

[0099] Based on the first enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the first positive sample. The first enhancement parameter group includes a first perturbation enhancement parameter and a first downsampling enhancement parameter; based on the second enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the second positive sample. The second enhancement parameter group includes a second perturbation enhancement parameter and a second downsampling enhancement parameter, and the second enhancement parameter group is different from the first enhancement parameter group; based on the third enhancement parameter group, perform perturbation enhancement processing, downsampling enhancement processing, and inversion enhancement processing on the reference data to obtain the negative sample. The third enhancement parameter group includes a third perturbation enhancement parameter and a third downsampling enhancement parameter. Among them, the above-mentioned perturbation enhancement processing is used to add perturbations based on the reference data according to the perturbation addition rule indicated by the perturbation enhancement parameter; the above-mentioned downsampling enhancement processing is used to extract action data based on the reference data according to the data extraction rule indicated by the downsampling enhancement parameter; the inversion enhancement processing is used to reverse the arrangement order of the action data in the reference data.

[0100] Specifically, the embodiments of the present application propose three data augmentation processing methods for motion sequence data, namely perturbation augmentation processing, downsampling augmentation processing, and inversion augmentation processing. Among them, perturbation augmentation processing and downsampling augmentation processing do not essentially change the semantics of the motion sequence data, while inversion augmentation processing can essentially change the semantics of the motion sequence data.

[0101] Specifically, the perturbation augmentation processing is a data augmentation processing method implemented according to the perturbation data augmentation strategy D pb (x’, p). Its basic principle is that the semantics of the motion sequence data are robust to small perturbations. In the embodiments of the present application, exemplarily, two random perturbations, data missing perturbation and data disorder perturbation, can be applied to each frame of action data in the reference data. It should be understood that if the reference data has been strengthened to obtain the strengthened reference data corresponding to the reference data, then at this time, two random perturbations can be applied to each frame of strengthened action data in the strengthened reference data. The addition methods of these two random perturbations are shown in Equation (1):

[0102]

[0103] where t pb is the probability of performing perturbation augmentation processing on the i-th frame of strengthened action data, t md is the probability of applying data missing perturbation; p i is the perturbation increase probability corresponding to the i-th frame of strengthened action data, which can be randomly assigned. When p i is less than t pb ·t md , data missing perturbation can be applied to the i-th frame of strengthened action data, that is, the i-th frame of strengthened action data is made missing; when p i is greater than or equal to t pb ·t md and less than t pb , data disorder perturbation can be applied to the i-th frame of strengthened action data, that is, the i-th frame of strengthened action data is replaced with the j-th frame of strengthened action data; when p i is greater than or equal to t pb , no perturbation augmentation processing is performed on the i-th frame of strengthened action data.

[0104] The downsampling augmentation processing is a data augmentation strategy D dsThe data augmentation processing method implemented by (x’, a, δ) has a basic principle that humans can successfully recognize motion sequences played at different playback rates, that is, the semantics of motion sequence data are largely independent of the playback rate. In the embodiments of the present application, exemplarily, the running sequence data can be downsampled based on a random playback rate and offset, and several frames of action data are extracted therefrom to achieve the effect of playing the motion sequence data at an accelerated speed; it should be understood that if the reference data has been enhanced to obtain the enhanced reference data corresponding to the reference data, then at this time, the enhanced reference data can be downsampled based on a random playback rate and offset, and several frames of enhanced action data are extracted therefrom to achieve the effect of playing the enhanced reference data at an accelerated speed. The implementation method of this downsampling enhancement processing is shown in Equation (2):

[0105]

[0106] where a is the offset and δ is the sampling interval corresponding to the downsampling process, and both can be set randomly. Through the above Equation (2), the a-th frame of enhanced action data s′ can be randomly sampled from the enhanced reference data x’ a , the (a + δ)-th frame of enhanced action data s′ a+δ , the (a + 2δ)-th frame of enhanced action data s′ a+2δ , ……, the (a + (n ds -1)δ)-th frame of enhanced action data

[0107] The reverse enhancement processing is a data augmentation processing method implemented according to the reverse data augmentation strategy D re (x’), and its basic principle is that the semantics of the motion sequence data played in reverse are different from those of the motion sequence data played forward. In the embodiments of the present application, exemplarily, the arrangement order of the action data in the reference data can be directly reversed to achieve the effect of playing the reference data in reverse; it should be understood that if the reference data has been enhanced to obtain the enhanced reference data corresponding to the reference data, then at this time, the arrangement order of each frame of enhanced action data in the enhanced reference data can be reversed to achieve the effect of playing the enhanced reference data in reverse. The implementation method of this reverse enhancement processing is shown in Equation (3):

[0108] D re (x′) = (s′ T , s′ T-1 , …, s′ 1 ) (3)

[0109] Through the above Equation (3), the arrangement order of each frame of enhanced action data in the enhanced reference data can be completely reversed, thereby changing the semantics of the enhanced reference data.

[0110] Taking the construction of the first positive sample, the second positive sample, and the negative sample corresponding to the reference data based on the enhanced reference data corresponding to the reference data as an example, the server can respectively construct the first positive sample v based on the reference enhanced data x' through the following formulas (4), (5), and (6). 1 + and the second positive sample v 2 + and the negative sample v r - :

[0111]

[0112]

[0113]

[0114] where p 1 is the first perturbation enhancement parameter, a 1 and δ 1 are the first downsampling enhancement parameters. The first perturbation enhancement parameter and the first downsampling enhancement parameter together form the first enhancement parameter group. p 2 is the second perturbation enhancement parameter, a 2 and δ 2 are the second downsampling enhancement parameters. The second perturbation enhancement parameter and the second downsampling enhancement parameter together form the second enhancement parameter group; it should be understood that there is at least one enhancement parameter in the second enhancement parameter group that is different from the corresponding enhancement parameter in the first enhancement parameter group. p 3 is the third perturbation enhancement parameter, a 3 and δ 3 are the third downsampling enhancement parameters. The third perturbation enhancement parameter and the third downsampling enhancement parameter together form the third enhancement parameter group; it should be understood that the enhancement parameters in the third enhancement parameter group may be the same as or different from the enhancement parameters in the first enhancement parameter group, and the present application does not make any limitations in this regard.

[0115] In this way, by constructing the first positive sample, the second positive sample, and the negative sample corresponding to the reference data through the above method, it can effectively ensure that the constructed first positive sample and the second positive sample are semantically similar, and the constructed first positive sample and the negative sample are semantically quite different, which is conducive to constructing a loss function based on the semantic relationship between the first positive sample, the second positive sample, and the negative sample, and realizing unsupervised training of the target feature encoder based on this loss function, so that the target feature encoder can accurately encode the features of various motion sequence data.

[0116] It should be understood that in practical applications, the server can also construct the first positive sample, the second positive sample, and the negative sample corresponding to the reference data in other ways, and the present application does not make any limitations on the construction methods of the first positive sample, the second positive sample, and the negative sample.

[0117] Step 602: Respectively perform feature encoding processing on the first positive sample and the negative sample through the first feature encoder to obtain the predicted encoding features of the first positive sample and the negative sample respectively; perform feature encoding processing on the second positive sample through the second feature encoder to obtain the predicted encoding feature of the second positive sample.

[0118] After the server constructs the first positive sample, the second positive sample, and the negative sample corresponding to the reference data, it can respectively perform feature encoding processing on the first positive sample and the negative sample through the pre-constructed first feature encoder, so as to obtain the predicted encoding features of the first positive sample and the negative sample; and can perform feature encoding processing on the second positive sample through the pre-constructed second feature encoder to obtain the predicted encoding feature of the second positive sample.

[0119] It should be noted that the model structures of the above first feature encoder and second feature encoder are the same, and their initial model parameters can be the same or different. The reason for setting two feature encoders and using the two feature encoders to process different training samples is to conform to the idea of contrastive learning; using the first feature encoder to process the first positive sample and the negative sample respectively to obtain the predicted encoding features of the first positive sample and the negative sample respectively, and using the second feature encoder to process the second positive sample to obtain the predicted encoding feature of the second positive sample. Furthermore, with the goal of making the predicted encoding features of the first positive sample and the second positive sample close to each other and making the predicted encoding features of the first positive sample and the negative sample far from each other, train the first feature encoder and the second feature encoder to achieve unsupervised training of the feature encoder, and at the same time ensure that the feature encoder can learn relatively rich knowledge during the training process, so as to obtain better training results.

[0120] Exemplarily, when the target feature encoder to be trained is used to process motion sequence data, the embodiments of the present application may use a spatio-temporal Transformer network as the above-mentioned first feature encoder and second feature encoder. The spatio-temporal Transformer network includes a spatial Transformer and a temporal Transformer. Among them, the spatial Transformer is used to model each frame of action data to extract low-level features embedding the relationships between body parts, that is, to independently calculate the correlations between each pair of joints in each frame of action data. The temporal Transformer is used to model the temporal trajectory of each frame of action data. By combining the outputs of the above-mentioned spatial Transformer and temporal Transformer, the predicted coding features extracted from the motion sequence data can be obtained.

[0121] It should be understood that in practical applications, the above-mentioned first feature encoder and second feature encoder may also be other model structures, and the present application does not make any limitation on the model structures of the above-mentioned first feature encoder and second feature encoder.

[0122] Step 603: Construct a target loss function according to the predicted coding features of the first positive sample, the second positive sample, and the negative sample respectively; based on the target loss function, adjust the model parameters of the first feature encoder; based on the model parameters of the first feature encoder, adjust the model parameters of the second feature encoder.

[0123] After the server obtains the predicted coding features of the first positive sample and the negative sample through the first feature encoder, and obtains the predicted coding features of the second positive sample through the second feature encoder, it can construct a target loss function according to the predicted coding features of the first positive sample, the second positive sample, and the negative sample respectively. Then, with the goal of minimizing the target loss function, the model parameters of the first feature encoder are adjusted, and based on the model parameters of the first feature encoder, the model parameters of the second feature encoder are also adjusted accordingly.

[0124] In a possible implementation manner, the server may specifically construct the target loss function in the following way: perform normalization processing on the predicted coding features of the first positive sample, the second positive sample, and the negative sample respectively to obtain the normalized coding features of the first positive sample, the second positive sample, and the negative sample respectively; then, determine a first loss value according to the difference between the normalized coding feature of the first positive sample and the normalized coding feature of the second positive sample; determine a second loss value according to the difference between the normalized coding feature of the first positive sample and the normalized coding feature of the negative sample; furthermore, determine the target loss function according to the above-mentioned first loss value and second loss value.

[0125] Exemplarily, taking the first positive sample, the second positive sample, and the negative sample as the first positive sample v mentioned above 1 + , the second positive sample v 2 + and the negative sample v r - as examples, after performing feature encoding processing on the first positive sample v 1 + , the predicted coding feature E(v 1 + ) corresponding to this first positive sample v will be obtained. After performing feature encoding processing on the second positive sample v 1 + , the predicted coding feature E(v 2 + ) corresponding to this second positive sample v will be obtained. After performing feature encoding processing on the negative sample v 2 + , the predicted coding feature E(v 2 + ) corresponding to this negative sample v will be obtained. Furthermore, the server can perform normalization processing on the predicted coding features E(v r - ), E(v r - ), and E(v r - ) respectively through the following formulas (7), (8), and (9): 1 + 2 + r - 1 + wherein,

[0126]

[0127]

[0128]

[0129] where is the normalized coding feature corresponding to the first positive sample v 1 + , is the normalized coding feature corresponding to the second positive sample v 2 + , is the normalized coding feature corresponding to the negative-positive sample v r - .

[0130] Furthermore, the server can construct the target loss function L according to the following formula (10) based on the normalized encoded features corresponding to the above-mentioned first positive sample v 1 + , second positive sample v 2 + and negative positive sample v r - respectively: s :

[0131]

[0132] wherein, is the first loss value determined according to the normalized encoded features corresponding to the first positive sample v 1 + and the second positive sample v 2 + respectively, is the second loss value determined according to the normalized encoded features corresponding to the first positive sample v 1 + and the negative sample v r - respectively, where τ is a hyperparameter.

[0133] In this way, by constructing the target loss function in the above manner, not only the predicted encoded features corresponding to the two positive samples are compared, but also the predicted encoded features corresponding to the positive sample and the negative sample are compared, which can better constrain the position of the representation of contrastive learning in the feature space, thus facilitating better training of the feature encoder based on this target loss function.

[0134] When the server adjusts the model parameters of the first feature encoder based on the target loss function, it can minimize this target loss function as the goal and adjust the model parameters of this first feature encoder. After adjusting the model parameters of this first feature encoder, use the momentum feature encoder to further adjust the model parameters of the second feature encoder through the following formula (11):

[0135] θ k ←αθ k-1 +(1-α)θ q (11)

[0136] wherein, θ on the left side of the arrow k is the model parameter of the adjusted second feature encoder, θ on the right side of the arrow k-1 is the model parameter of the second feature encoder before adjustment, θ q is the model parameter of the first feature encoder, and α∈[0,1) is the momentum coefficient.

[0137] In a scenario where the target feature encoder to be trained is used to perform feature encoding processing on motion sequence data, the target loss function constructed in the above manner can be regarded as a sequence-level loss function. In this scenario, embodiments of the present application can further construct a frame-level loss function (i.e., the reference loss function) based on the differences between the predicted coding features of action data, so as to jointly train the first feature encoder using the sequence-level loss function and the frame-level loss function.

[0138] When constructing the frame-level loss function, each frame of action data in the first positive sample can be regarded as the first action data, and each frame of action data in the second positive sample can be regarded as the second action data. For each frame of the first action data, the server can determine the positively correlated action data and negatively correlated action data corresponding to the first action data among each frame of the second action data included in the second positive sample; furthermore, for each frame of the first action data, a reference loss function corresponding to the first action data can be constructed according to the first action data, and the predicted coding features of each of its corresponding positively correlated action data and negatively correlated action data; here, the predicted coding feature of the first action data belongs to the predicted coding feature of the first positive sample, and the predicted coding features of the positively correlated action data and negatively correlated action data belong to the predicted coding feature of the second positive sample.

[0139] Specifically, since the first positive sample and the second positive sample are semantically related, for each frame of the first action data in the first positive sample, the second action data that is semantically related to the first action data can be found in the second positive sample as the positively correlated action data corresponding to the first action data. Here, the positively correlated action data can also be understood as the positive sample corresponding to the first action data. In addition, since the first positive sample and the second positive sample usually both include a large amount of action data, for each frame of the first action data in the first positive sample, the second action data that is significantly different in semantics from the first action data can also be found in the second positive sample as the negatively correlated action data corresponding to the first action data. Here, the negatively correlated action data can also be the negative sample corresponding to the first action data.

[0140] As an example, if the first positive sample and the second positive sample are obtained by performing downsampling enhancement processing on the reference data based on the first downsampling enhancement parameter and the second downsampling enhancement parameter respectively, the server can determine the corresponding positively correlated action data and negatively correlated action data for each frame of the first action data in the following manner: Based on the first upsampling restoration parameter corresponding to the first downsampling enhancement parameter, perform upsampling restoration processing on the first positive sample to obtain a first restored positive sample; Based on the second upsampling restoration parameter corresponding to the second downsampling enhancement parameter, perform upsampling restoration processing on the second positive sample to obtain a second restored positive sample; The number of first restored action data included in the first restored positive sample here is the same as the number of second restored action data included in the second restored positive sample; Furthermore, for each frame of the first restored action data, according to the arrangement position of the first restored action data in the first restored positive sample and the first reference frame interval, determine the positively correlated action data and negatively correlated action data corresponding to the first restored action data among the frames of the second restored action data included in the second restored positive sample.

[0141] Specifically, if the first positive sample and the second positive sample are obtained by performing downsampling enhancement processing on the reference data based on the first downsampling enhancement parameter and the second downsampling enhancement parameter respectively, then the number of first action data included in the first positive sample and the number of second action data included in the second positive sample may be different. To determine the positive and negative samples at the frame level, the server can perform upsampling restoration processing on the first positive sample based on the first upsampling restoration parameter corresponding to the first downsampling enhancement parameter to obtain a first restored positive sample; The upsampling restoration processing here is used to restore more action data based on the first action data included in the first positive sample, that is, to restore a first restored positive sample including more first restored action data; When specifically performing the upsampling restoration processing, interpolation processing can be performed based on the restoration ratio indicated by the first upsampling restoration parameter on the basis of the first action data already existing in the first positive sample, so as to restore more first restored action data. Similarly, for the second positive sample, the server can also perform upsampling restoration processing on the second positive sample correspondingly based on the second upsampling restoration parameter corresponding to the second downsampling enhancement parameter to obtain a second restored positive sample including more second restored action data. To facilitate the construction of positive and negative samples at the frame level, the number of restored action data included in the first restored positive sample and the second restored positive sample restored here should be equal.

[0142] Furthermore, for each frame of the first restored positive sample's first restored action data, the server can determine the arrangement position of the first restored action data in the first restored positive sample as the search center for the positively correlated action data in the second restored positive sample. Then, the server can determine the second restored action data whose distance from the search center is within the first reference frame interval as the positively correlated action data corresponding to the first restored action data, and determine the second restored action data whose distance from the search center is outside the first reference frame interval as the negatively correlated action data corresponding to the second restored action data. For example, when determining the corresponding positively and negatively correlated action data for the i-th frame of the first restored action data in the first restored positive sample, the server can determine the i-th frame of the second restored action data in the second restored positive sample as the search center. Then, the server can determine that the j-th frame of the second restored action data in the second restored positive sample that satisfies |i - j| < t nb is the positively correlated action data, and determine that the j-th frame of the second restored action data in the second restored positive sample that satisfies |i - j| ≥ t nb is the negatively correlated action data, where t nb is the first reference frame interval, which can be equal to 12, for example.

[0143] In this way, through the above method, based on the continuity and local invariance of actions in the motion sequence data, positive and negative samples at the frame level are constructed, which can ensure the reliability of the constructed positive and negative samples at the frame level, and is beneficial to better training the feature encoder based on the loss function constructed accordingly at the frame level.

[0144] It should be understood that in practical applications, the server can also use other methods to determine the positively and negatively correlated action data corresponding to each frame of the first action data. For example, the server can determine the positively correlated action data with relatively similar semantics and the negatively correlated action data with relatively different semantics corresponding to each frame of the first action data in the second positive sample according to the relationship between the first downsampling enhancement parameter and the second downsampling enhancement parameter. The present application does not limit the method for determining the positively and negatively correlated action data corresponding to the first action data in any way.

[0145] For each frame of the first action data, the server can extract the predicted coding features of the first action data from the predicted coding features of the first positive sample; for each frame of the positively correlated action data corresponding to the first action data, the server can extract the predicted coding features of the positively correlated action data from the predicted coding features of the second positive sample; for each frame of the negatively correlated action data corresponding to the first action data, the server can extract the predicted coding features of the negatively correlated action data from the predicted coding features of the second positive sample. Furthermore, the server can construct a reference loss function corresponding to the first action data, that is, a frame-level loss function, according to the predicted coding features of the first action data, the predicted coding features of each positively correlated action data, and the predicted coding features of each negatively correlated action data.

[0146] As an example, the server can construct a reference loss function corresponding to each frame of the first action data in the following way: determine a third loss value according to the difference between the predicted coding features of the first action data and the predicted coding features of each positively correlated action data; determine a fourth loss value according to the difference between the predicted coding features of the first action data and the predicted coding features of each negatively correlated action data; furthermore, determine the reference loss function corresponding to the first action data according to the third loss value and the fourth loss value.

[0147] Exemplarily, the predicted coding features of the first positive sample and the second positive sample can be respectively normalized to obtain the normalized coding features of the first positive sample and the normalized coding features of the second positive sample For the i-th frame of the first action data in the first positive sample, the server can extract its predicted coding features from the normalized coding features of the first positive sample For the positively correlated action data corresponding to the i-th frame of the first action data (that is, the j-th frame of the second action data in the second positive sample, where j satisfies |i - j| < t ), the server can extract its predicted coding features from the normalized coding features of the second positive sample nb ) For the negatively correlated action data corresponding to the i-th frame of the first action data (that is, the j-th frame of the second action data in the second positive sample, where j satisfies |i - j| ≥ t ), the server can extract its predicted coding features from the normalized coding features of the second positive sample nb ) For the negatively correlated action data corresponding to the i-th frame of the first action data (that is, the j-th frame of the second action data in the second positive sample, where j satisfies |i - j| ≥ t

[0148] Furthermore, the server can, through the following formula (12), according to the predicted coding features of the first action data and the predicted coding features of each positively correlated action data corresponding to the first action data And the prediction coding features of each negatively correlated action data corresponding to the first action data Construct a reference loss function L corresponding to the first action data f :

[0149]

[0150] Where, Ω + is the set of positively correlated action data corresponding to the first action data, Ω + ={j|t nb >|i - j|}, Ω - is the set of negatively correlated action data corresponding to the first action data, Ω - ={j|t nb ≤|i - j|}. is the third loss value determined according to the prediction coding features of the first action data and each positively correlated action data respectively, is the fourth loss value determined according to the prediction coding features of the first action data and each negatively correlated action data respectively, where τ is a hyperparameter.

[0151] In this way, by constructing the reference loss function at the frame level in the above manner, not only the prediction coding features of the first action data and the prediction coding features of each positively correlated action data corresponding thereto are compared, but also the prediction coding features of the first action data and the prediction coding features of each negatively correlated action data corresponding thereto are compared. This can better constrain the position of the feature representation of contrastive learning in the feature space at the frame level, thereby facilitating better training of the feature encoder based on this reference loss function.

[0152] When the server constructs the reference loss function at the frame level, when training the first feature encoder, the server can jointly adjust the parameters of the first feature encoder based on the target loss function and the reference loss function corresponding to each frame of the first action data. Exemplarily, the server can integrate the reference loss functions corresponding to each frame of the first action data to obtain a comprehensive reference loss function L F , and the integration method here can be, for example, summing or averaging the reference loss functions corresponding to each frame of the first action data; furthermore, the server can construct a comprehensive loss function L through the following formula (13) and adjust the model parameters of the first feature encoder based on this comprehensive loss function.

[0153] L = L s + ωL F (13)

[0154] Where, ω is a preset hyperparameter for fusing the comprehensive reference loss function.

[0155] In this way, by combining the loss function at the joint sequence level and the loss function at the frame level to train the feature encoder, two-layer motion contrast learning is achieved, which helps the trained feature encoder learn more information features, that is, helps the trained feature encoder to have better feature extraction performance.

[0156] Step 604: When the model training end condition is satisfied, determine the first feature encoder as the target feature encoder.

[0157] In practical applications, the server can repeatedly execute the above steps 601 to 603 based on different reference data, so as to achieve multiple rounds of iterative training for the first feature encoder and the second feature encoder. When it is determined that the model training end condition is satisfied, the server can use the first feature encoder obtained from the current training as the target feature encoder for actual application, that is Figure 2 the feature encoder used in the illustrated embodiment to perform feature encoding processing on the candidate data in the target dataset.

[0158] It should be understood that the reason for using the first feature encoder as the target feature encoder is that in the above model training process, this first feature encoder is the main feature encoder to be trained, and it is the feature encoder trained based on the constructed loss function; while the second feature encoder is a feature encoder adaptively trained to achieve contrast learning, and the model parameters of this second feature encoder are adjusted according to the adjustment of the first feature encoder; in other words, in the model training process, improving the performance of the first feature encoder is the essential model training goal, and adjusting the model parameters of the second feature encoder is a means adaptively executed to achieve this goal.

[0159] It should be understood that the above model training end condition is a condition used to measure whether to stop training the feature encoder, and it can be set according to actual needs; for example, this model training end condition can be that the number of training rounds for the first feature encoder reaches a preset number of rounds, or for another example, the performance of the first feature encoder reaches a preset performance requirement, etc. The present application does not make any limitation on this model training end condition here.

[0160] In a possible implementation manner, the server can apply the target feature encoder trained in the above manner to the target model, and this target model is the model that needs to be trained using the training data selected in the Figure 2 illustrated embodiment. When specifically training this target model, the server can obtain the annotation results corresponding to each training data in the training sample set, and then, based on each training data in the training sample set and its corresponding annotation result, train other network structures in the target model except the target feature encoder.

[0161] Exemplarily, taking the target model to be trained as an action recognition model as an example, the action recognition model is used to recognize the action types involved based on motion sequence data. The action recognition model includes a target feature encoder and a classifier. When training the action recognition model, the server can first obtain the annotation results corresponding to each training data in the training sample set. The annotation result corresponding to the training data is used to represent the action types involved in the training data. Furthermore, the server can train the classifier in the action recognition model based on each training data in the training sample set and its corresponding annotation result. During the process of training the classifier, the model parameters of the target feature encoder in the action recognition model remain unchanged.

[0162] The reason why the model parameters of the target feature encoder in the target model do not need to be adjusted during training is that the target feature encoder trained by the training method of the feature encoder introduced above has better feature encoding performance. The target feature encoder is trained based on the contrast learning mechanism and does not refer to specific annotation results during the training process. Therefore, the target feature encoder can universally and accurately extract the encoding features of various types of data, that is, the target feature encoder can be well adapted to various feature extraction tasks.

[0163] In addition, when applying the target feature encoder to the target model, since the model parameters of the target feature encoder do not need to be adjusted during the training of the target model, only the parameters of other structures except the target feature encoder need to be adjusted. Therefore, the time required to train the target model can be effectively shortened. And when the task to be executed by the target model needs to be adjusted, the annotation result adapted to the task corresponding to the training data can be directly obtained. Furthermore, based on the training data and its corresponding annotation result, other structures except the target feature encoder in the target model can be retrained, which can shorten the time for the target model to adapt to the new task and contribute to the rapid iterative update of the target model.

[0164] The above training method of the target feature encoder adopts the idea of contrast learning to train the target feature encoder. On the one hand, there is no need to specifically construct annotation samples, which can save the cost required for constructing annotation samples. On the other hand, the target feature encoder trained in an unsupervised manner can universally and accurately extract the encoding features of various types of data. And the embodiment of the present application also proposes a scheme of jointly training the target feature encoder using a double-layer loss function at the sequence level and the frame level, which can further ensure that the target feature encoder can learn richer information during the process of training the target feature encoder.

[0165] To facilitate a further understanding of the training data selection method provided by the embodiments of the present application, the following takes the training data to be selected for training an action recognition model as an example to provide an overall exemplary introduction to the training data selection method provided by the embodiments of the present application. It should be noted that the action recognition model is a model for recognizing the action types involved based on motion sequence data; in practical applications, the action recognition model can be applied, for example, in a game scenario, such as recognizing the actions made by a player in a motion-sensing game to trigger corresponding game operations.

[0166] In the scenario of selecting training data for training an action recognition model, the target data set can be expressed as D=(x 1 ,x 2 ,……,x N ), which is a motion sequence data set composed of N motion sequence data; where x i =(s i,1 ,s i,2 ,……,s i,T ) is a motion sequence data composed of T consecutive frames of action data for characterizing action postures, and s i,j ∈R J×3 (j = 1, 2, ……, T) represents the three-dimensional coordinates of J body joints in the action posture characterized by the j-th action data.

[0167] The embodiments of the present application aim to select a part of the motion sequence data from the target data set D to form a training sample set D train , and manually annotate the respective annotation results of each motion sequence data in the training sample set D train by an expert, that is, the expert configures a binary label vector c i ={0, 1} i,j for each action data s i,j in the motion sequence data x m , and c i,j,k = 1 (k = 1, 2, ……, m) indicates that the action posture characterized by the action data s i,j belongs to the k-th action type; then, using each motion sequence data in the training sample set D train and its respective corresponding annotation results, train the action recognition model, and then use the trained action recognition model to automatically annotate other motion sequence data in the target data set D.

[0168] Before selecting training data from the target data set, a two-layer contrastive learning mechanism can be used to train the target feature encoder first. The target feature encoder is used to perform feature encoding processing on the motion sequence data to obtain encoded features that can characterize the characteristics of the motion sequence data. Figure 7 This is a schematic diagram of the implementation principle of the two-layer contrastive learning mechanism provided by the embodiments of the present application.

[0169] In the embodiments of the present application, contrast learning can standardize the encoded feature space of motion sequence data in an unsupervised task, so that similar motion sequence data are close to each other in the encoded feature space, and motion sequence data with relatively large differences in features are far from each other in the encoded feature space. The embodiments of the present application achieve the above goals by minimizing the distance between two positive samples and maximizing the distance between a positive sample and a negative sample.

[0170] Considering that motion sequence data are usually captured at a high sampling rate (such as 120 FPS), the difference between the action data of adjacent frames is very small, which may lead to ambiguity, that is, the target feature encoder is confused when processing the action data of adjacent frames; to clarify this ambiguity, in the embodiments of the present application, for each frame of action data in the motion sequence data, within a reference time window centered on the action data, the context information of the action data is used to strengthen the semantics of the action data. The specific implementation formula is as follows:

[0171] s i,j ’ = (s i,j - s i,j-nl , …, s i,j - s i,j-l , s i,j , s i,j+l - s i,j , …, s i,j+nl - s i,j )

[0172] where n = t * r / l, t is the length of the reference time window, r is the sampling rate corresponding to the motion sequence data, and l is the sampling time interval corresponding to the motion sequence data. The embodiments of the present application use x i ’ = (s i,1 ’, s i,2 ’, …, s i,T ’) as the input of the trained feature encoder.

[0173] The embodiments of the present application can use a spatio-temporal Transformer network as the trained feature encoder. The spatio-temporal Transformer network includes a spatial Transformer and a temporal Transformer; among them, the spatial Transformer is used to model each frame of action data to extract low-level features that embed the relationships between body parts, that is, to independently calculate the correlation between each pair of joints in each frame of action data; the temporal Transformer is used to model the time trajectory of each frame of action data; by combining the outputs of the above spatial Transformer and temporal Transformer, the encoded features extracted from the motion sequence data can be obtained.

[0174] When training the target feature encoder based on the contrastive learning mechanism, two feature encoders are constructed, namely the first feature encoder Transformer q and the second feature encoder Transformer k. When training the target feature encoder, the first feature encoder Transformer q is used to process a positive sample and a negative sample constructed based on the motion sequence data, and the second feature encoder Transformer k is used to process another positive sample constructed based on the motion sequence data. Furthermore, according to the encoded features generated by the first feature encoder Transformer q and the second feature encoder Transformer k respectively, the model parameters of the first feature encoder Transformer q are adjusted, and then the model parameters of the second feature encoder Transformer k are adaptively adjusted according to the model parameters of the first feature encoder Transformer q. When it is determined that the training end condition is met, the trained first feature encoder Transformer q is used as the target feature encoder.

[0175] Under the contrastive learning mechanism, when adjusting the model parameters of the second feature encoder Transformer k according to the model parameters of the first feature encoder Transformer q, a momentum feature encoder can be used to update the model parameters of the first feature encoder Transformer k through the following formula:

[0176] θ k ←αθ k-1 +(1-α)θ q

[0177] where θ k is the adjusted model parameter of the second feature encoder Transformer k, θ k-1 is the model parameter of the second feature encoder Transformer k before adjustment, θ q is the model parameter of the first feature encoder Transformer q, and α∈[0,1) is the momentum coefficient.

[0178] The following introduces the sequence-level data augmentation method and loss function design provided by the embodiments of the present application. The embodiments of the present application provide three data augmentation methods to perform data augmentation processing on the original motion sequence data, so as to construct corresponding positive samples and negative samples.

[0179] 1) Perturbation augmentation processing. The perturbation augmentation processing is based on the perturbation data augmentation strategy D pbThe data augmentation processing method implemented by (x’, p) is based on the principle that the semantics of motion sequence data are robust to small perturbations. In the embodiments of this application, two random perturbations, namely data missing perturbation and data disorder perturbation, can be applied to each frame of action data in the motion sequence data. The addition methods of these two random perturbations are shown in the following formula:

[0180]

[0181] where t pb is the probability of performing perturbation augmentation processing on the i-th frame of action data, and t md is the probability of applying data missing perturbation; p i is the perturbation increase probability corresponding to the i-th frame of action data, which can be randomly assigned; in practical applications, t pb can be equal to 0.15, and t md can be equal to 0.9.

[0182] 2) Downsampling augmentation processing. The downsampling augmentation processing is a data augmentation processing method implemented according to the downsampling data augmentation strategy D ds (x’, a, δ). Its basic principle is that humans can successfully recognize motion sequences played at different playback rates, that is, the semantics of motion sequence data are largely independent of the playback rate. In the embodiments of this application, the motion sequence data can be downsampled based on a random playback rate and offset, and several frames of action data are extracted therefrom to achieve the effect of playing the motion sequence data at an accelerated speed. The implementation method of this downsampling augmentation processing is shown in the following formula:

[0183]

[0184] where a is the offset and δ is the sampling interval corresponding to the downsampling processing, and both can be randomly set.

[0185] 3) Reverse augmentation processing. The reverse augmentation processing is a data augmentation processing method implemented according to the reverse data augmentation strategy D re (x’). Its basic principle is that the semantics of the motion sequence data played in reverse are different from those played forward. In the embodiments of this application, negative samples can be constructed through reverse augmentation processing. The implementation method of this reverse augmentation processing is shown in the following formula:

[0186] D re (x′) = (s′ T , s′ T-1 , …, s′ 1 )

[0187] In the embodiments of this application, by combining the above three data augmentation methods, two corresponding positive samples v are constructed based on the motion sequence data 1+ and v 2 + and a negative sample v r - , and the specific construction method is as follows:

[0188]

[0189]

[0190]

[0191] The above two positive samples v are respectively subjected to feature encoding processing through the first feature encoder Transformer q and the second feature encoder Transformer k 1 + and v 2 + and a negative sample v r - , and then the obtained predicted coding features are normalized to obtain the following normalized coding features:

[0192]

[0193]

[0194]

[0195] Furthermore, a sequence-level loss function can be constructed according to the above normalized coding features through the following formula:

[0196]

[0197] By constructing the sequence-level loss function in the above manner, not only the normalized coding features corresponding to the two positive samples are compared, but also the normalized coding features corresponding to the positive sample and the negative sample are compared, so that the position of the representation of contrastive learning in the feature space can be better constrained.

[0198] In addition, the embodiment of the present application also designs a frame-level contrastive learning mechanism and a loss function. Specifically, the embodiment of the present application designs a frame-level loss function according to the continuity and local invariance of actions; the predicted coding features of each frame of action data in the motion sequence data are compared with the predicted coding features of other frames of action data in the motion sequence data. For the predicted coding features of each frame of action data s i ’s predicted coding features The predicted coding features of the action data with a relatively close distance can be defined as the positive sample, j ∈ Ω + and Ω+ = {j | t nb >| i - j |}, where t nb can be equal to 12, for example; in addition, the predicted coding features of action data with a relatively large distance from it can be used as negative samples, j ∈ Ω - and Ω - = {j | t nb ≤| i - j |}. Furthermore, a frame-level loss function can be constructed by the following formula:

[0199]

[0200] Finally, by combining the above sequence-level loss function and frame-level loss function, the final loss function is obtained:

[0201] L = L s + ωL F

[0202] Furthermore, based on this loss function, the first feature encoder Transformer q and the second feature encoder Transformer k can be trained according to the contrastive learning mechanism introduced above until the first feature encoder Transformer q that meets the training end condition is obtained, and it is used as the target feature encoder.

[0203] After completing the training of the target feature encoder, the target feature encoder can be used to perform feature encoding processing on each motion sequence data in the target dataset, so as to obtain the encoding features corresponding to each motion sequence data in the target dataset; the distribution of each motion sequence data in the target dataset in the original data space is as Figure 8 shown in (a) below, and the distribution of the encoding features corresponding to each motion sequence data in the target dataset in the encoding feature space is as Figure 8 shown in (b) below.

[0204] Furthermore, through the discriminator network including three-layer MLP designed in the embodiments of the present application, according to the encoding features corresponding to each motion sequence data in the above target dataset, representative motion sequence data can be selected from the target dataset and migrated to the training sample set D train as shown below. The embodiments of the present application use the design of a lightweight three-layer MLP, which is different from the previous complex discriminator network structure. This lightweight discriminator network can be trained quickly and can quickly predict the representativeness of motion sequence data, and only takes 1 minute of processing time for 100 pieces of data.

[0205] Specifically, the input of the discriminator network is the encoded features of each motion sequence data and the data type of the motion sequence data (indicating whether it has been migrated to the training sample set). The output of the discriminator network is the representative parameter corresponding to each motion sequence data that has not been migrated to the training sample set. This representative parameter can characterize the distance between the corresponding motion sequence data and the motion sequence data in the training sample set in the encoded feature space. The discriminator network in the embodiments of the present application can, based on the motion sequence data that has currently been selected as training data, further select the next most representative motion sequence data from the motion sequence data that has not been selected as training data and migrate it to the training sample set. The specific algorithm is as follows:

[0206] Data: Encoded feature space R (constituted by the encoded features of each motion sequence data), target data set D = (x 1 , x 2 , ……, x N ), training sample set Discriminator network C;

[0207] Result: Training sample set

[0208] 1. a ← randomly select a number {1, 2, …, N};

[0209] 2. Add (x a ), D ← D remove (x a );

[0210] 3. Before the number of motion sequence data included in reaches the preset number, repeat the following steps:

[0211] 4. Use the discriminator network C to learn the difference between and D on R:

[0212] 5. x a ← the most representative x i ∈ D;

[0213] 6. Add (x a ), D ← D remove (x a ).

[0214] Through the above motion sequence selection mechanism, a preset number of motion sequence data will be selected from the target data set as the training data to be labeled. The distribution of the training data in the encoded feature space is as shown in (c) in Figure 8 . Then, these training data can be pushed to relevant experts, and the relevant experts can configure the corresponding labeling results for them, as shown in Figure 8As shown in (d) of the figure. Furthermore, these training data and their corresponding annotation results can be used to train an action recognition model.

[0215] When the model training requirements change, the selected training data can be pushed to relevant experts, and the relevant experts can reconfigure the annotation results that match the changed requirements for it, and then retrain the action recognition model. It should be noted that when the action recognition model includes the target feature encoder described above, each time the action recognition model is trained, only the other structures of the action recognition model except the target feature encoder need to be trained. Experiments have proved that it only takes 0.007 hours to train the above action recognition model.

[0216] After completing the training of the action recognition model, the action recognition model can be used to automatically annotate the motion sequence data of other unselected training data in the target dataset. When the selected training data accounts for 15.69% of the target dataset, the F1 score of the trained action recognition model reaches 79.95%.

[0217] By comparing the distribution of the training data selected by the embodiments of the present application and the training data manually selected by experts in the encoded feature space, the Figure 9 shown comparison results are obtained. As Figure 9 shown in (a) and (b) of the figure, the training data selected by the method provided by the embodiments of the present application is more evenly distributed in the encoded feature space, more representative than the training data selected by experts. Among the training data selected by experts, there are multiple cases where the training data is highly similar, and it is easy to ignore some motion sequence data (such as Figure 9 the dotted area in (b) of the figure).

[0218] For the training data selection method described above, the present application also provides a corresponding training data selection device to enable the above training data selection method to be applied and implemented in practice.

[0219] See Figure 10 , Figure 10 is a schematic structural diagram of a training data selection device 1000 corresponding to the training data selection method shown above. As Figure 2 shown in the figure, the training data selection device 1000 includes: Figure 10 shown in the figure, the training data selection device 1000 includes:

[0220] A feature encoding module 1001, configured to perform feature encoding processing on each candidate data in the target dataset to obtain the encoded features of each candidate data in the target dataset;

[0221] An initial data selection module 1002 is configured to select at least one candidate data from the target data set as initial training data, and migrate the training data from the target data set to a training sample set; the training data included in the training sample set is unlabeled data for training a target model.

[0222] A data selection module 1003 is configured to repeatedly execute a training data selection operation until a data selection end condition is met; wherein, the training data selection operation includes: determining respective representative parameters corresponding to each candidate data in the target data set according to the respective encoding features of each candidate data in the target data set and the respective encoding features of each training data in the training sample set, where the representative parameters are used to characterize the distance between the corresponding candidate data and the training data in the training sample set in an encoding feature space; and selecting at least one candidate data from the target data set as training data according to the respective representative parameters corresponding to each candidate data in the target data set, and migrating the training data from the target data set to the training sample set.

[0223] Optionally, the data selection module 1003 is specifically configured to:

[0224] Regard each candidate data in the target data set and each training data in the training sample set as data to be processed.

[0225] Through a discriminator network, determine respective representative parameters corresponding to each data to be processed belonging to the target data set according to the respective encoding features of each data to be processed and the respective data types corresponding to each data to be processed; the data type is used to characterize that the corresponding data to be processed belongs to the target data set or the training sample set.

[0226] Optionally, the encoding features are generated by a target feature encoder, and the device further includes an encoder training module; the encoder training module includes:

[0227] A sample construction sub-module is configured to construct a first positive sample, a second positive sample, and a negative sample corresponding to reference data.

[0228] A sample processing sub-module is configured to respectively perform feature encoding processing on the first positive sample and the negative sample through a first feature encoder to obtain respective predicted encoding features of the first positive sample and the negative sample; and perform feature encoding processing on the second positive sample through a second feature encoder to obtain the predicted encoding feature of the second positive sample.

[0229] An encoder training sub-module, configured to construct a target loss function according to the predicted encoding features of the first positive sample, the second positive sample, and the negative sample respectively; based on the target loss function, adjust the model parameters of the first feature encoder; based on the model parameters of the first feature encoder, adjust the model parameters of the second feature encoder;

[0230] An encoder determination sub-module, configured to determine the first feature encoder as the target feature encoder when the model training end condition is satisfied.

[0231] Optionally, the reference data is motion sequence data, and the motion sequence data includes multiple frames of action data for reflecting action postures; the encoder training module further includes:

[0232] A data enhancement sub-module, configured to, for each frame of action data in the reference data, determine each frame of associated action data corresponding to the action data in the reference data according to the arrangement position of the action data in the reference data and the reference time window; determine the enhanced action data corresponding to the action data according to the difference between the action data and its corresponding each frame of associated action data, and the action data;

[0233] And determine the enhanced reference data corresponding to the reference data according to the enhanced action data corresponding to each frame of action data in the reference data.

[0234] Then the sample construction sub-module is specifically configured to:

[0235] Construct the first positive sample, the second positive sample, and the negative sample based on the enhanced reference data corresponding to the reference data.

[0236] Optionally, the data enhancement sub-module is specifically configured to:

[0237] For each frame of associated action data arranged before the action data in the reference data, subtract the associated action data from the action data to obtain the enhancement value of the associated action data; for each frame of associated action data arranged after the action data in the reference data, subtract the action data from the associated action data to obtain the enhancement value of the associated action data;

[0238] Arrange the enhancement values of each frame of the associated action data and the action data according to the arrangement positions of each frame of the associated action data and the action data in the reference data to obtain the enhanced action data corresponding to the action data.

[0239] Optionally, the reference data is motion sequence data, and the motion sequence data includes multiple frames of action data for reflecting action postures; specifically, the sample construction sub-module is configured to:

[0240] Based on the first enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the first positive sample; the first enhancement parameter group includes a first perturbation enhancement parameter and a first downsampling enhancement parameter;

[0241] Based on the second enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the second positive sample; the second enhancement parameter group includes a second perturbation enhancement parameter and a second downsampling enhancement parameter; the second enhancement parameter group is different from the first enhancement parameter group;

[0242] Based on the third enhancement parameter group, perform perturbation enhancement processing, downsampling enhancement processing, and inversion enhancement processing on the reference data to obtain the negative sample; the third enhancement parameter group includes a third perturbation enhancement parameter and a third downsampling enhancement parameter;

[0243] Wherein, the perturbation enhancement processing is used to add perturbations based on the reference data according to the perturbation addition rules indicated by the perturbation enhancement parameters; the downsampling enhancement processing is used to extract action data based on the reference data according to the data extraction rules indicated by the downsampling enhancement parameters; the inversion enhancement processing is used to reverse the arrangement order of the action data in the reference data.

[0244] Optionally, the encoder training sub-module is specifically configured to:

[0245] Perform normalization processing on the predicted coding features of the first positive sample, the second positive sample, and the negative sample respectively to obtain the normalized coding features of the first positive sample, the second positive sample, and the negative sample respectively;

[0246] Determine a first loss value according to the difference between the normalized coding features of the first positive sample and the normalized coding features of the second positive sample; determine a second loss value according to the difference between the normalized coding features of the first positive sample and the normalized coding features of the negative sample;

[0247] Determine the target loss function according to the first loss value and the second loss value.

[0248] Optionally, the reference data is motion sequence data, and the motion sequence data includes multiple frames of action data for reflecting action postures; the first positive sample includes multiple frames of first action data corresponding to the action data in the reference data, and the second positive sample includes multiple frames of second action data corresponding to the action data in the reference data;

[0249] The sample construction sub-module is further configured to determine, for each frame of the first action data, the positively correlated action data and the negatively correlated action data corresponding to the first action data among the frames of the second positive samples included in the second positive samples;

[0250] The encoder training sub-module is further configured to, for each frame of the first action data, construct a reference loss function corresponding to the first action data according to the first action data, and the predicted coding features of the respective positively correlated action data and negatively correlated action data corresponding thereto; the predicted coding feature of the first action data belongs to the predicted coding feature of the first positive sample, and the predicted coding features of the positively correlated action data and the negatively correlated action data belong to the predicted coding features of the second positive sample; and, based on the target loss function and the reference loss functions respectively corresponding to the frames of the first action data, adjust the model parameters of the first feature encoder.

[0251] Optionally, the first positive sample is obtained by performing down-sampling enhancement processing on the reference data using a first down-sampling enhancement parameter, and the second positive sample is obtained by performing down-sampling enhancement processing on the reference data using a second down-sampling enhancement parameter; the sample construction sub-module is specifically configured to:

[0252] Perform up-sampling restoration processing on the first positive sample based on a first up-sampling restoration parameter corresponding to the first down-sampling enhancement parameter to obtain a first restored positive sample; perform up-sampling restoration processing on the second positive sample based on a second up-sampling restoration parameter corresponding to the second down-sampling enhancement parameter to obtain a second restored positive sample; the number of the first restored action data included in the first restored positive sample is the same as the number of the second restored action data included in the second restored positive sample;

[0253] For each frame of the first restored action data, determine the positively correlated action data and the negatively correlated action data corresponding to the first restored action data among the frames of the second restored action data included in the second restored positive sample according to the arrangement position of the first restored action data in the first restored positive sample and a first reference frame interval.

[0254] Optionally, the encoder training sub-module is specifically configured to:

[0255] Determine a third loss value according to the difference between the predicted coding feature of the first action data and the predicted coding features of the respective positively correlated action data; determine a fourth loss value according to the difference between the predicted coding feature of the first action data and the predicted coding features of the respective negatively correlated action data;

[0256] Determine the reference loss function according to the third loss value and the fourth loss value.

[0257] Optionally, the target model includes the target feature encoder; the apparatus further includes: a model training module;

[0258] The model training module is configured to obtain the annotation result corresponding to each training data in the training sample set; and train the network structure in the target model except the target feature encoder based on each training data in the training sample set and its corresponding annotation result.

[0259] When the above training data selection apparatus selects training data from the target data set, it will select training data from the target data set according to the distance between the candidate data in the target data set and each training data that has been selected into the training sample set in the encoded feature space. The training data selected in this way is usually the candidate data that is more representative in the encoded feature space; the so-called representativeness can be understood as the selected training data can represent the candidate data group in the target data set with similar features, and the candidate data groups represented by each selected training data are different, that is, the selected training data is evenly and dispersedly distributed in the encoded feature space. The target model trained based on such training data usually has better model performance. The target model can process various types of data more accurately, and has better robustness and generalization; and compared with the solution of manually selecting training data by experts, the automatic selection training data solution provided by the embodiments of the present application has higher training data selection efficiency and lower implementation cost.

[0260] The embodiments of the present application further provide a computer device for selecting training data. The computer device may specifically be a terminal device or a server. Hereinafter, the terminal device and the server provided by the embodiments of the present application will be introduced from the perspective of hardware implementation.

[0261] See Figure 11 , Figure 11 is a schematic structural diagram of the terminal device provided by the embodiments of the present application. As Figure 11 shown, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc. Taking the terminal as a computer as an example:

[0262] Figure 11 Shows a block diagram of a part of the structure of a computer related to the terminal provided by the embodiments of the present application. Refer toFigure 11 , the computer includes components such as a Radio Frequency (RF) circuit 1110, a memory 1120, an input unit 1130 (which includes a touch panel 1131 and other input devices 1132), a display unit 1140 (which includes a display panel 1141), a sensor 1150, an audio circuit 1160 (which can be connected to a speaker 1161 and a microphone 1162), a wireless fidelity (WiFi) module 1170, a processor 1180, and a power supply 1190. Those skilled in the art can understand that Figure 11 the computer structure shown in

[0263] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functional applications and data processing of the computer by running the software programs and modules stored in the memory 1120. The memory 1120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer (such as audio data, a phone book, etc.). In addition, the memory 1120 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0264] The processor 1180 is the control center of the computer, connects various parts of the entire computer using various interfaces and lines, and executes various functions of the computer and processes data by running or executing the software programs and / or modules stored in the memory 1120, and by calling the data stored in the memory 1120. Optionally, the processor 1180 can include one or more processing units; preferably, the processor 1180 can integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1180.

[0265] In the embodiments of the present application, the processor 1180 included in the terminal is further used to execute the steps of any one of the implementation manners of the training data selection method provided in the embodiments of the present application.

[0266] See Figure 12 , Figure 12Schematic diagram of a server 1200 provided by an embodiment of the present application. The server 1200 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1222 (e.g., one or more processors) and a memory 1232, and one or more storage media 1230 (e.g., one or more mass storage devices) for storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage media 1230 may be transient storage or persistent storage. The programs stored in the storage media 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1222 may be configured to communicate with the storage media 1230 and execute a series of instruction operations in the storage media 1230 on the server 1200.

[0267] The server 1200 may further include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0268] The steps performed by the server in the above embodiments may be based on the Figure 12 server structure shown.

[0269] Among them, the CPU 1222 is used to execute the steps of any implementation manner of the training data selection method provided by the embodiment of the present application.

[0270] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute any implementation manner of the training data selection method described in the foregoing embodiments.

[0271] The embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any implementation manner of the training data selection method described in the foregoing embodiments.

[0272] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0273] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0274] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0275] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0276] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store computer programs.

[0277] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0278] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for selecting training data, characterized in that, the method includes: Performing feature encoding processing on each candidate data in the target dataset to obtain the encoded features of each candidate data in the target dataset; Selecting at least one candidate data from the target dataset as the initial training data, and migrating the training data from the target dataset to the training sample set; the training data included in the training sample set is the data to be labeled for training the target model, and the target model is a motion recognition model for identifying action types based on motion sequence data; Repeatedly performing the training data selection operation until the data selection end condition is met; Wherein, the training data selection operation includes: determining the respective representative parameters of each candidate data in the target dataset according to the encoded features of each candidate data in the target dataset and the encoded features of each training data in the training sample set, and the representative parameter is used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoded feature space; and selecting at least one candidate data from the target dataset as the training data according to the respective representative parameters of each candidate data in the target dataset, and migrating the training data from the target dataset to the training sample set; Wherein, the encoded feature is generated by a target feature encoder, and the target feature encoder is trained in the following manner: Constructing the first positive sample, the second positive sample and the negative sample corresponding to the reference data; Performing feature encoding processing on the first positive sample and the negative sample respectively through the first feature encoder to obtain the predicted encoded features of the first positive sample and the negative sample respectively; performing feature encoding processing on the second positive sample through the second feature encoder to obtain the predicted encoded features of the second positive sample; Constructing a target loss function according to the predicted encoded features of the first positive sample, the second positive sample and the negative sample respectively; adjusting the model parameters of the first feature encoder based on the target loss function; adjusting the model parameters of the second feature encoder based on the model parameters of the first feature encoder; When the model training end condition is met, determining the first feature encoder as the target feature encoder; Wherein, the reference data is motion sequence data, and the motion sequence data includes multiple frames of action data for reflecting action postures, and the method further includes: For each frame of action data in the reference data, determining the associated action data of each frame corresponding to the action data in the reference data according to the arrangement position of the action data in the reference data and the reference time window; determining the enhanced action data corresponding to the action data according to the difference between the action data and its corresponding associated action data of each frame and the action data; Determining the enhanced reference data corresponding to the reference data according to the enhanced action data corresponding to each frame of action data in the reference data; Then the constructing the first positive sample, the second positive sample and the negative sample corresponding to the reference data includes: Construct the first positive sample, the second positive sample, and the negative sample based on the enhanced reference data corresponding to the reference data.

2. The method according to claim 1, wherein, the determining of the respective representative parameters corresponding to the respective candidate data in the target data set according to the respective coding features of the respective candidate data in the target data set and the respective coding features of the respective training data in the training sample set includes: Regarding each candidate data in the target data set and each training data in the training sample set as data to be processed; Through a discriminator network, determine the respective representative parameters corresponding to the respective data to be processed belonging to the target data set according to the respective coding features of the respective data to be processed and the respective data types corresponding to the respective data to be processed; the data type is used to characterize that the corresponding data to be processed belongs to the target data set or the training sample set.

3. The method according to claim 1, wherein, the determining of the enhanced action data corresponding to the action data according to the difference between the action data and the respective frame-associated action data corresponding thereto, and the action data includes: For each frame-associated action data arranged before the action data in the reference data, subtract the associated action data from the action data to obtain the enhancement value of the associated action data; for each frame-associated action data arranged after the action data in the reference data, subtract the action data from the associated action data to obtain the enhancement value of the associated action data; Arrange the respective enhancement values of each frame of the associated action data and the action data according to the arrangement positions of each frame of the associated action data and the action data in the reference data to obtain the enhanced action data corresponding to the action data.

4. The method according to claim 1, wherein, the constructing of the first positive sample, the second positive sample, and the negative sample corresponding to the reference data includes: Based on the first enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the first positive sample; the first enhancement parameter group includes a first perturbation enhancement parameter and a first downsampling enhancement parameter; Based on the second enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the second positive sample; the second enhancement parameter group includes a second perturbation enhancement parameter and a second downsampling enhancement parameter; the second enhancement parameter group is different from the first enhancement parameter group; Based on the third enhancement parameter group, perform perturbation enhancement processing, downsampling enhancement processing, and inversion enhancement processing on the reference data to obtain the negative sample; the third enhancement parameter group includes a third perturbation enhancement parameter and a third downsampling enhancement parameter; Among them, the perturbation enhancement process is used to add perturbations based on the reference data according to the perturbation addition rules indicated by the perturbation enhancement parameters; the downsampling enhancement process is used to extract action data based on the reference data according to the data extraction rules indicated by the downsampling enhancement parameters; the inversion enhancement process is used to reverse the arrangement order of the action data in the reference data.

5. The method according to claim 1 or 4, wherein, constructing the target loss function according to the prediction coding features of the first positive sample, the second positive sample and the negative sample respectively includes: performing normalization processing on the prediction coding features of the first positive sample, the second positive sample and the negative sample respectively to obtain the normalized coding features of the first positive sample, the second positive sample and the negative sample respectively; determining a first loss value according to the difference between the normalized coding feature of the first positive sample and the normalized coding feature of the second positive sample; determining a second loss value according to the difference between the normalized coding feature of the first positive sample and the normalized coding feature of the negative sample; determining the target loss function according to the first loss value and the second loss value.

6. The method according to claim 1, wherein, the first positive sample includes multiple frames of first action data corresponding to the action data in the reference data, and the second positive sample includes multiple frames of second action data corresponding to the action data in the reference data; the method further includes: for each frame of the first action data, determining the positively correlated action data and the negatively correlated action data corresponding to the first action data among the frames of second action data included in the second positive sample; for each frame of the first action data, constructing a reference loss function corresponding to the first action data according to the prediction coding features of the first action data, its corresponding positively correlated action data and negatively correlated action data; the prediction coding feature of the first action data belongs to the prediction coding features of the first positive sample, and the prediction coding features of the positively correlated action data and the negatively correlated action data belong to the prediction coding features of the second positive sample; the adjusting the model parameters of the first feature encoder based on the target loss function includes: adjusting the model parameters of the first feature encoder based on the target loss function and the reference loss functions corresponding to each frame of the first action data respectively.

7. The method according to claim 6, wherein, the first positive sample is obtained by performing downsampling enhancement processing on the reference data using the first downsampling enhancement parameter, and the second positive sample is obtained by performing downsampling enhancement processing on the reference data using the second downsampling enhancement parameter; for each frame of the first action data, determining the positively correlated action data and the negatively correlated action data corresponding to the first action data among the frames of second action data included in the second positive sample includes: Based on the first upsampling restoration parameter corresponding to the first downsampling enhancement parameter, perform upsampling restoration processing on the first positive sample to obtain a first restored positive sample; based on the second upsampling restoration parameter corresponding to the second downsampling enhancement parameter, perform upsampling restoration processing on the second positive sample to obtain a second restored positive sample; the number of first restored action data included in the first restored positive sample is the same as the number of second restored action data included in the second restored positive sample; For each frame of the first restored action data, according to the arrangement position of the first restored action data in the first restored positive sample and the first reference frame interval, determine the positively correlated action data and negatively correlated action data corresponding to the first restored action data among each frame of second restored action data included in the second restored positive sample.

8. The method according to claim 6, wherein, The constructing the reference loss function corresponding to the first action data according to the prediction coding features of the first action data and its corresponding positively correlated action data and negatively correlated action data respectively includes: Determine a third loss value according to the difference between the prediction coding feature of the first action data and the prediction coding features of each of the positively correlated action data; determine a fourth loss value according to the difference between the prediction coding feature of the first action data and the prediction coding features of each of the negatively correlated action data; Determine the reference loss function according to the third loss value and the fourth loss value.

9. The method according to claim 1, wherein, The target model includes the target feature encoder; The method further includes: Obtain the annotation result corresponding to each training data in the training sample set; Based on each training data in the training sample set and its corresponding annotation result, train the network structure in the target model except the target feature encoder.

10. A training data selection device, wherein, The device includes: A feature encoding module, configured to perform feature encoding processing on each candidate data in the target data set to obtain the encoding feature of each candidate data in the target data set; An initial data selection module, configured to select at least one candidate data from the target data set as the initial training data, and migrate the training data from the target data set to the training sample set; the training data included in the training sample set is the data to be annotated for training the target model, and the target model is a motion recognition model for identifying action types based on motion sequence data; A data selection module, configured to repeatedly execute a training data selection operation until a data selection end condition is met; wherein, the training data selection operation includes: determining respective representative parameters corresponding to each candidate data in the target dataset according to the respective encoding features of each candidate data in the target dataset and the respective encoding features of each training data in the training sample set, where the representative parameters are used to characterize the distance between the corresponding candidate data and the training data in the training sample set in the encoding feature space; and selecting at least one candidate data from the target dataset as training data according to the respective representative parameters corresponding to each candidate data in the target dataset, and migrating the training data from the target dataset to the training sample set. Wherein, the encoding features are generated by a target feature encoder, and the device further includes an encoder training module, and the encoder training module includes: A sample construction sub-module, configured to construct a first positive sample, a second positive sample, and a negative sample corresponding to reference data. A sample processing sub-module, configured to respectively perform feature encoding processing on the first positive sample and the negative sample through a first feature encoder to obtain respective predicted encoding features of the first positive sample and the negative sample; and perform feature encoding processing on the second positive sample through a second feature encoder to obtain the predicted encoding feature of the second positive sample. An encoder training sub-module, configured to construct a target loss function according to the respective predicted encoding features of the first positive sample, the second positive sample, and the negative sample; adjust the model parameters of the first feature encoder based on the target loss function; and adjust the model parameters of the second feature encoder based on the model parameters of the first feature encoder. An encoder determination sub-module, configured to determine the first feature encoder as the target feature encoder when a model training end condition is met. Wherein, the reference data is motion sequence data, the motion sequence data includes multiple frames of action data for reflecting action postures, and the encoder training module further includes: A data enhancement sub-module, configured to, for each frame of action data in the reference data, determine respective frames of associated action data corresponding to the action data in the reference data according to the arrangement position of the action data in the reference data and a reference time window; determine enhanced action data corresponding to the action data according to the difference between the action data and its respective frames of associated action data and the action data; and determine enhanced reference data corresponding to the reference data according to the respective enhanced action data corresponding to each frame of action data in the reference data. Then the sample construction sub-module is specifically configured to: Construct the first positive sample, the second positive sample, and the negative sample based on the enhanced reference data corresponding to the reference data.

11. The device according to claim 10, wherein, The data selection module is specifically configured to: Regard each candidate data in the target dataset and each training data in the training sample set as data to be processed. Through the discriminator network, according to the encoding features of each piece of the data to be processed and the data type corresponding to each piece of the data to be processed, determine the representative parameters corresponding to each piece of the data to be processed belonging to the target data set; the data type is used to characterize that the corresponding data to be processed belongs to the target data set or the training sample set.

12. The apparatus according to claim 10, wherein, the data enhancement sub-module is specifically configured to: For each frame of associated action data arranged before the action data in the reference data, subtract the associated action data from the action data to obtain the enhancement value of the associated action data; for each frame of associated action data arranged after the action data in the reference data, subtract the action data from the associated action data to obtain the enhancement value of the associated action data; Arrange the enhancement values of each frame of the associated action data and the action data according to the arrangement positions of each frame of the associated action data and the action data in the reference data to obtain the enhanced action data corresponding to the action data.

13. The apparatus according to claim 10, wherein, the sample construction sub-module is specifically configured to: Based on the first enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the first positive sample; the first enhancement parameter group includes a first perturbation enhancement parameter and a first downsampling enhancement parameter; Based on the second enhancement parameter group, perform perturbation enhancement processing and downsampling enhancement processing on the reference data to obtain the second positive sample; The second enhancement parameter group includes a second perturbation enhancement parameter and a second downsampling enhancement parameter; The second enhancement parameter group is different from the first enhancement parameter group; Based on the third enhancement parameter group, perform perturbation enhancement processing, downsampling enhancement processing and inversion enhancement processing on the reference data to obtain the negative sample; The third enhancement parameter group includes a third perturbation enhancement parameter and a third downsampling enhancement parameter; wherein, the perturbation enhancement processing is used to add perturbations based on the reference data according to the perturbation addition rule indicated by the perturbation enhancement parameter; the downsampling enhancement processing is used to extract action data based on the reference data according to the data extraction rule indicated by the downsampling enhancement parameter; the inversion enhancement processing is used to reverse the arrangement order of the action data in the reference data.

14. The apparatus according to claim 10 or 13, wherein, the encoder training sub-module is specifically configured to: Perform normalization processing on the predicted encoding features of the first positive sample, the second positive sample and the negative sample respectively to obtain the normalized encoding features of the first positive sample, the second positive sample and the negative sample respectively; Determine a first loss value according to the difference between the normalized encoding feature of the first positive sample and the normalized encoding feature of the second positive sample; determine a second loss value according to the difference between the normalized encoding feature of the first positive sample and the normalized encoding feature of the negative sample; Determine the target loss function according to the first loss value and the second loss value.

15. The device according to claim 10, wherein, the first positive sample includes multiple frames of first action data corresponding to the action data in the reference data, and the second positive sample includes multiple frames of second action data corresponding to the action data in the reference data; the sample construction sub-module is further configured to, for each frame of the first action data, determine the positively correlated action data and the negatively correlated action data corresponding to the first action data from among the respective frames of second action data included in the second positive sample; the encoder training sub-module is further configured to, for each frame of the first action data, construct a reference loss function corresponding to the first action data according to the prediction coding features of the first action data and its respective positively correlated action data and negatively correlated action data; the prediction coding feature of the first action data belongs to the prediction coding features of the first positive sample, and the prediction coding features of the positively correlated action data and the negatively correlated action data belong to the prediction coding features of the second positive sample; and, based on the target loss function and the respective reference loss functions corresponding to each frame of the first action data, adjust the model parameters of the first feature encoder.

16. The device according to claim 15, wherein, the first positive sample is obtained by performing downsampling enhancement processing on the reference data using a first downsampling enhancement parameter, and the second positive sample is obtained by performing downsampling enhancement processing on the reference data using a second downsampling enhancement parameter; the sample construction sub-module is specifically configured to: perform upsampling restoration processing on the first positive sample based on a first upsampling restoration parameter corresponding to the first downsampling enhancement parameter to obtain a first restored positive sample; perform upsampling restoration processing on the second positive sample based on a second upsampling restoration parameter corresponding to the second downsampling enhancement parameter to obtain a second restored positive sample; the number of first restored action data included in the first restored positive sample is the same as the number of second restored action data included in the second restored positive sample; for each frame of the first restored action data, determine the positively correlated action data and the negatively correlated action data corresponding to the first restored action data from among the respective frames of second restored action data included in the second restored positive sample according to the arrangement position of the first restored action data in the first restored positive sample and a first reference frame interval.

17. The device according to claim 15, wherein, the encoder training sub-module is specifically configured to: determine a third loss value according to the difference between the prediction coding feature of the first action data and the respective prediction coding features of the positively correlated action data; determine a fourth loss value according to the difference between the prediction coding feature of the first action data and the respective prediction coding features of the negatively correlated action data; determine the reference loss function according to the third loss value and the fourth loss value.

18. The device according to claim 10, wherein, The target model includes the target feature encoder; the device further includes a model training module; The model training module is configured to obtain the annotation results corresponding to each training data in the training sample set; Based on each training data in the training sample set and its corresponding annotation result, train the network structure in the target model except the target feature encoder.

19. A computer device, characterized in that, the device includes a processor and a memory; the memory is used to store a computer program; the processor is configured to execute the training data selection method according to any one of claims 1 to 9 based on the computer program.

20. A computer-readable storage medium, characterized in that, the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the training data selection method according to any one of claims 1 to 9.

21. A computer program product, including a computer program or instruction, characterized in that, when the computer program or the instruction is executed by a processor, the training data selection method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Sample screening method and system, equipment and medium

    CN112508092A

  • Resume information processing method and device, electronic equipment and storage medium

    CN113806544A

  • Image retrieval model training method and device, equipment and storage medium

    CN114020950A