Classification model training method and device

By introducing K PU classification models in semi-supervised learning to generate pseudo-labels for unlabeled samples and combining labeled samples for model training, the problem of low pseudo-label quality on class imbalanced data sets is solved, and the accuracy of the classification model is improved.

CN120020898APending Publication Date: 2025-05-20HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311549737.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

When the existing semi-supervised learning method is trained on datasets with class imbalance, the generated pseudo-label quality is low, resulting in poor prediction classification accuracy of the classification model.

Method used

By introducing K positive sample label-free (PU) classification models, pseudo-labels are generated for unlabeled samples, and the classification model is trained based on the labeled sample subset and the pseudo-label sample subset, and the parameters of the feature extraction network and classifier are updated.

Benefits of technology

It effectively alleviates the biased problem of pseudo-labels, improves the labeling quality of pseudo-labels, and thus improves the semi-supervised training effect and classification accuracy of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020898A_ABST
    Figure CN120020898A_ABST
Patent Text Reader

Abstract

The invention provides a classification model training method and device, and the method comprises the steps: obtaining a target training data set, the target training data set comprises a labeled sample subset and an unlabeled sample subset, and the target training data set comprises K types of samples; respectively inputting the unlabeled samples into the K PU classification models to obtain output results of the K PU classification models; on the basis of output results of the K PU classification models, pseudo labels of unlabeled samples are obtained; and training the classification model based on the labeled sample subset and the pseudo-labeled sample subset, and updating parameters of the feature extraction network and parameters of the classifier. According to the method, the pseudo labels are generated for the unlabeled samples through the PU classification model, so that the problem that the pseudo labels are biased is effectively relieved, the labeling quality of the pseudo labels is improved, and semi-supervised training of the classification model is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular, to a method and apparatus for training a classification model. Background Art

[0002] Semi-supervised learning (SSL) has become a promising machine learning paradigm and is widely used in fields such as natural language processing and computer vision. It can use a small amount of labeled data and a large amount of unlabeled data to train a model, thereby reducing the cost and time of data annotation, improving the performance of the model on small samples or rare classes, and enhancing the robustness of the model to noisy or abnormal data.

[0003] However, most existing SSL methods are designed for data with a balanced class distribution, while in real-world scenarios, data usually exhibits an imbalanced distribution across classes. The class imbalance problem can severely degrade the performance of SSL methods because they tend to ignore data from minority classes, misclassify minority classes, or generate biased pseudo-labels for unlabeled data. Although some methods for solving the class-imbalanced semi-supervised learning (CISSL) problem have been proposed in recent years, the existing SSL solutions have the problem that when training on an imbalanced dataset, the quality of the generated pseudo-labels is low, which in turn leads to poor prediction classification accuracy of the trained classification model. Summary of the Invention

[0004] Embodiments of this application provide a method for training a classification model, which effectively solves the problem of poor prediction classification accuracy of the trained classification model when the training data distribution is unbalanced.

[0005] In a first aspect, the present application provides a method for training a classification model. The classification model to be trained includes a feature extraction network and a classifier. The method includes obtaining a target training data set, where the target training data set includes a labeled sample subset and an unlabeled sample subset. The labeled sample subset includes a plurality of labeled samples, and the unlabeled sample subset includes a plurality of unlabeled samples. The target training data set includes samples of K categories, where K is a positive integer greater than 1. The unlabeled samples are respectively input into K positive-unlabeled (PU) classification models to obtain the output results of the K PU classification models. Each of the K PU classification models in the K PU classification models is respectively used to estimate the categories of the samples in each of the K categories. Based on the output results of the K PU classification models, pseudo-labels of the unlabeled samples are obtained. The classification model to be trained is trained based on the labeled sample subset and the pseudo-label sample subset, and the parameters of the feature extraction network and the parameters of the classifier are updated. The pseudo-label sample subset includes a plurality of samples with pseudo-labels.

[0006] After the sample is input into the classification model to be trained, the feature extraction network extracts the features of the input sample to obtain the feature representation of the sample, and then the feature representation is input into the classifier. The classifier outputs the estimated category probability distribution of the input sample based on the feature representation. For example, probability 1, probability 2... probability N, indicating that the probability that the input sample is of category A is probability 1, the probability that it is of category B is probability 2... the probability that it is of category N is probability N. The estimated category of the input sample can be determined through the probability distribution. For example, if the maximum in the output probability distribution is the probability of category A being 0.9, then the category of the input sample is determined to be category A.

[0007] The method for training the classification model provided by the present application generates pseudo-labels for unlabeled samples through K PU classification models, effectively alleviating the problem of bias in pseudo-labels, thereby improving the labeling quality of pseudo-labels and facilitating the semi-supervised training of the classification model.

[0008] In a possible implementation, before the unlabeled samples are respectively input into the K PU classification models to obtain the output results of the K PU classification models, it further includes training the classification model based on the target training data set and updating the parameters of the feature extraction network. That is to say, first, semi-supervised training is carried out, and the classification model is trained using the labeled data subset and the unlabeled data subset, so that the feature extraction network learns high-quality sample feature representations.

[0009] In this possible implementation, each of the K PU classification models includes an updated feature extraction network and a PU classifier; before inputting the unlabeled samples into the K PU classification models respectively to obtain the output results of the K PU classification models, it also includes training each of the K PU classification models based on the target training dataset and updating the PU classifier of each PU classification model.

[0010] In other words, the K PU classification models are trained using the same batch of training datasets as the semi-supervised training classification models. The PU classification model includes the feature extraction network and the PU classifier after initial training and updating on the target training dataset. And during the training process of the PU classification model, the parameters of the feature extraction network remain unchanged, only the PU classifier is updated. Therefore, the training for the PU classification model can also be called the training for the PU classifier; in this way, it can ensure that the PU classifier adapts to the latest representation learned by the feature extraction network, so as to use this representation to generate higher-quality pseudo-labels for unlabeled samples.

[0011] In another possible implementation, the training method for any one of the K PU classification models is as follows: determine the PU training dataset corresponding to the currently trained PU classification model from the target training dataset, where the PU training dataset includes a positive sample subset and an unlabeled sample subset. The positive sample subset includes multiple positive samples, and the positive samples are samples of the target category in the labeled sample subset, and the target category is the category corresponding to the currently trained PU classification model; train the PU classification model based on the PU training dataset and update the parameters of the PU classifier. By training the PU classifier category by category, it can ensure that the learning of the PU classifier is not affected by the original data distribution, so as to ensure that the trained PU classifier is fair (that is, there is no bias against samples of a certain category), and further ensure the reliability of the labels generated by the PU classifier.

[0012] In another possible implementation, a specific implementation of training the PU classification model based on the PU training dataset and updating the parameters of the PU classifier is to input the samples in the PU training dataset into the PU classification model to obtain the category estimates of each sample in the PU training dataset by the PU classification model; based on the PU loss function, update the parameters of the PU classifier. The PU loss function includes a positive risk estimation term and a negative risk estimation term. The positive risk estimation term is used to punish the PU classifier for estimating positive samples as negative samples, and the negative risk estimation term is used to punish the PU classifier for estimating the samples in the unlabeled sample subset as positive samples.

[0013] Optionally, the PU loss function is also related to the proportion of samples of the target category in the labeled sample subset. Exemplarily, when training the k-th class PU classifier, the PU loss function is as follows:

[0014]

[0015] where π k is determined by the prior distribution of the k-th class in the labeled sample subset, represents the number of samples of the k-th class in the labeled sample subset, and n u represents the total number of samples in the unlabeled sample subset. The l(.,.) function is an arbitrary surrogate 0-1 loss function, such as the sigmoid loss function:

[0016] The first term (i.e., the term before the plus sign) in the PU loss function represents the positive risk estimate, which is used to penalize the PU classifier for assigning a low probability to a positive sample. The second term (i.e., the term after the plus sign) is used to represent the non-negative risk estimate, which is used to penalize the PU classifier for assigning a high probability to an unlabeled sample. At the same time, the max function ensures that the empirical risk estimate for unlabeled samples is greater than or equal to 0, thereby reducing the risk of model overfitting.

[0017] In another possible implementation, the unlabeled samples are respectively input into K PU classification models. Before obtaining the output results of the K PU classification models, it also includes performing a first data augmentation process on the unlabeled sample subset to increase the number of samples in the unlabeled sample subset, which is beneficial to the subsequent training of the classification model.

[0018] When the sample is an image, the first data augmentation process can be one or more of flipping, translation, rotation, and cropping for image data sample augmentation.

[0019] In another possible implementation, the output result of each PU classification model includes an estimated probability, and the estimated probability indicates the probability that each PU classification model estimates the input sample as a positive sample. Based on the output results of the K PU classification models, obtaining the pseudo-labels of the unlabeled samples includes determining the target PU classification model based on the output results of the K PU classification models, where the target PU classification model outputs the largest estimated probability; and labeling the class corresponding to the target PU classification model as the pseudo-label of the unlabeled sample. By setting a dedicated PU classification model for each class of samples in the target training dataset, the bias towards the majority class is effectively alleviated, and the quality of the generated pseudo-labels is increased.

[0020] In another possible implementation, when training a classification model based on a subset of labeled samples and a subset of pseudo-labeled samples, a specific implementation of updating the parameters of the feature extraction network and the parameters of the classifier is as follows: Input the labeled samples into the classification model, and output the class estimation corresponding to the labeled samples; Based on the labels of the labeled samples and the class estimation of the labeled samples, obtain the loss value of the first loss function; Input the samples with pseudo-labels into the classification model, and output the class estimation corresponding to the samples with pseudo-labels; Based on the pseudo-labels of the samples with pseudo-labels and the class estimation of the samples with pseudo-labels, obtain the loss value of the second loss function; Based on the loss value of the first loss function and the loss value of the second loss function, update the parameters of the feature extraction network and the parameters of the classifier.

[0021] Exemplarily, the loss function of the classification model is as follows:

[0022]

[0023] Where L is the standard cross-entropy loss function, α(.) is the data augmentation function, and h(.) is the original classifier. is the feature extraction network, used to extract features. represents the labeled samples, represents the sample corresponding true label. represents the unlabeled samples, represents the sample pseudo-labels generated by the PU classifier.

[0024] Optionally, before training the classification model based on the subset of labeled samples and the subset of pseudo-labeled samples and updating the parameters of the feature extraction network and the classifier, a second data augmentation process is also included for the subset of pseudo-labeled samples. For example, when the sample is an image, the second data augmentation process can be one or more of flipping, translation, rotation, and cropping for image data sample augmentation.

[0025] In another possible implementation, the training method of the classification model provided in this application further includes training the classification model to be trained through multiple training cycles to obtain a trained target classification model. Among them, after each current training cycle is completed, samples with reliable pseudo-labels are determined from the subset of pseudo-labeled samples to obtain a subset of reliable pseudo-labeled samples, and the output result of the PU classification model corresponding to the samples with reliable pseudo-labels is greater than or equal to a preset threshold; Based on the subset of reliable pseudo-labeled samples, data augmentation is performed on the subset of labeled samples to increase the number of samples in the subset of labeled samples, thereby improving the training quality of the classification model.

[0026] In this possible implementation, to ensure the reliability of generating pseudo-labels, only the samples whose output results of the PU classifier are greater than or equal to the threshold τ are selected to form the pseudo-label list for each class k, and then a reliable pseudo-label sample set is obtained. Through the reliable pseudo-label sample subset, the labeled sample subset is augmented with data to increase the number of samples in the labeled sample subset, thereby improving the training quality of the classification model.

[0027] In another possible implementation, a specific implementation of augmenting the labeled sample subset with data based on the reliable pseudo-label sample subset is as follows: Determine the number of samples of the first class in the labeled sample subset, where the number of samples of the first class is the largest; Based on the difference between the number of samples of each class in the labeled sample subset and the number of samples of the first class, determine the sampling number of samples of each class; Based on the sampling number of samples of each class, sample the samples of each class in the reliable pseudo-label sample subset to obtain an augmented data set; Supplement the augmented data set to the labeled sample subset to obtain the labeled data subset after data augmentation.

[0028] In this possible implementation, there is no need to use complex class rebalancing rules. For each class k, only T k confident samples need to be selected from the pseudo-label list of this class, where where represents the number of samples of the k-th class in the labeled data subset, represents the number of samples of the class with the largest number of samples in the labeled sample subset, ρ belongs to (0,1) and is the sampling ratio, which is used to control the trade-off between class balance and pseudo-label quality, thus solving the class imbalance problem of the training data set and improving the training quality of the classification model.

[0029] In another possible implementation, the full training data set corresponding to the classification model includes a labeled sample set and an unlabeled sample set. The labeled sample set includes multiple labeled samples, and the unlabeled sample set includes multiple unlabeled samples; A specific implementation of obtaining the target training data set for the current training cycle is as follows: Use a random sampler to collect samples from the labeled sample set and the unlabeled sample set respectively to obtain the labeled sample subset and the unlabeled sample subset for the current batch.

[0030] In some other examples, the target training data set can also be directly the full training data set, that is, the training is not carried out in a batch training manner, and the classification model to be trained is directly trained on the full data set to obtain the trained target classification model.

[0031] Second aspect, the present application provides a training device for a classification model. The classification model includes a feature extraction network and a classifier. The device includes an acquisition module, a PU classification module, a pseudo-label annotation module, and a training module. Among them, the acquisition module is used to acquire a target training data set, which includes a labeled sample subset and an unlabeled sample subset. The labeled sample subset includes multiple labeled samples, and the unlabeled sample subset includes multiple unlabeled samples. The target training data set includes samples of K categories, and K is a positive integer greater than 1. The PU classification module is used to input the unlabeled samples into K PU classification models respectively to obtain the output results of the K PU classification models. Each PU classification model in the K PU classification models is used to estimate the categories of samples in each of the K categories. The pseudo-label annotation module is used to obtain the pseudo-labels of the unlabeled samples based on the output results of the K PU classification models. The training module is used to train the classification model based on the labeled sample subset and the pseudo-label sample subset, and update the parameters of the feature extraction network and the classifier. The pseudo-label sample subset includes multiple samples with pseudo-labels.

[0032] In a possible implementation, the training module is further used to: train the classification model based on the target training data set and update the parameters of the feature extraction network.

[0033] In another possible implementation, each PU classification model in the K PU classification models includes an updated feature extraction network and a PU classifier. The training device for the classification model provided by the present application further includes a PU learning module, which is used to train each PU classification model in the K PU classification models based on the target training data set and update the PU classifiers of each PU classification model.

[0034] In another possible implementation, the i-th PU classification model is any one of the K PU classification models. The i-th PU classification model includes an updated feature extraction network and the i-th PU classifier. The specific implementation of the PU learning module for training the i-th PU classification model is: determine the i-th PU training data set corresponding to the i-th PU classification model from the target training data set, where the i-th PU training data set includes a positive sample subset and an unlabeled sample subset. The positive sample subset includes multiple positive samples, and the positive samples are samples of the target category in the labeled sample subset. The target category is the category corresponding to the i-th PU classification model. Train the i-th PU classification model based on the i-th PU training data set and update the parameters of the i-th PU classifier.

[0035] In another possible implementation, to train the i-th PU classification model based on the i-th PU training dataset and update the parameters of the i-th PU classifier, a specific implementation is as follows: Input the samples in the i-th PU training dataset into the i-th PU classification model to obtain the class estimates of each sample in the i-th PU training dataset by the i-th PU classification model; Based on the PU loss function, update the parameters of the i-th PU classifier. The PU loss function includes a positive risk estimation term and a negative risk estimation term. The positive risk estimation term is used to penalize the i-th PU classifier for estimating positive samples as negative samples, and the negative risk estimation term is used to penalize the i-th PU classifier for estimating unlabeled samples as positive samples.

[0036] In another possible implementation, the PU loss function is also related to the proportion of samples of the target class in the subset of labeled samples.

[0037] In another possible implementation, the training device for the classification model provided in this application further includes a first data augmentation module, which is used to perform first data augmentation processing on the subset of unlabeled samples.

[0038] In another possible implementation, the output result of each PU classification model includes an estimated probability, which indicates the probability that the input sample of each PU classification model is a positive sample; The pseudo-label annotation module is specifically used for: Based on the output results of K PU classification models, determine the target PU classification model, and the target PU classification model outputs the largest estimated probability; Label the class corresponding to the target PU classification model as the pseudo-label of the unlabeled sample.

[0039] In another possible implementation, the training module is specifically used for: Input the labeled samples into the classification model and output the class estimates corresponding to the labeled samples; Based on the labels of the labeled samples and the class estimates of the labeled samples, obtain the loss value of the first loss function; Input the samples with pseudo-labels into the classification model and output the class estimates corresponding to the samples with pseudo-labels; Based on the pseudo-labels of the samples with pseudo-labels and the class estimates of the samples with pseudo-labels, obtain the loss value of the second loss function; Based on the loss value of the first loss function and the loss value of the second loss function, update the parameters of the feature extraction network and the parameters of the classifier.

[0040] In another possible implementation, the training device for the classification model provided in this application further includes a second data augmentation module, which is used to perform second data augmentation processing on the subset of pseudo-label samples.

[0041] In another possible implementation, the training device for the classification model provided in this application further includes a data augmentation module. The data augmentation module is used to determine samples with reliable pseudo-labels from the subset of pseudo-label samples after each training cycle, so as to obtain a subset of reliable pseudo-label samples. The output result of the PU classification model corresponding to the samples with reliable pseudo-labels is greater than or equal to a preset threshold; based on the subset of reliable pseudo-label samples, the subset of labeled samples is augmented with data.

[0042] In another possible implementation, a specific implementation of augmenting the subset of labeled samples with data based on the subset of reliable pseudo-label samples is to determine the number of samples of the first category in the subset of labeled samples, where the number of samples of the first category is the largest; based on the difference between the number of samples of each category in the subset of labeled samples and the number of samples of the first category, determine the sampling number of samples of each category; based on the sampling number of samples of each category, perform sample sampling on samples of each category in the subset of reliable pseudo-label samples to obtain an augmented data set; supplement the augmented data set to the subset of labeled samples to obtain an augmented subset of labeled data with data.

[0043] In another possible implementation, the full training data set corresponding to the classification model includes a labeled sample set and an unlabeled sample set. The labeled sample set includes multiple labeled samples, and the unlabeled sample set includes multiple unlabeled samples; the obtaining module is specifically used to use a random sampler to respectively collect samples from the labeled sample set and the unlabeled sample set to obtain the current batch of subset of labeled samples and the current batch of subset of unlabeled samples.

[0044] In a third aspect, an embodiment of this application provides a computing device, including a memory and a processor. Instructions are stored in the memory, and when the instructions are executed by the processor, the method described in the first aspect is implemented.

[0045] In a fourth aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0046] In a fifth aspect, an embodiment of this application further provides a computer program or a computer program product. The computer program or the computer program product includes instructions, and when the instructions are executed, the computer is made to execute the method described in the first aspect.

[0047] In a sixth aspect, an embodiment of this application further provides a chip, including at least one processor and a communication interface. The processor is used to execute the method described in the first aspect. Description of the Drawings

[0048] Figure 1 Show a schematic diagram of an artificial intelligence main framework;

[0049] Figure 2 It is a system architecture diagram of the sample processing system provided by the embodiment of the present application;

[0050] Figure 3 It shows an implementation architecture diagram of a training method for a classification model provided by the embodiment of the present application;

[0051] Figure 4 It is a flowchart of a training method for a classification model provided by the embodiment of the present application;

[0052] Figure 5 It is a schematic flowchart of a classification method provided by the embodiment of the present application;

[0053] Figure 6 It is a schematic structural diagram of a training device for a classification model provided by the embodiment of the present application;

[0054] Figure 7 It is a schematic structural diagram of a computing device provided by the embodiment of the present application. Detailed implementation manners

[0055] The term "and / or" mentioned in this article is an association relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0056] The terms "first", "second", etc. in the specification and claims of this article are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first memory chain data and the second memory chain data are used to distinguish different memory chain data, rather than to describe a specific order of the memory chain data.

[0057] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0058] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements, etc.

[0059] First, a description is made from the overall working process of the artificial intelligence system. Please refer to Figure 1 , Figure 1Shows a schematic diagram of an artificial intelligence agent framework, which describes the overall workflow of an artificial intelligence system and is applicable to the general requirements of the artificial intelligence field.

[0060] The above artificial intelligence topic framework is elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0061] (1) Infrastructure:

[0062] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips, and the aforementioned intelligent chips include but are not limited to hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), and field programmable gate array (FPGA); the basic platform includes platform guarantees and supports related to distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0063] (2) Data

[0064] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, and humidity.

[0065] (3) Data processing

[0066] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making and other methods.

[0067] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0068] Inference refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, and using formalized information for machine thinking and problem-solving based on an inference control strategy. The typical functions are search and matching.

[0069] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, ranking, prediction, etc.

[0070] (4) General capabilities

[0071] After the data undergoes the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0072] (5) Intelligent products and industry applications

[0073] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and realize landing applications. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city, intelligent terminals, etc.

[0074] The training method and device for the classification model provided by the embodiments of the present application are mainly applied to training a classification model by using a small amount of labeled sample data and a large amount of unlabeled sample data by means of semi-supervised learning. The trained classification model can be applied to the above various application fields to achieve classification or recognition functions. The processing objects of the trained classification model can be image samples, text samples, voice samples, etc. As an example, for example, in the field of intelligent terminals, a trained classification model can be configured on the intelligent terminal to achieve text classification functions. As another example, for example, in the field of autonomous driving, a trained classification model can be configured on an autonomous driving vehicle to achieve image classification functions. As yet another example, for example, in the field of intelligent security, a trained classification model can be configured in a monitoring system to achieve image recognition functions. The classification models mentioned in the foregoing examples can be trained by using the training method of the classification model provided by the present application during the training stage, thereby improving the classification accuracy of the trained classification model. It should be understood that the examples here are only for facilitating the understanding of the application scenarios of the embodiments of the present application, rather than an exhaustive list of the application scenarios of the embodiments of the present application.

[0075] The embodiments of the present application will be described below in conjunction with the accompanying drawings. Those skilled in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0076] To facilitate the understanding of this solution, in the embodiments of the present application, first, in conjunction with Figure 2 a sample processing system provided by the embodiments of the present application will be introduced. Please first refer to the attached Figure 2 , Figure 2 which is a system architecture diagram of a sample processing system provided by the embodiments of the present application. In Figure 2 , the sample processing system 200 includes an execution device 210, a training device 220, a database 230, a client device 240, a data storage system 250, and a data acquisition device 260. The execution device 210 includes a computing module 211.

[0077] Among them, the data acquisition device 260 is used to collect training data. The training data in the embodiments of the present application includes a labeled sample set and an unlabeled sample set. The labeled sample set includes a plurality of unannotated sample data, and the unlabeled sample set includes a plurality of unannotated sample data.

[0078] After the training data is collected, the data acquisition device 260 stores these training data in the database 230, and the training device 220 generates a target model / rule 201 based on the training data maintained in the database 230. How the training device 220 obtains the target model / rule 201 based on the training data will be described in more detail below. The target model / rule 201 can implement the classification function of the classification model provided by the embodiments of the present application, that is, reason about the input sample and output the category of the input sample. For example, taking a certain image as the input of the classification model and outputting the category of the image.

[0079] In actual applications, the training data maintained in the database 230 may not necessarily come from the collection of the data acquisition device 260, and it may also be received from other devices. In addition, it should be noted that the training device 220 may not necessarily train the target model / rule 201 completely based on the training data maintained in the database 230, and it may also obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation on the embodiments of the present application.

[0080] The target model / rule 201 trained according to the training device 220 can be applied to different systems or devices, such as applied to Figure 2In the execution device 210, the execution device 210 can be a terminal, such as a mobile phone, a tablet computer, a laptop computer, augmented reality (AR), virtual reality (VR), a wearable device, an intelligent robot, a vehicle-mounted terminal, etc., or can also be a server or the cloud, etc.

[0081] In Figure 2 , the execution device 210 configures an input / output (I / O) interface 212 for data exchange with external devices. The user can input data to the I / O interface 212 through the client device 240. In this embodiment, the data can include an image to be classified, text, speech, etc. input by the user.

[0082] The preprocessing modules 213 and 214 are used to preprocess the input data (such as the image input by the user) received by the I / O interface 212.

[0083] The computing module 211 is used to perform calculations and other related processing on the data input from the preprocessing modules 213 and 214 according to the above-mentioned target model / rule 201.

[0084] When the execution device 210 preprocesses the input data, or when the computing module 211 of the execution device 210 performs calculations and other related processing, the execution device 210 can call the data, code, etc. of the database storage system 250 for corresponding processing, or can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 250.

[0085] Finally, the I / O interface 212 returns the processing result (such as a classification result or an identification result, etc.) to the client device 240 and provides it to the user. It should be understood that corresponding to different classification tasks, the target model / rule 201 is different, and its processing result is also different accordingly.

[0086] It is worth noting that the training device 220 can generate the corresponding target model / rule 201 for the downstream system. The corresponding target model / rule 201 can achieve the above-mentioned goal or complete the above-mentioned task, so as to provide the required result for the user. It should be noted that the training device 220 can also generate the corresponding preprocessing model for the target model / rule 201 corresponding to different downstream systems, such as the corresponding preprocessing models in the preprocessing modules 213 and / or 214, etc.

[0087] In Figure 2In the situation shown, the user can manually specify data in the input execution device 210 (for example, input an image or a piece of text), for example, operate in the interface provided by the I / O interface 212. In another situation, the client device 240 can automatically input data (for example, input an image or a piece of text) to the I / O interface 212 and obtain the result. If the client device 240 needs user authorization to automatically input data, the user can set the corresponding permissions in the client device 240. The user can view the result output by the execution device 210 in the client device 240 (for example, the output result can be a classification result or an identification result, etc.), and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 240 can also be used as a data collection end to collect, such as Figure 2 the input data (the image, text, or voice to be classified, etc.) input to the I / O interface 212 and the matching result output by the target model / rule 201 as new training sample data, and store it in the database 230.

[0088] It should be noted that Figure 2 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 2 , the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210. For another example, Figure 2 the shown system architecture may include more or fewer modules or components. For example, in some other examples, the system architecture may not include the preprocessing module 213 and the preprocessing module 214. After receiving the input data through the I / O, the execution device directly inputs the input data into the computing module for model inference without preprocessing the input data.

[0089] As Figure 2 shown, according to the target model / rule 201 trained by the training device 220, the target model / rule 201 can be the classification model in the embodiment of the present application. Specifically, the classification model provided by the embodiment of the present application is a neural network, such as a convolutional neural network (CNN), a deep convolutional neural network (DCNN), etc.

[0090] In related technologies, when training a model with semi-supervised learning on class-imbalanced training data, there are often various problems, resulting in poor classification accuracy of the classification model obtained through semi-supervised learning. For example, in one related technology, an annotated dataset is augmented by iteratively using a basic semi-supervised learning model, and the pseudo-label samples generated from the unannotated dataset are added to the annotated dataset, where the pseudo-label samples of the minority classes are more frequently selected according to an estimated class distribution. This solution also proposes a progressive distribution alignment for adaptively adjusting the rebalancing intensity.

[0091] However, this solution relies on labeled data for distribution alignment, and has limited effectiveness for extremely imbalanced or opposite distribution cases. Additionally, there is no independent pseudo-label generation module, resulting in biased pseudo-label generation.

[0092] Related technology two proposes a novel co-learning framework that decouples representation learning and classifier learning, and tightly couples them through a shared encoder and pseudo-label generation.

[0093] The auxiliary classifier and the original classifier in this solution adopt the same structure and training method, without dealing with unlabeled data. At the same time, when training the auxiliary classifier, it is assumed that the data classes are evenly distributed, resulting in poor generalization ability for location test data in the real scenario.

[0094] In summary, the semi-supervised learning of the classification models in related technologies cannot solve the problems of class imbalance and label scarcity, and cannot achieve high-performance classification in the case of class imbalance.

[0095] In view of this, the embodiments of the present application propose a training method for a classification model, introducing K PU classifiers for samples of K classes in the training dataset, generating pseudo-labels for unlabeled samples through the PU classifiers, effectively alleviating the problem of biased pseudo-labels, thereby improving the annotation quality of the pseudo-labels and enhancing the classification accuracy of the trained classification model. Since this method affects both the training stage and the inference stage, and the implementation processes of the training stage and the inference stage are different, the specific implementation processes of the foregoing two stages will be described separately below.

[0096] I. Training stage

[0097] In the embodiments of the present application, the training stage refers to the process in which the training device 220 described above Figure 2 performs a training operation on the target model / rule 201. Hereinafter, only taking the sample to be processed as an image sample as an example, the training method of the classification model provided by the embodiments of the present application will be introduced. It should be understood that when the sample to be processed is a sample in other formats, such as a text sample or a sound sample, etc., it can be analogously applied, and will not be elaborated here.

[0098] Figure 3Shows an implementation architecture diagram of a training method for a classification model provided by an embodiment of the present application. As Figure 3 shown, the implementation architecture of the training method for the classification model provided by the embodiment of the present application mainly includes three parts, a semi-supervised learning module (Semi-Supervised learning module), a pseudo-label generation module (pseudo-label generation module), a PU classifier training module (PU classifier learning module), and a labeled set expansion module (Labeled set expansion). First, the semi-supervised learning module uses the labeled sample set D @ and the unlabeled sample set D A to train the feature extraction network so that it learns a high-quality representation; then the PU classifier training module uses the PU learning algorithm on the labeled sample set D @ and the unlabeled sample set D A to train and obtain a PU classification model; the pseudo-label generation module uses the trained PU classification model to generate pseudo-labels for the unlabeled samples to obtain a pseudo-label sample set. Finally, the semi-supervised learning module trains and obtains a classification model with good classification performance based on the labeled sample set and the pseudo-label sample set. After the current training iteration is completed, the labeled set expansion module selects pseudo-labels with relatively high reliability to expand the labeled sample set and increase the number of samples in the labeled sample set, which is beneficial to the subsequent training of the classification model.

[0099] The training architecture provided by the embodiment of the present application can be abbreviated as PU-SSL.

[0100] Figure 4 Is a flowchart of a training method for a classification model provided by an embodiment of the present application. This method can be implemented in Figure 2 the training device 220 to train and obtain a highly robust classification model. As Figure 4 shown, the training method for the classification model provided by the embodiment of the present application includes at least steps S401 to S405.

[0101] In step S401, obtain the target training data set for the current training cycle.

[0102] In AI model training, the training data set is usually divided into multiple batches (batch), and then the training data is fed to the AI model to be trained batch by batch, and the parameters of the AI model are updated according to gradient descent.

[0103] In the embodiments of the present application, when the training device executes the semi-supervised training task, it will use a sampler to sample multiple batches of training data from the database. For example, the full amount of data corresponding to the classification model includes 10,000 image samples. 2,000 pictures are sampled in each batch, and the 2,000 sampled pictures are used as a batch to feed the classification model to be trained for semi-supervised learning.

[0104] The full amount of training data set stored in the database includes a labeled sample set D C and an unlabeled sample set D D , the labeled sample set D C includes multiple labeled samples, and the unlabeled sample set D D includes multiple unlabeled samples. Exemplarily, the labeled sample set D C includes a small number of labeled samples. The small number of labeled samples can be those for which it is difficult or expensive to obtain labeled data in the real world scenario. For example, medical data and biochemical experiment data often have limited labeled samples and a long acquisition process; of course, the small number of labeled samples can also be a small number of labeled samples that are manually labeled to reduce the workload of manual labeling; the unlabeled sample set D D includes a large number of unlabeled samples.

[0105] In order to train a classifier that can recognize multiple categories, the training data set needs to include samples of multiple categories. However, due to the existence of a large number of unlabeled samples, it is difficult to ensure that the sample quantity distributions of multiple categories are balanced. Therefore, during the actual training process of the classification model, it often needs to learn and predict classification on a training data set with unbalanced data distribution.

[0106] For example, as Figure 3 shown, the full amount of training data set corresponding to the classification model includes samples of multiple categories (for example, K categories), see Figure 3 , Figure 3 in the Labeled Set:D C The bar chart shows the quantity distribution of samples of K categories in the labeled sample set D C , and the Unlabeled Set:D E The bar chart shows the quantity distribution of samples of K categories in the unlabeled sample set D A . It can be seen that the quantity distributions of samples of K categories in the labeled sample set D C and the unlabeled sample set D A are not balanced. The quantity distributions of samples of some categories are large, while those of some categories are small.

[0107] The training device respectively samples from the labeled sample set D through the sampler CSample multiple labeled samples to obtain a subset of labeled samples, and sample from the unlabeled sample set D D to obtain a subset of unlabeled samples. The subset of labeled samples and the subset of unlabeled samples constitute the target training dataset for the current training cycle.

[0108] It should be noted that in the embodiments of this application, the training iteration process of the classification model using a batch of training data is called a training cycle, and the training iteration of the classification model using another batch of training data is called another training cycle. The target training dataset for the current training cycle can be understood as the current batch of training dataset collected by the sampler from the full training dataset. After the training device uses the current batch of training dataset to perform iterative update of the classification model, it is called the end of one training cycle.

[0109] An appropriate sampling algorithm can be configured in the sampler to sample the current batch of training dataset from the full training dataset. For example, a random sampling algorithm can be configured in the sampler. At this time, the sampler can be called a random sampler (Random Sampler). The random sampler randomly samples multiple labeled samples from the labeled sample set D C to obtain a subset of labeled samples, and randomly samples multiple unlabeled samples from the unlabeled sample set D D to obtain a subset of unlabeled samples.

[0110] In one example, after obtaining the current batch of training dataset, for example, it includes a subset of labeled samples (such as Figure 3 the l batch in) and a subset of unlabeled samples (such as Figure 3 the u batch in), use the current batch of subset of labeled samples and subset of unlabeled samples to preliminarily train the classification model, and update the parameters (such as weight parameters) of the backbone network, so that the backbone network can learn the feature representation of the samples. The backbone network is used to extract features from the input samples, so the backbone network can also be called a feature extraction network.

[0111] In one example, after the classification model is iteratively updated through the current batch of training dataset for preliminary training, as Figure 3 shown, perform exponential moving average (EMA) on the updated parameters, and then save the parameters of the feature extraction network. In this way, on the one hand, it speeds up the training speed of the feature extraction model, and on the other hand, it increases the robustness of the feature extraction network. The feature extraction network after EMA processing can be called the first feature extraction network.

[0112] In step S402, the unlabeled samples are respectively input into K PU classification models to obtain the output results of the K PU classification models.

[0113] After the classification model is preliminarily trained, a first feature extraction network is obtained. The parameters of the first feature extraction network are fixed, and then K PU classifiers are respectively mounted after the feature extraction network to form K PU classification models. Based on the current batch of training data sets, the K PU classification models are respectively trained to update each PU classifier, and each PU classification model for each category is obtained.

[0114] For example, the current batch of training data sets includes samples of K categories, such as samples of the first category, samples of the second category... samples of the k-th category. The K PU classification models include PU classification models of K categories, such as the PU classification model of the first category, the PU classification model of the second category... the PU classification model of the k-th category. The PU classification model of the first category is used to identify whether the input sample is of the first category, the PU classification model of the second category is used to identify whether the input sample is of the second category, and so on. The PU classification model of the k-th category is used to identify whether the input sample is of the k-th category.

[0115] In one example, the K categories in the current batch of training data sets can be determined by the sample categories distributed in the current batch of labeled sample sets. For example, by checking that the labeled sample set includes samples of K categories, it is determined that the current batch of training data sets also includes samples of K categories.

[0116] The specific implementation of the training of the PU classification model for any one category (such as the first category) among the PU classification models of K categories is as follows: First, select the samples of the first category from the labeled sample set in the current batch of training data sets as the positive sample set. The positive sample set and the unlabeled sample set in the current batch constitute the PU training data set of the PU classification model of the first category.

[0117] Any suitable PU learning algorithm can be used to perform PU learning on the PU training data set to obtain a PU classifier. For example, the samples in the PU training data set are used as the input of the PU classification model of the first category. The first feature extraction network extracts the features of the input samples to obtain the feature representation of the samples, and then the feature representation is used as the input of the PU classifier of the first category to output the probability that the input sample is a positive sample, or the probability that the input sample is of the first category. Furthermore, the category estimation of the input sample is obtained. For example, when the probability output by the PU classifier is 0.9, the category estimation of the input sample is the first category. The PU loss of the PU loss function is calculated according to the category estimation of the sample and the input sample, and then the parameters of the PU classifier are updated according to the PU loss.

[0118] Exemplarily, the PU loss is a form of non - negative risk estimation that separately estimates the expected losses of positive samples and unlabeled samples. For each class k among the K classes, a PU classifier g k is introduced, and the PU loss is defined as follows:

[0119]

[0120] where π k is determined by the prior distribution of the k - th class in the labeled sample set of the current batch, represents the number of samples of the k - th class in the labeled sample set, n u represents the total number of samples in the unlabeled sample subset, and the l(.,.) function is an arbitrary surrogate 0 - 1 loss function, such as the sigmoid loss function:

[0121] The first term (i.e., the term before the plus sign) in the PU loss function represents the positive risk estimation, which is used to penalize the PU classifier for assigning a low probability to positive samples. The second term (i.e., the term after the plus sign) is used to represent the non - negative risk estimation, which is used to penalize the PU classifier for assigning a high probability to unlabeled samples. At the same time, the max function ensures that the empirical risk estimation for unlabeled samples is greater than or equal to 0, thereby reducing the risk of model overfitting.

[0122] It can be understood that the meaning of "penalize" mentioned above is that when the PU classifier assigns a low probability (e.g., 0.3) to positive samples, the positive risk estimation term will increase, and / or when the PU classifier assigns a high probability (e.g., 0.7) to negative samples, the negative risk estimation term will increase, thereby causing the loss function to increase.

[0123] In an embodiment of the present application, when training the PU classifier G, the parameters of the first feature extraction network are kept fixed, and only the parameters (such as weight parameters) of the PU classifier G are updated. In this way, it is ensured that the PU classifier G can adapt to the latest representation learned by the feature extraction network, so as to use this representation to generate higher - quality pseudo - labels for unlabeled samples.

[0124] In one example, after each PU classifier G among the K PU classifiers G is trained and iteratively updated through the training data set of the current batch, the updated parameters are processed by EMA, and then the parameters of the PU classifier G are saved. On the one hand, this speeds up the training speed of the PU classifier G, and on the other hand, it increases the robustness of the PU classifier G.

[0125] Optionally, the K PU classifiers can be trained for each category using the unlabeled sample set and the labeled sample set of the current batch, so as to ensure that the learning of the PU classifier is not affected by the original data distribution, so that the learned PU classifier is fair, and further ensure the reliability of the pseudo-labels generated by the PU classifier.

[0126] After training K classification models using the training data set of the current batch, the unlabeled samples in the unlabeled data set are respectively input into the K PU classification models, and the K PU classification models output K results.

[0127] For example, referring to Figure 3 the pseudo-label generation module part in, the unlabeled samples in the unlabeled data set (u batch) are input into the first feature extraction network, and the first feature extraction network extracts features from the input samples to obtain the feature representation z of the unlabeled samples u , and then the feature representation z of the unlabeled samples u are respectively input into K PU classifiers G, and each PU classifier in the K PU classifiers outputs the estimated probability that the input unlabeled sample is a positive sample, obtaining K estimated probabilities.

[0128] In another example, before inputting the unlabeled data set into the K PU classification models to generate pseudo-labels, the unlabeled data set is subjected to data augmentation processing, and then the unlabeled samples in the unlabeled sample set after data augmentation processing are input into the K PU classification models to obtain the pseudo-labels of the unlabeled samples. To increase the number of samples in the unlabeled data set, which is beneficial to the training of subsequent classification models. For example:

[0129]

[0130] Among them, represents the pseudo-label of the unlabeled sample generated by the K PU classification models, α represents a data augmentation function, and g k represents the PU classifier corresponding to the k-th class.

[0131] Optionally, the method for data augmentation of the unlabeled samples in the unlabeled data set can be one or more of flipping, translation, rotation, and cropping of the unlabeled sample image data samples.

[0132] In step S403, based on the output results of the K PU classification models, the pseudo-labels of the unlabeled samples are obtained.

[0133] After inputting the feature representation of the unlabeled sample into the K PU classifiers, the output results of the K PU classifiers are obtained. For example, the feature representation z of the unlabeled sample uAfter inputting K PU classifiers respectively, the probability that the first-class classifier G outputs the unlabeled sample as the first class is 0.1, the probability that the second-class classifier G outputs the unlabeled sample as the second class is 0.3, the probability that the third-class classifier G outputs the unlabeled sample as the third class is 0.4, the probability that the fourth-class classifier G outputs the unlabeled sample as the fourth class is 0.7…, and the probability that the k-th class classifier G outputs the unlabeled sample as the k-th class is 0.9. The class corresponding to the PU classifier with the largest estimated probability (or score) in the output results is used as the pseudo-label of the unlabeled sample. For example, the estimated probability of the k-th class is 0.9, which is the highest among the K estimated probabilities output by the K PU classifiers. The class corresponding to the k-th PU classifier, that is, the k-th class, is used as the pseudo-label of the unlabeled sample, and then the unlabeled sample is labeled with the pseudo-label, that is, the pseudo-label of the unlabeled sample is the k-th class. Repeat the above process to determine the pseudo-labels of each unlabeled sample and label each unlabeled sample with a pseudo-label.

[0134] In step S404, the classification model is trained based on the labeled sample subset and the pseudo-labeled sample subset, and the parameters of the feature extraction network and the classifier are updated.

[0135] By the above steps, the unlabeled samples in the unlabeled sample set of the current batch are labeled with pseudo-labels to obtain a pseudo-labeled sample set, and then the classification model is trained based on the labeled sample set and the pseudo-labeled sample set of the current batch, and the parameters of the first feature extraction network and the classifier are updated according to the loss function.

[0136] Optionally, before training the classification model, the pseudo-labeled sample set is subjected to data augmentation processing to expand the number of samples in the pseudo-labeled sample set. For example, one or more of flipping, translation, rotation, and cropping are performed on the pseudo-labeled samples for image data sample augmentation processing.

[0137] In another example, the data augmentation methods used for the data augmentation processing of the unlabeled sample set and the pseudo-labeled sample set are different. For example, if the flipping data augmentation method is used for the data augmentation processing of the unlabeled sample set, the rotation data augmentation method is used for the data augmentation processing of the pseudo-labeled sample set. Or, for the data augmentation processing of the unlabeled sample set and the pseudo-labeled sample set, a certain data augmentation method randomly selected from a variety of data augmentation methods is adopted for data augmentation processing; for example, the translation data augmentation method is randomly used for the data augmentation processing of the unlabeled sample set, and the cropping data augmentation method is randomly used for the data augmentation processing of the pseudo-labeled sample set.

[0138] Input the samples in the labeled sample set and the pseudo-labeled sample set after data augmentation processing into the classification model to be trained. The first feature extraction network in the classification model extracts features from the input samples to obtain the feature representations of the samples, and then inputs the feature representations into the classifier. The classifier outputs the class estimation of the input samples. Based on the class estimation and the labels (true labels or pseudo-labels) of the input samples, determine the loss of the loss function, and update the parameters of the first feature extraction network and the parameters of the classifier in the classification model by minimizing the loss.

[0139] In one example, a loss function designed for the classification model to be trained is as follows:

[0140]

[0141] where L is the standard cross-entropy loss function, α(.) is the data augmentation function, and h(.) is the original classifier (i.e., the classifier in the classification model to be trained). is the feature extraction network used to extract features. represents the labeled sample, represents the sample corresponding true label. represents the unlabeled sample, represents the sample pseudo-label generated by the PU classifier.

[0142] By minimizing the loss value of the above loss function, update the parameters (such as weight parameters) of the first feature extraction network and the parameters (such as weight parameters) of the classifier in the classification model to be trained.

[0143] In step S405, after training for multiple training cycles, obtain the trained classification model.

[0144] Repeat the above steps S401 to S404, and after training for multiple training cycles, when the training completion condition is met, obtain the trained classification model to achieve high-precision classification recognition of the classification model.

[0145] Optionally, the training completion condition can be that the loss function converges or reaches a preset number of training cycles, etc. The training completion condition can be set according to actual needs, and this application does not limit it.

[0146] The training method of the classification model provided by the embodiments of the present application respectively introduces a dedicated PU classifier for samples of each category in the training dataset, and uses the PU classifier to generate pseudo-labels for unlabeled samples, rather than using the classifier in the classification model to be trained to generate pseudo-labels, which improves the quality of pseudo-label generation. At the same time, during the training process, the training of the PU classifier and the training of the feature extraction network in the classification model to be trained are decoupled, and then integrated through EMA. In each training iteration, first train and update the parameters of the feature extraction network in the classification model to be trained, then fix the parameters of the feature extraction network, extract the feature representation of the samples through the feature extraction network, and then input the feature representation into the PU classifier to train and update the parameters of the PU classifier, so that the PU classifier adapts to the latest representation learned by the feature extraction network, thereby using this representation to generate high-quality pseudo-labels for labeled samples.

[0147] In another example, after a training cycle is completed, it further includes using the pseudo-label dataset to augment the labeled sample set. For example, select high-confidence pseudo-labels from the pseudo-label sample subset as reliable pseudo-labels to obtain a reliable pseudo-label sample subset; based on the reliable pseudo-label sample subset, augment the labeled sample subset to increase the number of samples in the labeled sample subset, thereby improving the training quality of the classification model.

[0148] For example, select the pseudo-labels with the output results of the PU classifier higher than the preset threshold τ (such as 0.9) as reliable pseudo-labels to form a pseudo-label list for each category k, and then obtain a reliable pseudo-label sample set. Through the reliable pseudo-label sample set, augment the labeled sample subset to increase the number of samples in the labeled sample subset, thereby improving the training quality of the classification model.

[0149] In another example, in order to balance the class distribution in the labeled dataset, a sampling method is designed to sample from the reliable pseudo-label sample set, and the sampled pseudo-label samples are supplemented into the labeled sample set for augmentation. This sampling method can be, for example: first determine the number of samples in the category with the largest number of samples in the labeled sample set Then count the number of samples of each category in the labeled sample set The sampling quantity of each category Then sample T k reliable pseudo-label samples of each category from the reliable pseudo-label sample set.

[0150] For example, there are a total of 500 samples in the labeled sample set of the current batch, including 100 samples of class 1, 200 samples of class 2, 50 samples of class 3, and 150 samples of class 4. There are a total of 800 samples in the reliable pseudo-label sample set, including 150 samples of class 1, 500 samples of class 2, 100 samples of class 3, and 50 samples of class 4. Then the sampling quantity T of the samples of class 1 + is 200 - 100 = 100, and the sampling quantity T of the samples of class 2 J is 200 - 200 = 0, and the sampling quantity T of the samples of class 3 O is 200 - 50 = 150, and the sampling quantity T of the samples of class 4 P is 200 - 150 = 50; then samples are sampled from the reliable pseudo-label sample set according to the sampling quantity of each sample. For example, 100 pseudo-label samples are sampled from the samples of class 1 in the reliable pseudo-label sample set; the sampling quantity T of the samples of class 2 J = 0, so there is no need to sample the samples of class 2; since the number of samples of class 3 in the reliable pseudo-label sample set is 100, which is less than the sampling quantity 150 of the samples of class 3, all 100 samples of class 3 in the reliable pseudo-label sample set are sampled, and 50 pseudo-label samples are collected from the samples of class 4 in the reliable pseudo-label sample set. In this way, without using complex class rebalancing rules, it is possible to balance the data distribution of the labeled sample set during the data augmentation process.

[0151] In another example, the sampling quantity of the samples of each category from the reliable pseudo-label sample set can also be determined according to the following sampling formula: where, represents the number of samples of the k-th category in the labeled dataset, represents the number of samples of the category with the largest number of samples in the labeled sample set, ρ belongs to (0, 1) and is the sampling ratio, which is a hyperparameter set by humans and is used to control the trade-off between class balance and pseudo-label quality. In this way, the class imbalance problem of the training dataset is solved, and the training quality of the classification model is improved.

[0152] The embodiment of the present application effectively alleviates the problem of data distribution imbalance in the training sample set by selecting reliable pseudo-label samples to supplement the labeled sample set for the next stage of semi-supervised learning, and further improves the training quality of the trained classification model.

[0153] The code implementation of the training method of the classification model provided by the embodiment of the present application is as follows:

[0154]

[0155]

[0156] It can be understood that the above only gives a feasible specific implementation of a training method for a classification model provided in the embodiments of the present application, and there may be other specific implementations. For example, the training method for the classification model provided in the embodiments of the present application may also not adopt the method of batch training, but directly use the full training dataset to train the classification model. One training iteration update of the classification model through the full training dataset is called a training cycle. For another example, multiple PU classification models can also be pre-trained in the full training dataset to obtain multiple PU classification models. That is to say, the PU classification models do not participate in the training together with the classification model; when pseudo-labels need to be labeled for unlabeled samples during the training process of the classification model, the samples in the unlabeled sample set are directly input into the multiple PU classification models to generate pseudo-labels for the unlabeled samples.

[0157] II. Inference stage

[0158] In the embodiments of the present application, the inference stage refers to the process in which the execution device 210 in the above Figure 2 performs a classification operation using the trained target model / rule 201. And taking only the sample to be processed as an image sample as an example, the classification method provided in the embodiments of the present application is introduced. It should be understood that when the sample to be processed is a sample in other formats, such as a text sample or a sound sample, it can be analogously applied and will not be elaborated here. Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a classification method provided in the embodiments of the present application. The method at least includes steps S501 to S503.

[0159] In step S501, an image to be classified is obtained.

[0160] In the embodiments of the present application, the execution device will first obtain the data to be processed, such as the image data to be classified, during one inference process. Specifically, the execution device can obtain the image to be classified by receiving the image to be classified sent by other communication devices (such as client devices); it can also select the image to be classified for the user from the image library stored in the execution device; or there are N types of images pre-stored on the execution device, and the image to be classified is collected from the pre-stored N types of images. The embodiments of the present application do not make specific limitations on the specific acquisition method of the image to be classified.

[0161] In step S502, the image to be classified is used as the input of the feature extraction network to obtain the feature representation of the image to be classified.

[0162] The execution device uses the image to be classified as the input of the classification model. The classification model first extracts features from the image to be classified through the feature extraction network to obtain the feature representation of the image to be classified.

[0163] In step S503, the feature representation of the image to be classified is input into the classifier to obtain the estimated category of the image to be classified.

[0164] In the classification model, the feature extraction network is followed by a classifier. The feature representation output by the feature extraction network serves as the input to the classifier. The classifier outputs the probability distribution of the image to be classified for each category based on the feature representation of the image to be classified. The category of the image to be classified can be determined according to the probability distribution. For example, the classifier outputs the probability distribution of the image to be classified as follows: the probability of the image to be classified belonging to the first category is 0.9, the probability of the second category is 0.1, the probability of the third category is 0.2,..., the probability of the k-th category is 0.1. Since the probability of the image to be classified belonging to the first category is the highest, the classification model outputs the category of the image to be classified as the first category.

[0165] The feature extraction network and the classifier in the above-mentioned classification model are trained by the training method of the classification model described in the embodiments of the present application. The specific training process is as described above and will not be elaborated here.

[0166] To verify the training effect of the training method of the classification model provided by the embodiments of the present application, tests are carried out on multiple data sets. The multiple data sets include, for example, CIFAR-10 / 100LT, Small-ImageNET-127, and Food-101-LT.

[0167]

[0168]

[0169] Table 1

[0170]

[0171] Table 2

[0172] Table 1 and Table 2 respectively show the training effect tests of the training method of the classification model provided by the embodiments of the present application (i.e., PU-SSL) and other semi-supervised training algorithms on different data sets. It can be seen that the training method of the classification model provided by the embodiments of the present application can achieve the best results compared with other semi-supervised algorithms on different data sets and class imbalance ratios.

[0173]

[0174] Table 3

[0175] Table III shows the performance of the training method of the classification model provided in the embodiments of the present application and other semi-supervised training algorithms on datasets with various imbalanced distributions. It can be seen that the training method of the classification model in the embodiments of the present application is robust to datasets with various imbalanced distributions. Regardless of the training distribution used, the PU-SSL in the embodiments of the present application can always maintain a more balanced performance and a higher average precision under various test imbalance ratios, which demonstrates the robustness of the training method of the classification model provided in the embodiments of the present application to different training data and test data distributions.

[0176] Based on the same concept as the embodiment of the training method of a classification model described above, an embodiment of the present application further provides a training device 600 for a classification model. The training device 600 for the classification model can be deployed on any device, equipment, platform or device cluster with computing capabilities to execute, implement the training method of the classification model provided in the embodiments of the present application, so as to achieve high robustness and high classification accuracy of the semi-supervised training algorithm on a dataset with an imbalanced data distribution. The training device 600 for the classification model includes units or modules for implementing Figure 3 and 4 each step in the training method of the classification model shown.

[0177] Figure 6 FIG. is a schematic structural diagram of a training device for a classification model provided in an embodiment of the present application. As Figure 6 shown, the training device 600 for the classification model at least includes an acquisition module 601, a PU classification module 602, a pseudo-label annotation module 603, and a training module 604. Among them, the acquisition module 601 is used to acquire a target training dataset for the current training cycle. The target training dataset includes a labeled sample subset and an unlabeled sample subset. The labeled sample subset includes a plurality of labeled samples, and the unlabeled sample subset includes a plurality of unlabeled samples. The target training dataset includes samples of K categories, and K is a positive integer greater than 1; the PU classification module 602 is used to input the unlabeled samples into K PU classification models respectively, and obtain the output results of the K PU classification models. Each PU classification model among the K PU classification models is respectively used to estimate the categories of samples in each of the K categories; the pseudo-label annotation module 603 is used to obtain pseudo-labels of the unlabeled samples based on the output results of the K PU classification models; the training module 604 is used to train the classification model based on the labeled sample subset and the pseudo-label sample subset, and update the parameters of the feature extraction network and the parameters of the classifier. The pseudo-label sample subset includes a plurality of samples with pseudo-labels; after training for multiple training cycles, a trained classification model is obtained.

[0178] In a possible implementation, the training module 604 is further used to: train the classification model based on the target training dataset and update the parameters of the feature extraction network.

[0179] In another possible implementation, each of the K PU classification models includes an updated feature extraction network and a PU classifier; the training device 600 for the classification model provided in this application further includes a PU learning module 605, and this PU learning module is used to train each of the K PU classification models based on the target training dataset and update the PU classifiers of each of the PU classification models.

[0180] In another possible implementation, the i-th PU classification model is any one of the K PU classification models. The i-th PU classification model includes an updated feature extraction network and the i-th PU classifier; the specific implementation of the PU learning module for training the i-th PU classification model is: determining the i-th PU training dataset corresponding to the i-th PU classification model from the target training dataset, where the i-th PU training dataset includes a positive sample subset and an unlabeled sample subset, the positive sample subset includes a plurality of positive samples, and the positive samples are samples of the target category in the labeled sample subset, and the target category is the category corresponding to the i-th PU classification model; training the i-th PU classification model based on the i-th PU training dataset and updating the parameters of the i-th PU classifier.

[0181] In another possible implementation, a specific implementation of training the i-th PU classification model based on the i-th PU training dataset and updating the parameters of the i-th PU classifier is: inputting the samples in the i-th PU training dataset into the i-th PU classification model to obtain the class estimation of each sample in the i-th PU training dataset by the i-th PU classification model; updating the parameters of the i-th PU classifier based on the PU loss function, and the PU loss function includes a positive risk estimation term and a negative risk estimation term. The positive risk estimation term is used to punish the i-th PU classifier for estimating positive samples as negative samples, and the negative risk estimation term is used to punish the i-th PU classifier for estimating unlabeled samples as positive samples.

[0182] In another possible implementation, the PU loss function is also related to the proportion of samples of the target category in the labeled sample subset.

[0183] In another possible implementation, the training device 600 for the classification model provided in this application further includes a first data augmentation module 606, and this first data augmentation module is used to perform first data augmentation processing on the unlabeled sample subset.

[0184] In another possible implementation, the output result of each PU classification model includes an estimated probability, which indicates the probability that the input sample of each PU classification model is a positive sample; specifically, the pseudo-label annotation module 603 is configured to: based on the output results of the K PU classification models, determine a target PU classification model, where the estimated probability output by the target PU classification model is the largest; and label the class corresponding to the target PU classification model as the pseudo-label of the unlabeled sample.

[0185] In another possible implementation, specifically, the training module 604 is configured to: input the labeled samples into the classification model, and output the class estimation corresponding to the labeled samples; based on the labels of the labeled samples and the class estimations of the labeled samples, obtain the loss value of the first loss function; input the samples with pseudo-labels into the classification model, and output the class estimation corresponding to the samples with pseudo-labels; based on the pseudo-labels of the samples with pseudo-labels and the class estimations of the samples with pseudo-labels, obtain the loss value of the second loss function; and based on the loss value of the first loss function and the loss value of the second loss function, update the parameters of the feature extraction network and the parameters of the classifier.

[0186] In another possible implementation, the classification model training device 600 provided in this application further includes a second data augmentation module 607, which is configured to perform second data augmentation processing on the pseudo-label sample subset.

[0187] In another possible implementation, the classification model training device 600 provided in this application further includes a data augmentation module 608, which is configured to, after the current training cycle is completed, determine samples with reliable pseudo-labels from the pseudo-label sample subset to obtain a reliable pseudo-label sample subset, where the output results of the PU classification models corresponding to the samples with reliable pseudo-labels are greater than or equal to a preset threshold; and based on the reliable pseudo-label sample subset, perform data augmentation on the labeled sample subset.

[0188] In another possible implementation, a specific implementation of performing data augmentation on the labeled sample subset based on the reliable pseudo-label sample subset is to determine the number of samples of the first class in the labeled sample subset, where the number of samples of the first class is the largest; based on the difference between the number of samples of each class in the labeled sample subset and the number of samples of the first class, determine the sampling number of samples of each class; based on the sampling number of samples of each class, perform sample sampling on samples of each class in the reliable pseudo-label sample subset to obtain an augmented data set; and supplement the augmented data set to the labeled sample subset to obtain an augmented labeled data subset.

[0189] In another possible implementation, the full training dataset corresponding to the classification model includes a labeled sample set and an unlabeled sample set. The labeled sample set includes multiple labeled samples, and the unlabeled sample set includes multiple unlabeled samples. The obtaining module 601 is specifically configured to use a random sampler to collect samples from the labeled sample set and the unlabeled sample set respectively, so as to obtain a current batch of labeled sample subset and unlabeled sample subset.

[0190] The training device 600 of the classification model according to the embodiment of the present application can correspond to executing the method described in the embodiment of the present application, and the above and other operations and / or functions of each module in the training device 600 of the classification model are respectively for realizing Figure 3 and 4 the corresponding processes of each method in, for the sake of brevity, will not be described in detail here.

[0191] The embodiment of the present application also provides a computing device, including at least one processor, a memory, and a communication interface. The processor is configured to execute Figure 3-5 the

[0192] Figure 7 method.

[0193] As Figure 7 shown, the computing device 700 includes at least one processor 701, a memory 702, and a communication interface 703. Among them, the processor 701, the memory 702, and the communication interface 703 are communicatively connected, and can be communicatively connected in a wired manner (such as a bus), or can be communicatively connected in a wireless manner. The communication interface 703 is used to send and / or receive data sent by other devices; the memory 702 stores computer instructions, and the processor 701 executes the computer instructions to execute the training method of the classification model in the foregoing method embodiment, so as to achieve high robustness and high classification accuracy of the semi-supervised training algorithm on a dataset with unbalanced data distribution.

[0194] It should be understood that in the embodiment of the present application, the processor 701 may be a central processing unit CPU, and the processor 701 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0195] The memory 702 may include a read-only memory and a random access memory, and provide instructions and data to the processor 701. The memory 702 may also include a non-volatile random access memory.

[0196] The memory 702 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0197] It should be understood that the computing device 700 according to the embodiments of the present application may execute the method implemented in the embodiments of the present application. Figure 3-5 For the detailed description of the implementation of the method shown above, please refer to the above text. For the sake of brevity, it will not be repeated here.

[0198] Embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer instructions are executed by a processor, the method mentioned above is implemented.

[0199] Embodiments of the present application provide a chip, which includes at least one processor and an interface. The at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the method mentioned above.

[0200] Embodiments of the present application provide a computer program or a computer program product, which includes instructions that, when executed, cause a computer to perform the methods mentioned above.

[0201] Those of ordinary skill in the art should further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0202] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0203] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A classification model training method, characterized in that: The classification model to be trained includes a feature extraction network and a classifier, wherein the feature extraction network is used to extract a feature representation of a first input sample, and the classifier is used to output an estimated category probability distribution of the first input sample based on the feature representation, wherein the first input sample is a sample input to the feature extraction network, and the method includes: Acquire a target training data set, wherein the target training data set includes a labeled sample subset and an unlabeled sample subset, wherein the labeled sample subset includes a plurality of labeled samples, and the unlabeled sample subset includes a plurality of unlabeled samples, and the target training data set includes samples of K categories, where K is a positive integer greater than 1; Inputting the unlabeled samples into K PU classification models respectively to obtain output results of the K PU classification models, wherein each PU classification model in the K PU classification models is used to perform category estimation on samples of each category in the K categories; Based on the output results of the K PU classification models, obtaining pseudo labels for the unlabeled samples; The classification model to be trained is trained based on the labeled sample subset and the pseudo-label sample subset, and the parameters of the feature extraction network and the parameters of the classifier are updated. The pseudo-label sample subset includes a plurality of samples with pseudo-labels.

2. The method according to claim 1, characterized in that: The step of inputting the unlabeled samples into K PU classification models respectively to obtain output results of the K PU classification models further includes: The classification model to be trained is trained based on the target training data set, and the parameters of the feature extraction network are updated.

3. The method according to claim 2, characterized in that Each of the K PU classification models comprises the updated feature extraction network and a PU classifier, wherein the PU classifier is used to output a probability that the second input sample is a positive sample based on a feature representation of the second input sample extracted by the updated feature extraction network, and the second input sample is a sample input to the updated feature extraction network; The step of inputting the unlabeled samples into K PU classification models respectively to obtain output results of the K PU classification models further includes: Each PU classification model in the K PU classification models is trained based on the target training data set, and the PU classifier of each PU classification model is updated.

4. The method according to claim 3, characterized in that The i-th PU classification model is any one of the K PU classification models, and the i-th PU classification model includes the updated feature extraction network and the i-th PU classifier; Training each PU classification model among the K PU classification models based on the target training data set, and updating the PU classifier of each PU classification model, including: Determine an i-th PU training data set corresponding to the i-th PU classification model from the target training data set, wherein the i-th PU training data set includes a positive sample subset and the unlabeled sample subset, the positive sample subset includes samples of the target category in the labeled sample subset, and the target category is the category corresponding to the i-th PU classification model; The i-th PU classification model is trained based on the i-th PU training data set, and parameters of the i-th PU classifier are updated.

5. The method according to claim 4, characterized in that The training of the i-th PU classification model based on the i-th PU training data set and updating the parameters of the i-th PU classifier include: Inputting the samples in the i-th PU training data set into the i-th PU classification model to obtain a category estimation of each sample in the i-th PU training data set by the i-th PU classification model; Based on the PU loss function, update the parameters of the i-th PU classifier, the PU loss function includes a positive risk estimation term and a negative risk estimation term, the positive risk estimation term is used to punish the i-th PU classifier for estimating the positive sample as a negative sample, and the negative risk estimation term is used to punish the i-th PU classifier for estimating the unlabeled sample as a positive sample.

6. The method according to claim 5, characterized in that The PU loss function is also related to the proportion of samples of the target category in the labeled sample subset.

7. The method according to any one of claims 1 to 6, characterized in that: The output result of each PU classification model includes an estimated probability, where the estimated probability indicates a probability that each PU classification model estimates that the third input sample is a positive sample, where the third input sample is a sample input to the PU classification model; The obtaining the pseudo labels of the unlabeled samples based on the output results of the K PU classification models includes: Determine a target PU classification model based on the output results of the K PU classification models, wherein the estimated probability output by the target PU classification model is the largest; The category corresponding to the target PU classification model is marked as a pseudo label of the unlabeled sample.

8. The method according to any one of claims 1 to 7, characterized in that: The step of training the classification model to be trained based on the labeled sample subset and the pseudo-label sample subset, and updating the parameters of the feature extraction network and the parameters of the classifier, comprises: Inputting the labeled samples into the classification model to be trained, and outputting the category estimation corresponding to the labeled samples; Obtaining a loss value of a first loss function based on the label of the labeled sample and the category estimation of the labeled sample; Inputting the samples with pseudo labels into the classification model to be trained, and outputting the category estimation corresponding to the samples with pseudo labels; Obtaining a loss value of a second loss function based on the pseudo-labels of the samples with the pseudo-labels and the category estimation of the samples with the pseudo-labels; Based on the loss value of the first loss function and the loss value of the second loss function, the parameters of the feature extraction network and the parameters of the classifier are updated.

9. The method according to any one of claims 1 to 8, characterized in that: Also includes: The classification model to be trained is trained through multiple training cycles to obtain a trained target classification model, wherein after each training cycle is completed, a sample with a reliable pseudo-label is determined from the pseudo-label sample subset to obtain a reliable pseudo-label sample subset, and an output result of the PU classification model corresponding to the sample with a reliable pseudo-label is greater than or equal to a preset threshold; Based on the reliable pseudo-label sample subset, data expansion is performed on the labeled sample subset.

10. The method according to claim 9, characterized in that The step of performing data expansion on the labeled sample subset based on the reliable pseudo-label sample subset includes: Determine the number of samples of the first category in the labeled sample subset, where the number of samples of the first category is the largest; Determine the sampling quantity of the samples of each category based on the difference between the quantity of the samples of each category in the labeled sample subset and the quantity of the samples of the first category; Based on the sampling number of samples of each category, sample samples of each category in the reliable pseudo-label sample subset are sampled to obtain an expanded data set; The expanded data set is added to the labeled sample subset to obtain a labeled data subset after data expansion.

11. The method according to any one of claims 1 to 10, characterized in that: The full training data set corresponding to the classification model to be trained includes a labeled sample set and an unlabeled sample set, wherein the labeled sample set includes a plurality of labeled samples, and the unlabeled sample set includes a plurality of unlabeled samples; The step of obtaining a target training data set includes: A random sampler is used to collect samples from the labeled sample set and the unlabeled sample set respectively to obtain a labeled sample subset and an unlabeled sample subset of the current batch.

12. A training device for a classification model, characterized in that: The classification model to be trained includes a feature extraction network and a classifier, wherein the feature extraction network is used to extract a feature representation of a first input sample, and the classifier is used to output an estimated category probability distribution of the first input sample based on the feature representation, and the first input sample is a sample input to the feature extraction network. The device includes: An acquisition module is used to acquire a target training data set, wherein the target training data set includes a labeled sample subset and an unlabeled sample subset, wherein the labeled sample subset includes a plurality of labeled samples, and the unlabeled sample subset includes a plurality of unlabeled samples, and the target training data set includes samples of K categories, wherein K is a positive integer greater than 1; A PU classification module, used for inputting the unlabeled samples into K PU classification models respectively to obtain output results of the K PU classification models, wherein each PU classification model of the K PU classification models is used for performing category estimation on samples of each category in the K categories respectively; A pseudo label marking module, used for obtaining a pseudo label of the unlabeled sample based on the output results of the K PU classification models; A training module is used to train the classification model to be trained based on the labeled sample subset and the pseudo-label sample subset, and update the parameters of the feature extraction network and the parameters of the classifier. The pseudo-label sample subset includes multiple samples with pseudo labels.

13. The device according to claim 12, characterized in that The training module is also used to: The classification model is trained based on the target training data set, and the parameters of the feature extraction network to be trained are updated.

14. The device according to claim 13, characterized in that Each of the K PU classification models comprises the updated feature extraction network and a PU classifier, wherein the PU classifier is used to output a probability that the second input sample is a positive sample based on a feature representation of the second input sample extracted by the updated feature extraction network, and the second input sample is a sample input to the updated feature extraction network; The device also includes: The PU learning module is used to train each PU classification model in the K PU classification models based on the target training data set, and update the PU classifier of each PU classification model.

15. The device according to claim 14, characterized in that The i-th PU classification model is any one of the K PU classification models, and the i-th PU classification model includes the updated feature extraction network and the i-th PU classifier; The PU learning module is specifically used for: Determine an i-th PU training data set corresponding to the i-th PU classification model from the target training data set, wherein the i-th PU training data set includes a positive sample subset and the unlabeled sample subset, the positive sample subset includes samples of a target category in the labeled sample subset, and the target category is a category corresponding to the i-th PU classification model; The i-th PU classification model is trained based on the i-th PU training data set, and parameters of the i-th PU classifier are updated.

16. The device according to claim 15, characterized in that The training of the i-th PU classification model based on the i-th PU training data set and updating the parameters of the i-th PU classifier include: Inputting the samples in the i-th PU training data set into the i-th PU classification model to obtain a category estimation of each sample in the i-th PU training data set by the i-th PU classification model; Based on the PU loss function, update the parameters of the i-th PU classifier, the PU loss function includes a positive risk estimation term and a negative risk estimation term, the positive risk estimation term is used to punish the i-th PU classifier for estimating the positive sample as a negative sample, and the negative risk estimation term is used to punish the i-th PU classifier for estimating the unlabeled sample as a positive sample.

17. The device according to any one of claims 12 to 16, characterized in that: The output result of each PU classification model includes an estimated probability, where the estimated probability indicates a probability that each PU classification model estimates that the third input sample is a positive sample, where the third input sample is a sample input to the PU classification model; The pseudo-label marking module is specifically used for: Determine a target PU classification model based on the output results of the K PU classification models, wherein the estimated probability output by the target PU classification model is the largest; The category corresponding to the target PU classification model is marked as a pseudo label of the unlabeled sample.

18. The device according to any one of claims 12 to 17, characterized in that: Also includes: A data expansion module, configured to determine, after each training cycle is completed, samples of reliable pseudo-labels from the pseudo-label sample subset to obtain a reliable pseudo-label sample subset, wherein an output result of a PU classification model corresponding to the samples of the reliable pseudo-labels is greater than or equal to a preset threshold; Based on the reliable pseudo-label sample subset, data expansion is performed on the labeled sample subset.

19. A computing device comprising a memory and a processor, characterized in that: Instructions are stored in the memory, and when the instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.