Model training method and homogeneous population screening method and device
Through collaborative guidance training and semi-supervised learning, the models in noise learning are trained, which solves the problems of low sample data utilization and poor model effectiveness, and achieves more efficient model training and higher accuracy.
Patent Information
- Application Number
- CN202510088122.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing noise-free learning methods have low sample data utilization in model training and poor model training, especially in the absence of high-accuracy labeling samples.
The initial model is trained using a collaborative guidance training method, and label data without using sample data is generated, and the model is further trained through semi-supervised learning idea to improve the utilization rate of sample data and the accuracy of the model.
This improves the sample data utilization rate in model training and increases the training data scale, thereby improving the model training effect and accuracy.
Smart Images

Figure CN120014685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a model training method and a homogeneous population screening method and device. Background Art
[0002] Deep learning models based on artificial neural networks are widely used in various task scenarios. Deep learning models often require large-scale, high-accuracy labeled sample data for training to achieve good performance. However, in real-world scenarios, the labeling cost of sample data is high, and it is difficult to obtain large-scale, high-accuracy labeled samples.
[0003] Learning with Noisy Label refers to using sample data containing noisy labels to train deep learning models. The accuracy of noise labels is very low. Learning with noisy labels aims to filter out relatively accurate sample data from noise samples for model training. This model training method has a low utilization rate of sample data and poor model training effect. Summary of the invention
[0004] In order to improve the accuracy and precision of the model in the noisy learning process, the embodiments of this specification provide a model training method and device, a homogeneous population screening method and device, an electronic device, a storage medium, and a computer program product.
[0005] In a first aspect, an embodiment of the present specification provides a model training method, comprising:
[0006] Acquire a sample data set, wherein the sample data set includes sample data with noise labels;
[0007] Using part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain a first intermediate model and a second intermediate model;
[0008] For sample data not used in the collaborative guidance training, generating label data corresponding to the unused sample data;
[0009] Based on the unused sample data and label data thereof, the first intermediate model and the second intermediate model are trained to obtain a first sub-model and a second sub-model.
[0010] In some embodiments, using part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain the first intermediate model and the second intermediate model includes:
[0011] Screening out a first low-noise sample from the sample data set using the first initial model, and training the second initial model using the first low-noise sample to obtain the second intermediate model;
[0012] A second low-noise sample is screened out from the sample data set using the second initial model, and the first initial model is trained using the second low-noise sample to obtain the first intermediate model.
[0013] In some implementations, for sample data not used in the collaborative guidance training, generating label data corresponding to the unused sample data includes:
[0014] Acquire a first high-noise sample, where the first high-noise sample is sample data in the sample data set except the first low-noise sample;
[0015] For each sample data in the first high-noise sample, extract features of the sample data using the first intermediate model to obtain a first feature vector;
[0016] Based on the label data corresponding to a preset number of feature vectors adjacent to the first feature vector, the label data corresponding to the sample data is determined.
[0017] In some implementations, the process of training the first intermediate model includes:
[0018] For each sample data in the first high-noise sample, input the sample data into the first intermediate model to obtain an output result predicted by the first intermediate model;
[0019] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the first intermediate model are adjusted until a convergence condition is met to obtain the first sub-model.
[0020] In some implementations, for sample data not used in the collaborative guidance training, generating label data corresponding to the unused sample data includes:
[0021] Acquire a second high noise sample, where the second high noise sample is sample data in the sample data set except the second low noise sample;
[0022] Constructing a reference model corresponding to the second intermediate model, wherein the reference model represents an average parameter model during the training process of the second intermediate model;
[0023] For each sample data in the second high-noise sample, the sample data is input into the reference model to obtain label data corresponding to the sample data output by the reference model.
[0024] In some implementations, the process of training the second intermediate model includes:
[0025] For each sample data in the second high noise sample, input the sample data into the second intermediate model to obtain an output result predicted by the second intermediate model;
[0026] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the second intermediate model are adjusted until the convergence condition is met to obtain the second sub-model.
[0027] In a second aspect, the embodiments of this specification provide a method for screening a homogeneous population, comprising:
[0028] Obtain seed population data and full population data;
[0029] Using a pre-trained homogeneous population model to extract features from the seed population data to obtain seed population features, and extracting features from the full population data to obtain full population features; wherein the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are obtained according to the model training method described in any of the above embodiments;
[0030] Feature matching is performed on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
[0031] In some embodiments, using a pre-trained homogeneous population model to extract features from the seed population data to obtain seed population features, and extracting features from the full population data to obtain full population features, includes:
[0032] Using the first sub-model to extract features from the seed population data to obtain a first feature component, using the second sub-model to extract features from the seed population data to obtain a second feature component, and fusing the first feature component and the second feature component to obtain the seed population feature;
[0033] The first sub-model is used to extract features from the entire population data to obtain a third feature component, the second sub-model is used to extract features from the entire population data to obtain a fourth feature component, and the third feature component and the fourth feature component are fused to obtain the features of the entire population.
[0034] In some implementations, the feature matching of the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data includes:
[0035] Determine the distance between each user feature in the seed population feature and each user feature in the full population feature, and obtain the homogeneous population feature corresponding to the seed population feature by matching from the full population feature according to the distance, wherein the homogeneous population feature represents the feature corresponding to the homogeneous population data.
[0036] In a third aspect, the embodiments of this specification provide a model training device, including:
[0037] A data set module is configured to obtain a sample data set, wherein the sample data set includes sample data with noise labels;
[0038] A first training module is configured to use part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain a first intermediate model and a second intermediate model;
[0039] A label generation module is configured to generate label data corresponding to the sample data not used in the collaborative guidance training;
[0040] The second training module is configured to train the first intermediate model and the second intermediate model based on the unused sample data and label data thereof to obtain a first sub-model and a second sub-model.
[0041] In some embodiments, the first training module is configured to:
[0042] Screening out a first low-noise sample from the sample data set using the first initial model, and training the second initial model using the first low-noise sample to obtain the second intermediate model;
[0043] A second low-noise sample is screened out from the sample data set using the second initial model, and the first initial model is trained using the second low-noise sample to obtain the first intermediate model.
[0044] In some implementations, the tag generation module is configured to:
[0045] Acquire a first high-noise sample, where the first high-noise sample is sample data in the sample data set except the first low-noise sample;
[0046] For each sample data in the first high-noise sample, extract features of the sample data using the first intermediate model to obtain a first feature vector;
[0047] Based on the label data corresponding to a preset number of feature vectors adjacent to the first feature vector, the label data corresponding to the sample data is determined.
[0048] In some embodiments, the second training module is configured to:
[0049] For each sample data in the first high-noise sample, input the sample data into the first intermediate model to obtain an output result predicted by the first intermediate model;
[0050] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the first intermediate model are adjusted until a convergence condition is met to obtain the first sub-model.
[0051] In some implementations, the tag generation module is configured to:
[0052] Acquire a second high noise sample, where the second high noise sample is sample data in the sample data set except the second low noise sample;
[0053] Constructing a reference model corresponding to the second intermediate model, wherein the reference model represents an average parameter model during the training process of the second intermediate model;
[0054] For each sample data in the second high-noise sample, the sample data is input into the reference model to obtain label data corresponding to the sample data output by the reference model.
[0055] In some embodiments, the second training module is configured to:
[0056] For each sample data in the second high noise sample, input the sample data into the second intermediate model to obtain an output result predicted by the second intermediate model;
[0057] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the second intermediate model are adjusted until the convergence condition is met to obtain the second sub-model.
[0058] In a fourth aspect, the embodiments of this specification provide a homogeneous population screening device, comprising:
[0059] A data acquisition module is configured to acquire seed population data and full population data;
[0060] a feature extraction module, configured to use a pre-trained homogeneous population model to perform feature extraction on the seed population data to obtain seed population features, and to perform feature extraction on the full population data to obtain full population features; wherein the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are obtained according to the model training method described in any one of the first aspects above;
[0061] The population matching module is configured to perform feature matching on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
[0062] In some embodiments, the feature extraction module is configured to:
[0063] Using the first sub-model to extract features from the seed population data to obtain a first feature component, using the second sub-model to extract features from the seed population data to obtain a second feature component, and fusing the first feature component and the second feature component to obtain the seed population feature;
[0064] The first sub-model is used to extract features from the entire population data to obtain a third feature component, the second sub-model is used to extract features from the entire population data to obtain a fourth feature component, and the third feature component and the fourth feature component are fused to obtain the features of the entire population.
[0065] In some implementations, the population matching module is configured to:
[0066] Determine the distance between each user feature in the seed population feature and each user feature in the full population feature, and obtain the homogeneous population feature corresponding to the seed population feature by matching from the full facial features according to the distance, wherein the homogeneous population feature represents the feature corresponding to the homogeneous population data.
[0067] In a fifth aspect, an embodiment of the present specification provides an electronic device, including:
[0068] processor;
[0069] A memory stores computer instructions, wherein the computer instructions, when executed by a processor, implement the method described in any of the above embodiments.
[0070] In a sixth aspect, an embodiment of this specification provides a computer program product, and the computer program product is used to implement the method described in any of the above embodiments.
[0071] The model training method implemented in the present specification includes using a sample data set containing noise labels to collaboratively guide the training of a first initial model and a second initial model, and then based on the idea of semi-supervised learning, generating corresponding label data for high-noise data not used in the collaborative guidance training, thereby using the full amount of sample data to train the model, greatly improving the utilization rate of the sample data, and increasing the scale of training data, thereby improving the model training effect and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0073] Figure 1 It is a flowchart of the model training method in some implementation methods of this specification.
[0074] Figure 2 It is a schematic diagram of the model training method in some implementation methods of this specification.
[0075] Figure 3 It is a schematic diagram of the model training method in some implementation methods of this specification.
[0076] Figure 4 It is a flowchart of the model training method in some implementation methods of this specification.
[0077] Figure 5 It is a schematic diagram of the model training method in some implementation methods of this specification.
[0078] Figure 6 It is a flowchart of the model training method in some implementation methods of this specification.
[0079] Figure 7 It is a schematic diagram of the model training method in some implementation methods of this specification.
[0080] Figure 8 is a flow chart of a homogeneous population screening method in some embodiments of the present specification.
[0081] Fig. 9 It is a schematic diagram of the principle of the homogeneous population screening method in some embodiments of the present specification.
[0082] Fig.10 It is a schematic diagram of the principle of the homogeneous population screening method in some embodiments of the present specification.
[0083] Fig.11 It is a structural block diagram of the model training device in some implementation methods of this specification.
[0084] Fig.12 It is a structural flowchart of the homogeneous population screening method in some embodiments of this specification.
[0085] Fig.13 It is a structural block diagram of an electronic device in some implementation modes of this specification. DETAILED DESCRIPTION
[0086] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0087] Homogeneous groups refer to groups of people who have high similarity in one or more characteristics, which can be age, gender, hobbies, consumption habits, user behavior, etc. The differences between individuals in a homogeneous group are small, and homogeneous groups show high consistency in behavior patterns, consumption habits, psychological needs, etc.
[0088] Homogeneous population diffusion (User Look-Alike / User Expansion) refers to finding people who are homogeneous with the seed population in the entire population. Homogeneous population diffusion helps analyze user behavior and assist business decisions, so it plays an important role in big data analysis.
[0089] For example, taking the financial service platform as an example, the seed population can be the lost users in the past period of time. By screening the homogeneous population of the seed population from the entire population, the homogeneous population with the lost users can be found among the current users. These homogeneous populations show the characteristics of easy churn. Therefore, by analyzing the behavioral patterns, balance calculation and other dimensions of the homogeneous population, the abnormal attribution of lost users can be accurately found, thereby assisting the financial service platform to make targeted decisions and improve user stickiness.
[0090] At present, the homogeneous population diffusion method is mainly divided into the following two stages:
[0091] 1. Model training
[0092] The goal of model training is to obtain a deep learning model through training with a large amount of sample data. The role of the deep learning model is to extract the feature vectors of population data and accurately capture the user characteristics among homogeneous populations.
[0093] 2. Crowd Matching
[0094] After the user feature vector is extracted through the deep learning model, the user feature vector of the seed population is matched with the user feature vector of the entire population, so as to screen out people with the same characteristics as the seed population from the entire population.
[0095] During the model training phase, deep learning models require large-scale, high-accuracy labeled sample data for training to achieve better performance. However, in real scenarios, the labeling cost of sample data is high. Often only a small amount of sample data is labeled. There is also a large amount of unlabeled data or data with low label accuracy, making it difficult to obtain large-scale, high-accuracy labeled sample data.
[0096] Learning with Noisy Label refers to using sample data containing noisy labels to train deep learning models. Noisy labels refer to sample data with low label accuracy. For example, when a computer program is used to automatically label sample data in batches, the label data generated by the program has low accuracy, which is called a noisy label. The algorithm that uses these sample data containing noisy labels to train learning models is called a noisy learning algorithm.
[0097] Traditional noisy learning algorithms aim to filter out relatively accurate sample data (called clean samples) from the full amount of noisy label data for model training. This algorithm only uses part of the filtered sample data for model training, and the utilization rate of sample data is not high, which limits the performance improvement of deep learning models.
[0098] Based on this, the embodiments of this specification provide a model training method and device, a homogeneous population screening method and device, an electronic device, a storage medium, and a computer program product, which aim to label unused sample data with pseudo labels (PseudoLabels) based on noisy learning and combined with the idea of semi-supervised learning, so as to realize the training of deep learning models in combination with the full amount of sample data, improve sample utilization, and thus improve model accuracy.
[0099] Semi-supervised learning refers to the process of training a deep learning model using a small amount of labeled sample data and a large amount of unlabeled sample data. Semi-supervised learning can use both labeled and unlabeled samples to train the model to improve the accuracy of the model. Pseudo-labeling is a technique for labeling sample data in semi-supervised learning. Pseudo-labels are generated by predicting unlabeled sample data, and then these pseudo-labels are used together with sample data to train the model.
[0100] Figure 1 A flowchart of a model training method in some embodiments of this specification is shown below. Figure 1 Provide explanation.
[0101] like Figure 1 As shown, in some embodiments, the model training method exemplified in this specification includes:
[0102] S110: Obtain a sample data set.
[0103] A sample data set refers to a collection of all sample data used for model training. For example, in a homogeneous population screening task scenario, the sample data set is the full population data.
[0104] It is worth noting that in the implementation manner of this specification, the label data (Label) corresponding to the sample data in the sample data set is not necessarily accurate, that is, the sample data set contains sample data with noisy labels.
[0105] For example, taking the full population data as an example, a sample data in the sample data set represents the user data corresponding to a certain user. The label data (Label) corresponding to the sample data indicates whether the user corresponding to the sample data is a homogeneous population of the seed population. If the label data is positive, it means that the sample data is a positive sample, that is, the user corresponding to the sample data is a homogeneous population. If the label data is negative, it means that the sample data is a negative sample, that is, the user corresponding to the sample data is not a homogeneous population.
[0106] It can be understood that during the model training process, the model parameters need to be continuously optimized according to the loss term between the output results and the label data. Therefore, the accuracy of the label data corresponding to the sample data will directly affect the training effect of the model.
[0107] In the implementation manner of this specification, the label data corresponding to the sample data is not necessarily accurate, that is, the label data corresponding to some sample data may be accurate, while the label data corresponding to some sample data may be wrong. The sample data with wrong label data is the sample data with noisy labels.
[0108] In some implementations, sample data with accurate label data can be obtained, for example, sample data with accurate label data can be obtained through manual labeling, or sample data with accurate label data that has been labeled can be obtained from public data sets. It is understandable that this part of sample data is often small in size and difficult to support large-scale model training tasks.
[0109] Therefore, a large amount of unlabeled sample data can be further obtained. Since the sample data has not been labeled, these sample data have no label data. Then, it can be assumed that these sample data are all negative samples or all positive samples, that is, it is assumed that the sample data corresponding to these sample data are all positive or negative. In this case, there must be some sample data whose label data is correct, and there are also some sample data whose label data is wrong. Afterwards, the sample data with the accurate label data and the sample data with the assumed label are mixed to construct a sample data set containing a large number of noise labels.
[0110] S120. Use part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain a first intermediate model and a second intermediate model.
[0111] Co-teaching is a noisy learning algorithm that is specifically designed to handle situations where there are a large number of noisy labels in the training data. The basic principles of the Co-teaching algorithm are explained below.
[0112] First, the Co-teaching algorithm needs to build two neural network models, the model structures of the two neural network models are exactly the same, but the initial model parameters are different. Therefore, in the implementation of this specification, a first initial model and a second initial model can be pre-built, the network structures of the first initial model and the second initial model are the same, and different initial model parameters are initialized for the two initial models.
[0113] Then, when using the sample data set for collaborative guidance training, in each training iteration, the first initial model will select a portion of clean sample data that it considers to be the most reliable (i.e., less noisy) from the current batch of sample data, and pass this portion of clean sample data to the second initial model for training. At the same time, the second initial model will also select a portion of clean sample data that it considers to be the most reliable (i.e., less noisy) from the current batch of sample data, and pass this portion of clean sample data to the first initial model for training.
[0114] It can be seen that in the collaborative guidance training process, the first initial model selects some clean samples from the full sample data set and passes these clean samples to the second initial model for training. At the same time, the second initial model selects some clean samples from the full sample data set and passes these clean samples to the first initial model for training. That is, no matter the first initial model or the second initial model, the sample data used for training are the sample data selected from the sample data set, rather than the full sample data.
[0115] See also Figure 2 As shown, in the implementation mode of this specification, the sample data set U is taken as an example. In the collaborative guidance training process, the first initial model g1 selects the clean samples it considers from the sample data set U. Defining clean samples is the first low noise sample Remove the first low noise sample from the sample data set U The other sample data is the first high noise sample That is, the first low-noise sample It is the clean sample selected by the first initial model g1 from the sample data set U, and the first high noise sample That is, the remaining noise samples in the sample data set U.
[0116] Similarly, the second initial model g2 selects the clean samples it considers from the sample data set U Defining clean samples The second lowest noise sample Remove the second low noise sample from the sample data set U The other sample data is the second highest noise sample That is, the second low noise sample It is the clean sample selected by the second initial model g2 from the sample data set U, and the second high noise sample That is, the remaining noise samples in the sample data set U.
[0117] Then, using the first low-noise sample The second initial model g2 is trained to obtain the second intermediate model g2'. The first initial model g1 is trained to obtain the first intermediate model g1'. That is, the two models exchange their selected clean samples with each other, and then use these clean samples to train and update each other's parameters.
[0118] Specifically, taking the training process of the first initial model g1 as an example, for the second low-noise sample Any sample data in the sample data can be input into the first initial model g1 to obtain the output result predicted by the first initial model g1. It can be understood that the goal of model training is to align the output result of the first initial model g1 with the label data corresponding to the actual sample data. Therefore, the loss (Loss) between the output result and the label data of the sample data can be calculated, and then the model parameters of the first initial model g1 can be adjusted according to the back propagation of the loss to complete an iterative tuning process. The above only takes one sample as an example. Using the second low-noise sample The above training process is repeated to continuously optimize the model parameters of the first initial model g1 until the convergence condition is met, thereby obtaining the trained first intermediate model g1'.
[0119] Similarly, for the training process of the second initial model g2, the first low-noise sample Take any sample data in the dataset and input the sample data into the second initial model g2 to get the output result predicted by the second initial model g2. Then, calculate the loss between the output result and the label data of the sample data, and then adjust the model parameters of the second initial model g2 according to the back propagation of the loss, and then complete an iterative tuning process. The above training process is repeated to continuously optimize the model parameters of the second initial model g2 until the convergence condition is met to obtain the trained second intermediate model g2'.
[0120] Model convergence conditions include but are not limited to: the number of training rounds reaches a preset number of rounds, the model accuracy reaches a preset threshold, the training data reaches a preset amount of data, etc. This manual does not impose any restrictions on this.
[0121] It can be understood from the above that the first low noise sample and the second low noise sample Both are clean samples with high accuracy of label data selected by the Co-teaching algorithm. Therefore, the first intermediate model and the second intermediate model obtained by training are models with relatively better effects.
[0122] S130 . For sample data not used in collaborative guidance training, generate label data corresponding to the unused sample data.
[0123] S140. Based on the unused sample data and label data thereof, the first intermediate model and the second intermediate model are trained to obtain a first sub-model and a second sub-model.
[0124] Combination Figure 2 As described above, in the Co-teaching algorithm process, only the first low-noise sample is used. and the second low noise sample Model training is performed, that is, only a part of the sample data in the sample data set U is used, and for the first high noise sample and the second highest noise sample It is then screened out as noise, resulting in a low actual sample utilization rate. and the second low noise sample The proportion of data in the sample data set U is low, and the amount of data actually involved in model training is small, which leads to the failure of the model to converge well and poor model accuracy.
[0125] Therefore, in the embodiment of this specification, for the first high noise sample not used in the co-teaching training process and the second highest noise sample Instead of directly screening out the samples with inaccurate label data, more accurate label data is generated for them. Then, these high-noise samples are used to further train the model, making full use of the full amount of sample data to improve the model effect.
[0126] And in some embodiments, for the first high noise sample and the second highest noise sample Different algorithms are used to generate the first high noise sample and the second highest noise sample For example, see Figure 3 As shown, the first high noise sample The corresponding label data is generated by the first labeling algorithm, and the second high noise sample The corresponding label data is generated by a second marking algorithm, and the first marking algorithm is different from the second marking algorithm.
[0127] The purpose of this is to use pseudo-label data (Pseudo Labels) generated by different marking algorithms to make full use of the differences between the two models, so that the model can capture more comprehensive feature information during feature extraction, thereby improving the model's feature expression effect and output result accuracy.
[0128] Figure 4 and Figure 5 The process and principle diagram of the first marking algorithm are shown below. Figure 4 and Figure 5 Provide explanation.
[0129] like Figure 4 As shown, in some embodiments, a first high noise sample is generated The process of labeling data includes:
[0130] S410: Acquire a first high noise sample.
[0131] Combination Figure 3 As shown, the first high noise sample is It represents the first low-noise sample in the sample data set U except the first initial model g1. The remaining sample data can be understood as sample data with inaccurate label data.
[0132] S420 . For each sample data in the first high-noise sample data, use the first intermediate model to perform feature extraction on the sample data to obtain a first feature vector.
[0133] See also Figure 5 As shown, the first high noise sample is Taking any sample data i in as an example, the sample data i is input into the first intermediate model g1', so that the first intermediate model g1' extracts the feature vector of the sample data i and outputs the first feature vector Zi of high-dimensional representation.
[0134] S430 : Determine label data corresponding to the sample data based on label data corresponding to a preset number of feature vectors adjacent to the first feature vector.
[0135] Continue to refer to Figure 5 In this example, the k-nearest neighbor algorithm can be used to determine the label data corresponding to the sample data i.
[0136] Specifically, after obtaining the first eigenvector Zi corresponding to the sample data i, the first eigenvector Zi and the first low-noise sample data i can be calculated. The distance between the feature vectors of , which can be expressed as cosine distance, is then sorted from low to high according to the distance, and the feature vectors of the top k clean samples are selected. For example, Figure 5 In this example, the value of k is 4, that is, the feature vectors of four clean samples adjacent to the first feature vector Zi are selected.
[0137] In some implementations, the label data corresponding to the sample data i may be a hard false label, which refers to directly assigning a specific positive or negative type of label data to the sample data i. Figure 5 In this example, the label data of the k=4 feature vectors adjacent to the first feature vector Zi are 3 positive and 1 negative, so the label data with the largest proportion can be directly determined as the hard false label data corresponding to the sample data i through the recommendation method, that is, Figure 5 In the example, the label data corresponding to sample data i is positive.
[0138] In other embodiments, the label data corresponding to the sample data i may be a soft false label, which is not a label type directly assigned to the sample data i, but means that the label data of the sample data i is a probability distribution, which indicates the possibility that the sample data i belongs to each category. Figure 5 In the example, the label data of k=4 feature vectors adjacent to the first feature vector Zi are 3 positive and 1 negative, so it can be determined that the soft false label corresponding to the sample data i is 75% positive and 25% negative.
[0139] The above only takes the first high noise sample Taking any sample data i in as an example, the process of determining the label data of sample data i is explained. For the first high noise sample For each sample data of Each sample data in generates corresponding label data, which will not be described in detail in this manual.
[0140] Combination Figure 3 As shown, for the first high noise sample After generating the label data, the first high-noise sample containing more accurate label data is obtained Then we can use the first high noise sample The first intermediate model g1' is further trained.
[0141] Specifically, for the first high noise sample Any sample data in the first intermediate model g1' can be input into the first intermediate model g1' to obtain the output result predicted by the first intermediate model g1'. It can be understood that the goal of model training is to align the output result of the first intermediate model g1' with the label data corresponding to the actual sample data. Therefore, the loss (Loss) between the output result and the label data of the sample data can be calculated, and then the model parameters of the first intermediate model g1' can be adjusted according to the back propagation of the loss to complete an iterative tuning process. The above only takes one sample as an example. Using the first high noise sample The above training process is repeated to continuously optimize the model parameters of the first intermediate model g1' until the convergence condition is met to obtain the trained first sub-model G1.
[0142] Figure 6 and Figure 7 The process and principle diagram of the second marking algorithm are shown below. Figure 6 and Figure 7 Provide explanation.
[0143] like Figure 6 As shown, in some embodiments, a second high noise sample is generated The process of labeling data includes:
[0144] S610: Obtain a second highest noise sample.
[0145] Combination Figure 3 As shown, the second highest noise sample is It represents the second low-noise sample in the sample data set U except the second initial model g2 The remaining sample data can be understood as sample data with inaccurate label data.
[0146] S620: Construct a reference model corresponding to the second intermediate model.
[0147] In this example, the moving average method is used to generate the reference model, for example Figure 7In this example, the first low noise sample is used During multiple rounds of iterative training of the second initial model g2, the model parameters of several consecutive rounds of iterations can be averaged to obtain a reference model g2'', that is, the reference model g2'' represents the average model of multiple consecutive rounds of iterations during the training of the first initial model g2.
[0148] S630: For each sample data in the second high noise sample, the sample data is input into a reference model to obtain label data corresponding to the sample data output by the reference model.
[0149] In the implementation mode of this specification, the second highest noise sample Taking any sample data j in as an example, the sample data j is input into the reference model. The reference model extracts features from the sample data j, and then outputs the label data corresponding to the sample data j according to the characteristic vector prediction. The label data can be the hard false label data or the soft false label data mentioned above. Those skilled in the art can understand this in combination with the foregoing and will not elaborate on it.
[0150] The above only uses the second highest noise sample Taking any sample data j in as an example, the process of determining the label data of sample data j is explained. Repeat the above process for each sample data of the second highest noise sample. Each sample data in generates corresponding label data, which will not be described in detail in this manual.
[0151] Combination Figure 3 As shown, for the second highest noise sample After generating the label data, the second highest noise sample containing more accurate label data is obtained Then we can use the second highest noise sample The second intermediate model g2' is further trained.
[0152] Specifically, for the second highest noise sample Any sample data in the sample data can be input into the second intermediate model g2' to obtain the output result predicted by the second intermediate model g2'. It can be understood that the goal of model training is to align the output result of the second intermediate model g2' with the label data corresponding to the actual sample data. Therefore, the loss (Loss) between the output result and the label data of the sample data can be calculated, and then the model parameters of the second intermediate model g2' can be adjusted according to the back propagation of the loss to complete an iterative tuning process. The above only takes one sample as an example. Using the second high noise sample The above training process is repeated to continuously optimize the model parameters of the second intermediate model g2' until the convergence condition is met to obtain the trained second sub-model G2.
[0153] In one example, during the above model training process, the loss between the calculated output result and the label data of the sample data can be calculated using the KL (Kullback-Leibler) loss function. Of course, other loss functions can also be used, and this specification does not limit this.
[0154] Combination Figure 2 It can be seen that in the traditional Co-teaching algorithm, only part of the clean sample data is used for model training, which leads to low utilization of sample data and small sample data scale, resulting in poor model effect. Figure 3 As shown in the figure, based on the idea of semi-supervised learning, corresponding label data is generated for high-noise data, so that the model can be trained using the full amount of sample data, which greatly improves the utilization rate of sample data and increases the scale of training data, thereby improving the model training effect and accuracy. In addition, by using different labeling algorithms to generate label data for high-noise samples of the two models, the difference between the two models is fully improved, so that the model can capture more comprehensive feature information during feature extraction, thereby improving the model's feature expression effect and output result accuracy.
[0155] The above describes the method and principle of model training. After the model is trained through the above method, the model can be used to screen homogeneous populations. The following describes the method for screening homogeneous populations.
[0156] like Figure 8 As shown, in some embodiments, the homogeneous population screening method exemplified in this specification includes:
[0157] S810. Obtain seed population data and full population data.
[0158] Based on the above, we can see that the seed population refers to the target population that needs to be found from the entire population with similar populations.
[0159] For example, in a sample scenario, taking a financial service platform as an example, the seed population represents one or more user groups that have been lost in the historical period. The goal of screening homogeneous populations is to find people who are homogeneous with the seed population from all users. These homogeneous populations are highly similar to lost users, which means that these homogeneous populations may also be lost in the future. Therefore, it is necessary to screen out homogeneous populations. By analyzing the behavioral patterns, balance calculation and other dimensions of homogeneous populations, the abnormal attribution of lost users can be accurately found, thereby assisting the financial service platform to make targeted decisions and improve user stickiness.
[0160] In this example, the seed population data refers to the data of lost users in the historical period, and the full population data refers to the full user data of the financial service platform.
[0161] S820. Use a pre-trained homogeneous population model to perform feature extraction on the seed population data to obtain seed population features, and perform feature extraction on the full population data to obtain full population features.
[0162] In the implementation of this specification, the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are the first sub-model G1 and the second sub-model G2 obtained by the model training method of any of the aforementioned implementations. For the process of model training, those skilled in the art can refer to the aforementioned, and this specification will not elaborate on it.
[0163] See also Fig. 9 As shown, in this example, the seed population data is S and the full population data is U. First, the seed population data S is input into the first sub-model G1 and the second sub-model G2 respectively. The first sub-model G1 extracts the features of the seed population data S to obtain the first feature component and the second eigencomponent
[0164] It can be understood that since the first sub-model G1 and the second sub-model G2 maintain certain differences through the above training method, the first characteristic component and the second eigencomponent It can capture data features of different dimensions, so that fusing feature components of multiple dimensions can obtain feature vectors with richer information.
[0165] Therefore, in the embodiment of this specification, the first characteristic component is extracted by the first sub-model G1 and the second sub-model G2 respectively. and the second eigencomponent After that, the first feature component can be fused through the feature fusion algorithm and the second eigencomponent Get the seed population characteristics Z s In one example, the feature fusion algorithm may adopt additive fusion or weighted summation, etc., which can be understood by those skilled in the art and will not be described in detail in this specification.
[0166] The same is true for the full population data U. The full population data U is input into the first sub-model G1 and the second sub-model G2 respectively. The first sub-model G1 extracts features from the full population data U to obtain the third feature component. and the fourth eigencomponent
[0167] It can be understood that since the first sub-model G1 and the second sub-model G2 maintain certain differences through the above training method, the third characteristic component and the fourth eigencomponent It can capture data features of different dimensions, so that fusing feature components of multiple dimensions can obtain feature vectors with richer information.
[0168] Therefore, in the embodiment of this specification, the third characteristic component is extracted by the first sub-model G1 and the second sub-model G2 respectively. and the fourth eigencomponent After that, the third feature component can be fused through the feature fusion algorithm and the fourth eigencomponent Get the seed population characteristics Z U In one example, the feature fusion algorithm may adopt additive fusion or weighted summation, etc., which can be understood by those skilled in the art and will not be described in detail in this specification.
[0169] S830: Perform feature matching on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
[0170] In the implementation mode of this specification, after obtaining the seed population feature Z s and the overall population characteristics Z U After that, we can use feature matching to identify the features of all people Z. U Matching and seed population characteristics Z s Similar homogeneous population characteristics.
[0171] For example, in an example, the k nearest neighbor matching algorithm can be used to match the seed population feature Z s and the overall population characteristics Z U For feature matching, see Fig.10 As shown, taking any seed population feature Z s For example, we can calculate the seed population characteristic Z s With each full population characteristic Z U The distance, and then select the k closest full-scale population features Z U As a feature of homogeneous population, the value of k can be selected according to the needs of specific scenarios, and this specification does not impose any restrictions on this.
[0172] For example, in the financial service platform scenario of the aforementioned example, the seed population is taken as an example of lost users, and the homogeneous population corresponding to the screened seed population represents users who are homogeneous with the lost users.
[0173] From the above, it can be seen that in the implementation of this specification, the feature matching algorithm is used to screen homogeneous populations, which can better reflect the overall distribution of the seed population compared to the score sorting algorithm, and improve the homogeneity of the recalled homogeneous population and the seed population. Moreover, the nearest neighbor matching algorithm is used for feature matching, which can effectively utilize the representation space of the model and further improve the accuracy of the recalled homogeneous population.
[0174] In some embodiments, this specification provides a model training device, such as Fig.11 As shown, the device comprises:
[0175] The data set module 10 is configured to obtain a sample data set, wherein the sample data set includes sample data with noise labels;
[0176] A first training module 20 is configured to use part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain a first intermediate model and a second intermediate model;
[0177] The label generation module 30 is configured to generate label data corresponding to the sample data not used in the collaborative guidance training;
[0178] The second training module 40 is configured to train the first intermediate model and the second intermediate model based on the unused sample data and label data thereof to obtain a first sub-model and a second sub-model.
[0179] In some embodiments, the first training module 20 is configured to:
[0180] Screening out a first low-noise sample from the sample data set using the first initial model, and training the second initial model using the first low-noise sample to obtain the second intermediate model;
[0181] A second low-noise sample is screened out from the sample data set using the second initial model, and the first initial model is trained using the second low-noise sample to obtain the first intermediate model.
[0182] In some implementations, the label generation module 30 is configured to:
[0183] Acquire a first high-noise sample, where the first high-noise sample is sample data in the sample data set except the first low-noise sample;
[0184] For each sample data in the first high-noise sample, extract features of the sample data using the first intermediate model to obtain a first feature vector;
[0185] Based on the label data corresponding to a preset number of feature vectors adjacent to the first feature vector, the label data corresponding to the sample data is determined.
[0186] In some embodiments, the second training module 40 is configured to:
[0187] For each sample data in the first high-noise sample, input the sample data into the first intermediate model to obtain an output result predicted by the first intermediate model;
[0188] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the first intermediate model are adjusted until a convergence condition is met to obtain the first sub-model.
[0189] In some implementations, the label generation module 30 is configured to:
[0190] Acquire a second high noise sample, where the second high noise sample is sample data in the sample data set except the second low noise sample;
[0191] Constructing a reference model corresponding to the second intermediate model, wherein the reference model represents an average parameter model during the training process of the second intermediate model;
[0192] For each sample data in the second high-noise sample, the sample data is input into the reference model to obtain label data corresponding to the sample data output by the reference model.
[0193] In some embodiments, the second training module 40 is configured to:
[0194] For each sample data in the second high noise sample, input the sample data into the second intermediate model to obtain an output result predicted by the second intermediate model;
[0195] Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the second intermediate model are adjusted until the convergence condition is met to obtain the second sub-model.
[0196] In some embodiments, the present specification provides a homogeneous population screening device, such as Fig.12 As shown, the device comprises:
[0197] The data acquisition module 50 is configured to acquire seed population data and full population data;
[0198] The feature extraction module 60 is configured to use a pre-trained homogeneous population model to perform feature extraction on the seed population data to obtain seed population features, and to perform feature extraction on the full population data to obtain full population features; wherein the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are obtained according to the model training method described in any one of the first aspects above;
[0199] The population matching module 70 is configured to perform feature matching on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
[0200] In some embodiments, the feature extraction module 60 is configured to:
[0201] Using the first sub-model to extract features from the seed population data to obtain a first feature component, using the second sub-model to extract features from the seed population data to obtain a second feature component, and fusing the first feature component and the second feature component to obtain the seed population feature;
[0202] The first sub-model is used to extract features from the entire population data to obtain a third feature component, the second sub-model is used to extract features from the entire population data to obtain a fourth feature component, and the third feature component and the fourth feature component are fused to obtain the features of the entire population.
[0203] In some implementations, the population matching module 70 is configured to:
[0204] Determine the distance between each user feature in the seed population feature and each user feature in the full population feature, and obtain the homogeneous population feature corresponding to the seed population feature by matching from the full facial features according to the distance, wherein the homogeneous population feature represents the feature corresponding to the homogeneous population data.
[0205] In some implementations, this specification provides an electronic device, including:
[0206] processor;
[0207] A memory stores computer instructions, wherein the computer instructions, when executed by a processor, implement the method described in any of the above embodiments.
[0208] In some embodiments, this specification provides a computer program product, which is used to implement the method described in any of the above embodiments.
[0209] Specifically, Fig.13 A schematic diagram of the structure of an electronic device 600 suitable for implementing the method disclosed in the present invention is shown. Fig.13 The electronic device shown can realize the corresponding functions of the above-mentioned processor and storage medium.
[0210] like Fig.13 As shown, the electronic device 600 includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the memory 602 or the program loaded from the storage part 608 into the memory 602. In the memory 602, various programs and data required for the operation of the electronic device 600 are also stored. The processor 601 and the memory 602 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0211] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.
[0212] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the embodiments. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A model training method, characterized in that: include: Acquire a sample data set, wherein the sample data set includes sample data with noise labels; Using part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain a first intermediate model and a second intermediate model; For sample data not used in the collaborative guidance training, generating label data corresponding to the unused sample data; Based on the unused sample data and label data thereof, the first intermediate model and the second intermediate model are trained to obtain a first sub-model and a second sub-model.
2. The method according to claim 1, characterized in that The method of using part of the sample data in the sample data set to perform collaborative guidance training on the first initial model and the second initial model to obtain the first intermediate model and the second intermediate model includes: Screening out a first low-noise sample from the sample data set using the first initial model, and training the second initial model using the first low-noise sample to obtain the second intermediate model; A second low-noise sample is screened out from the sample data set using the second initial model, and the first initial model is trained using the second low-noise sample to obtain the first intermediate model.
3. The method according to claim 2, characterized in that For sample data not used in the collaborative guidance training, generating label data corresponding to the sample data not used, including: Acquire a first high-noise sample, where the first high-noise sample is sample data in the sample data set except the first low-noise sample; For each sample data in the first high-noise sample, extract features of the sample data using the first intermediate model to obtain a first feature vector; Based on the label data corresponding to a preset number of feature vectors adjacent to the first feature vector, the label data corresponding to the sample data is determined.
4. The method according to claim 3, characterized in that The process of training the first intermediate model includes: For each sample data in the first high-noise sample, input the sample data into the first intermediate model to obtain an output result predicted by the first intermediate model; Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the first intermediate model are adjusted until a convergence condition is met to obtain the first sub-model.
5. The method according to claim 2, characterized in that: For sample data not used in the collaborative guidance training, generating label data corresponding to the sample data not used, including: Acquire a second high noise sample, where the second high noise sample is sample data in the sample data set except the second low noise sample; Constructing a reference model corresponding to the second intermediate model, wherein the reference model represents an average parameter model during the training process of the second intermediate model; For each sample data in the second high-noise sample, the sample data is input into the reference model to obtain label data corresponding to the sample data output by the reference model.
6. The method according to claim 5, characterized in that The process of training the second intermediate model includes: For each sample data in the second high noise sample, input the sample data into the second intermediate model to obtain an output result predicted by the second intermediate model; Based on the difference between the output result and the label data corresponding to the sample data, the model parameters of the second intermediate model are adjusted until the convergence condition is met to obtain the second sub-model.
7. A method for screening homogeneous populations, characterized in that: include: Obtain seed population data and full population data; Using a pre-trained homogeneous population model, feature extraction is performed on the seed population data to obtain seed population features, and feature extraction is performed on the full population data to obtain full population features; wherein the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are obtained according to the model training method according to any one of claims 1 to 6; Feature matching is performed on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
8. The method according to claim 7, characterized in that The seed population data is subjected to feature extraction using a pre-trained homogeneous population model to obtain seed population features, and the full population data is subjected to feature extraction to obtain full population features, including: Using the first sub-model to extract features from the seed population data to obtain a first feature component, using the second sub-model to extract features from the seed population data to obtain a second feature component, and fusing the first feature component and the second feature component to obtain the seed population feature; The first sub-model is used to extract features from the entire population data to obtain a third feature component, the second sub-model is used to extract features from the entire population data to obtain a fourth feature component, and the third feature component and the fourth feature component are fused to obtain the features of the entire population.
9. The method according to claim 7, characterized in that: The feature matching of the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data includes: Determine the distance between each user feature in the seed population feature and each user feature in the full population feature, and obtain the homogeneous population feature corresponding to the seed population feature by matching from the full population feature according to the distance, wherein the homogeneous population feature represents the feature corresponding to the homogeneous population data.
10. A homogeneous population screening device, characterized in that: include: A data acquisition module is configured to acquire seed population data and full population data; a feature extraction module, configured to use a pre-trained homogeneous population model to perform feature extraction on the seed population data to obtain seed population features, and to perform feature extraction on the full population data to obtain full population features; wherein the homogeneous population model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are obtained according to the model training method according to any one of claims 1 to 6; The population matching module is configured to perform feature matching on the seed population features and the full population features to determine homogeneous population data corresponding to the seed population data.
11. An electronic device, characterized in that: include: processor; A memory storing computer instructions, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that The computer program product is used to implement the method according to any one of claims 1 to 9.