A model training method and device, a computer storage medium and an apparatus

By employing hierarchical sampling and model fusion, multiple models are trained and fused using sample datasets with varying difficulty levels. This approach addresses the overfitting and underfitting issues of neural network models, thereby improving prediction performance and generalization ability.

CN113762579BActive Publication Date: 2025-11-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110018592.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-07
Publication Date
2025-11-18
Estimated Expiration
2041-01-07

AI Technical Summary

Technical Problem

In existing technologies, neural network models are prone to overfitting and underfitting when randomly sampling sample data, resulting in poor prediction performance and weak generalization ability.

Method used

By acquiring at least two sample datasets, each containing sample data with different difficulty ratios, training a preset model separately, and then merging the trained models to form the target model.

Benefits of technology

It improves the model's prediction accuracy and generalization ability, avoids the problem of overfitting a single model, makes full use of the features of the sample data, and enhances the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113762579B_ABST
    Figure CN113762579B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model training method and device, a computer storage medium and equipment, the method comprising: obtaining at least two sample data sets; wherein the at least two sample data sets each comprise sample data with different difficulty ratios; training at least two preset models using the at least two sample data sets respectively to obtain at least two trained models; and fusing the at least two trained models to obtain a target model. In this way, by using at least two sample data sets with different difficulty ratios, hierarchical sampling of sample data is achieved, thereby improving the prediction effect of the target model obtained ultimately; moreover, the at least two sample data sets are used for model training respectively, and then the target model is obtained through model fusion, thereby improving the generalization ability of the target model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning, and particularly relates to a model training method and device, a computer storage medium and equipment. BACKGROUND

[0002] At present, when a prediction or ranking problem is solved by using machine learning, a neural network model needs to be trained by using sample data, and then the trained neural network model is used to solve an actual problem. However, in the related art, sample data is obtained by random sampling, and different sample data is not distinguished, so the neural network model is prone to overfitting and underfitting, resulting in poor prediction effect of the trained neural network model. SUMMARY

[0003] The present application provides a model training method and device, a computer storage medium and equipment, which obtain a target model by hierarchical sampling and model fusion, so as to improve the prediction accuracy and generalization ability of the target model.

[0004] The technical scheme of the present application is implemented as follows:

[0005] In a first aspect, the present application provides a model training method, which comprises the following steps:

[0006] obtaining at least two sample data sets; wherein the sample data included in each of the at least two sample data sets is different in difficulty ratio;

[0007] training at least two preset models by using the at least two sample data sets respectively, to obtain at least two trained models;

[0008] model fusion is performed on the at least two trained models to obtain a target model.

[0009] In a second aspect, the present application provides a model training device, which comprises an obtaining unit, a training unit and a fusion unit, wherein:

[0010] The obtaining unit is configured to obtain at least two sample data sets; wherein the sample data included in each of the at least two sample data sets is different in difficulty ratio;

[0011] The training unit is configured to train at least two preset models by using the at least two sample data sets respectively, to obtain at least two trained models;

[0012] The fusion unit is configured to perform model fusion on the at least two trained models to obtain a target model.

[0013] In a third aspect, an embodiment of the present application provides a model training apparatus, the model training apparatus comprising a memory and a processor; wherein

[0014] The memory is configured to store a computer program capable of running on the processor.

[0015] The processor is configured to execute the steps of the method according to the first aspect when running the computer program.

[0016] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores a model training program, and the model training program, when executed by at least one processor, implements the steps of the method according to the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a model training device, which comprises at least the model training apparatus according to the second aspect or the third aspect.

[0018] The embodiments of the present application provide a model training method, apparatus, computer storage medium and device. At least two sample data sets are obtained, wherein the sample data included in each of the at least two sample data sets has different difficulty ratios. At least two preset models are trained by using the at least two sample data sets respectively, to obtain at least two trained models. The at least two trained models are fused to obtain a target model. In this way, by using the at least two sample data sets with different difficulty ratios, hierarchical sampling of sample data is realized, so as to improve the prediction effect of the target model obtained finally. Moreover, the at least two sample data sets are used for model training respectively, and then the target model is obtained by model fusion, which also improves the generalization ability of the target model. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A working process diagram of a similar population expansion model provided by an embodiment of the present application;

[0020] Figure 2 A flowchart of a model training method provided by an embodiment of the present application;

[0021] Figure 3 A flowchart of another model training method provided by an embodiment of the present application;

[0022] Figure 4 A flowchart of still another model training method provided by an embodiment of the present application;

[0023] Figure 5 A working process diagram of a model training method provided by an embodiment of the present application;

[0024] Figure 6Another model training method provided by the embodiment of the present application provides a working process diagram of the model training method;

[0025] Figure 7 The composition structure diagram of the model training device provided by the embodiment of the present application;

[0026] Figure 8 The composition structure diagram of another model training device provided by the embodiment of the present application

[0027] Figure 9 The hardware structure diagram of the model training device provided by the embodiment of the present application;

[0028] Figure 10 The composition structure diagram of another model training device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application. In addition, it should be noted that, for the convenience of description, only the parts related to the application are shown in the drawings.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0031] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.

[0032] It should be noted that the terms "first", "second", "third" involved in the embodiments of the present application are only to distinguish similar objects, and do not represent the specific order of the objects. It can be understood that "first", "second", "third" can be interchanged in specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0033] The look-alike model is a popular audience extension technology. Specifically, the audience extension technology refers to finding more similar people with potential correlations based on seed users by using the look-alike model. The look-alike model can be applied in the field of advertising. The core idea is to circle a part of seed users according to historical behaviors for a certain item, and then find similar people to the seed users by using the look-alike model to recommend the item to them.

[0034] The look-alike model can be used for precise marketing in different scenarios to help customers obtain target people, such as education, finance, and automobile. See Figure 1 FIG. 1 shows a working process diagram of a look-alike model provided by an embodiment of the present application. As shown in Figure 1 The working process of the look-alike model can be divided into user association, sample sampling, user feature extraction, model training, and model prediction, and specifically includes the following steps.

[0035] S101: Obtain positive samples and negative samples provided by a customer in a specific scenario.

[0036] It should be noted that the positive samples and negative samples provided by the customer are obtained. Generally, the positive sample people are the target people who actually have a preset conversion behavior (for example, a purchase behavior), and the negative sample people are users who do not have a preset expected behavior (for example, no purchase, no click, etc.).

[0037] S102: Associate with the company to filter cold users.

[0038] It should be noted that the positive samples and negative samples provided by the customer are obtained. Generally, the positive sample people are the target people who actually have a preset conversion behavior (for example, a purchase behavior), and the negative sample people are users who do not have a preset expected behavior (for example, no purchase, no click, etc.).

[0039] The specific association method can be to use a mobile phone identification code (pin) to perform a database collision, etc. Then, according to the association result, the seed people and the negative sample people are determined in the user group of the company.

[0040] S103: Determine the seed people.

[0041] It should be noted that according to the user association result, the user associated with the positive sample is taken as the seed people (or called the positive sample people), and the user associated with the negative sample is taken as the negative sample people.

[0042] S104: Screen the user group of the company.

[0043] It should be noted that if the negative sample population is small, the user group of the company can also be screened, and the users who are neither the positive samples provided by the customer company nor the negative samples provided by the customer company are added to the negative sample population.

[0044] S105: Determine the negative sample population.

[0045] Here, step S102 and step S104 can be executed in parallel, and the execution order of the two is not distinguished. Specifically, after step S102, the positive samples can be determined, and then the seed population is determined. After steps S102 and S104, the negative samples can be determined, and then the negative sample population is determined.

[0046] That is, according to the results of user association and screening the user group of the company, the negative sample population is determined. In this way, through user association and screening, the positive sample population and the negative sample population are finally obtained.

[0047] S106: User feature extraction.

[0048] It should be noted that according to the user feature data of a certain company, the seed population and the negative sample population can be extracted, and the user features of each user in the positive sample population and the negative sample population are obtained, for example, the extracted user features can be user portrait and user behavior.

[0049] S107: Model training of the extracted features using a preset model training module.

[0050] It should be noted that the extracted features are trained using a preset model training module (such as a Look-alike model), that is, the Look-alike model is trained according to the user features of the positive sample population and the negative sample population.

[0051] S108: Obtain the to-be-predicted population.

[0052] It should be noted that the user data of the to-be-predicted population provided by the customer company or the user data of the to-be-predicted population obtained by other means is obtained for subsequent prediction.

[0053] S109: Model prediction, output population package.

[0054] It should be noted that the trained Look-alike model is used to predict and sort the to-be-predicted population, and according to the customer's demand, the population package with high demand for target goods is output.

[0055] In the related art, for the Look-alike model, the traditional data sampling processing scheme has certain limitations, such as random sampling in the data sampling process, no distinction between samples, which leads to easy overfitting and underfitting of the model, affecting the final effect; in data sampling optimization, generally remove noise sample data or strengthen the role of typical positive and negative samples, ignoring other data distribution, so although it can improve the accuracy of the model under experimental conditions, but in the actual application environment, it deviates from the actual data distribution, so that the actual application generalization performance of the model is not good. In addition to this, for some neural network models with relatively simple structure (such as similar person group expansion model), generally single model, there is a great risk of overfitting on small data sets, and the above reasons all lead to poor prediction effect of the trained neural network model.

[0056] The embodiment of the application provides a model training method, and the basic idea of the method is: obtaining at least two sample data sets; wherein the difficulty ratio of the sample data included in each of the at least two sample data sets is different; training at least two preset models through the at least two sample data sets respectively to obtain at least two training models; and model fusion is performed on the at least two training models to obtain a target model. In this way, by using at least two sample data sets with different difficulty ratios, hierarchical sampling of sample data is realized, thereby improving the prediction effect of the finally obtained target model; moreover, without using at least two sample data sets to train models respectively and then obtaining a target model through model fusion, the role of typical positive and negative samples is excessively strengthened, the sample data is fully utilized, and the generalization ability of the target model is also improved; in addition, by training at least two models for fusion, the problem of overfitting of a single model can also be avoided.

[0057] The embodiments of the application will be described in detail below with reference to the accompanying drawings.

[0058] In an embodiment of the application, see Figure 2 which shows a flowchart of a model training method provided by an embodiment of the application. As Figure 2 shown, the method can include:

[0059] S201: Obtain at least two sample data sets.

[0060] It should be noted that the embodiment of the application provides a model training method, which can be applied to ranking models, especially Look-alike models. In addition, the idea of the model training method can also be extended to gesture recognition models, security detection types and models in other fields.

[0061] It should be noted that the model training process refers to determining a plurality of parameters of a preset model by using a large amount of sample data, so that the trained model can calculate new data to obtain a prediction result of the new data. In the related technical solution, all sample data are treated equally, but the effect of model training is not good. In the model training process, some sample data are easy for the preset model to learn (that is, the preset model can easily extract the features of the sample data), and this part of sample data can be referred to as easy sample; and some other data are difficult for the preset model to learn (that is, the preset model cannot easily extract the features of the sample data), and this part of sample data can be referred to as hard sample; the sample between the easy sample and the hard sample can be referred to as medium sample. Therefore, in order to distinguish the easy sample, the hard sample and the medium sample in the model training process, the present application embodiment introduces a parameter of "training difficulty category" for each sample data. That is, the training difficulty category is used to indicate the difficulty level of training the sample data.

[0062] Based on such an idea, in the present application embodiment, at least two sample data sets need to be obtained, and the specific number of sample data sets can be determined according to actual application scenarios. In addition, the sample data included in each of the at least two sample data sets has different difficulty ratios, and the difficulty ratio refers to the number ratio of sample data under different training difficulty categories. For example, the training difficulty categories include category A and category B, the at least two sample data sets include sample data set 1 and sample data set 2, the sample data set 1 includes 100 sample data of category A and 200 sample data of category B, and the data difficulty ratio of the sample data set 1 is 100:200; the sample data set 2 includes 200 sample data of category A and 100 sample data of category B, and the data difficulty ratio of the sample data set 2 is 200:100.

[0063] That is, for the sample data set, the sample data contained therein comes from different training difficulty categories (or different levels), thereby realizing hierarchical sampling of sample data, avoiding treating all sample data equally, and thereby improving the effect of model training.

[0064] Further, in some embodiments, for obtaining the at least two sample data sets, as shown in Figure 3 The step can include S301-S303. Details are as follows:

[0065] S301: Obtain a plurality of sample data, and calculate data scores of the plurality of sample data respectively;

[0066] It should be noted that, since the at least two sample data sets each contain different proportions of data difficulty, sampling needs to be performed according to the difficulty of the sample data, so as to form the at least two sample data sets.

[0067] Specifically, first, according to the original plurality of sample data, the data score of the plurality of sample data is calculated, and the data score represents the difficulty of the sample data in training; second, according to the data score of each sample data, the training difficulty category to which the sample data belongs is determined; finally, according to the training difficulty category to which the sample data belongs, data stratified sampling is performed, so as to obtain at least two sample data sets with different proportions of data difficulty.

[0068] That is, the data score of the sample data is used to indicate the difficulty of the sample data in training, so that the training category of the sample data can be determined, and at least two sample data sets with different proportions of difficulty can be determined.

[0069] Further, in some embodiments, the calculation of the data score of each of the plurality of sample data can include:

[0070] Grouping the plurality of sample data to obtain N to-be-calculated data sets; wherein N is an integer greater than or equal to 2;

[0071] From the N to-be-calculated data sets, a first to-be-calculated data set and a second to-be-calculated data set are determined; wherein the first to-be-calculated data set refers to any one of the N to-be-calculated data sets, and the second to-be-calculated data set includes all to-be-calculated data sets in the N to-be-calculated data sets except the first to-be-calculated data set;

[0072] Training the preset score model using the second to-be-calculated data set to obtain a target score model;

[0073] Model testing is performed on the target score model using the first to-be-calculated data set to determine the data score of each sample data in the first to-be-calculated data set;

[0074] In the determination of the data score of each sample data in the N to-be-calculated data sets.

[0075] It should be noted that in order to determine the data score of the sample data, all the sample data is grouped to obtain N data sets to be calculated. Here, N is an integer greater than or equal to 2, and the specific value can be determined according to the actual use scene, for example, N can be 10, 20, and the embodiment of the present application takes N=10 for subsequent description. In addition, the grouping process of the sample data can be random, but it is better to ensure that the proportion of positive and negative samples in each data set to be calculated is approximately the same, that is, the proportion of positive and negative samples in all data sets to be calculated is the same value (such as 3:7).

[0076] After obtaining the N data sets to be calculated, determining the data score of each sample data can include the following steps:

[0077] (1) In the N data sets to be calculated, one of the data sets to be calculated is determined as the first data set to be calculated, and the remaining data sets to be calculated are determined as the second data set to be calculated;

[0078] (2) training the preset score model using the sample data in the second data set to be calculated to obtain a target score model;

[0079] (3) testing the target score model using the sample data in the first data set to be calculated, and determining the data score of each sample data in the first data set to be calculated according to the test result.

[0080] Specifically, model testing refers to using the target score model to calculate the sample data in the first data set to be calculated, and outputting the predicted value of the sample data. Then, by comparing the predicted value of the sample data with the true label value of the sample data, the data score of the sample data is determined. Here, the true label value is used to indicate whether the sample data is a positive sample or a negative sample.

[0081] In the N data sets to be calculated, each data set to be calculated is sequentially determined as the first data set to be calculated, and then the data score of each sample data in the first data set to be calculated is determined through the foregoing steps (1)-(3), so that the data score of each sample data in the N data sets to be calculated can be determined.

[0082] It should be noted that the target score model can include M target score sub-models. Training the preset score model using the second data set to be calculated to obtain a target score model can include:

[0083] grouping the second data set to be calculated to obtain M data subsets to be calculated; wherein M is an integer greater than or equal to 1;

[0084] The M target scoring sub-models are obtained by training the preset scoring model using the M data subsets to be calculated.

[0085] It should be noted that, in order to improve the accuracy of data scoring, the target scoring model can include M target scoring sub-models. M is an integer greater than or equal to 1, and the specific value of M can be determined according to the actual application scenario, for example, M can be 10, 20.

[0086] When the target scoring model includes M target scoring sub-models, the second data set to be calculated needs to be grouped to obtain M training data subsets, and the M training data subsets are used to train the preset scoring model respectively, so as to obtain the M target scoring sub-models.

[0087] Taking M=10 as an example, the second data set to be calculated is grouped into training data subset 1, training data subset 2, …, and training data subset 10. Then, the training data subset 1 is used to train the preset scoring model to obtain the target scoring sub-model 1; the training data subset 2 is used to train the preset scoring model to obtain the target scoring sub-model 2, …, and the training data subset 1 is used to train the preset scoring model to obtain the target scoring sub-model 10. Finally, the target scoring sub-model 1, the target scoring sub-model 2, …, and the target scoring sub-model 10 constitute the target scoring model. Similarly, the positive and negative sample data ratios in the M training data subsets are preferably approximate.

[0088] In this way, by grouping the second data set to be calculated, M target scoring sub-models are trained and obtained, and then the M target scoring sub-models are used to test the sample data in the first data set to be calculated to determine the data scores of the sample data in the first data set to be calculated.

[0089] Further, in some embodiments, the model testing of the target scoring model using the first data set to be calculated to determine the respective data scores of each sample data in the first data set to be calculated can include:

[0090] The test sample data is input into the M target scoring models, and M model test results are output; wherein the test sample data refers to any one sample data in the first data set to be calculated;

[0091] Based on the M model test results, the data score of the test sample data is determined.

[0092] It should be noted that the to-be-tested sample data is input as an input value into the M target scoring sub-models to obtain M model test results. Here, the to-be-tested sample data refers to any one sample data in the first to-be-calculated data set. Then, according to the M model test results, the data score of the to-be-tested sample data is determined.

[0093] Further, the determining of the data score of the to-be-tested sample data based on the M model test results can include:

[0094] determining a maximum value, a minimum value, a median value, an average value, and a standard deviation from the M model test results, and determining a true label value of the to-be-tested sample data;

[0095] calculating an absolute value of a difference between the average value and the true label value to obtain a first difference value;

[0096] calculating an absolute value of a difference between the median value and the true label value to obtain a second difference value;

[0097] calculating a difference between the maximum value and the minimum value to obtain a third difference value;

[0098] performing weighted summation calculation on the first difference value, the second difference value, the third difference value, and the standard deviation to obtain the data score of the to-be-tested sample data.

[0099] It should be noted that for the to-be-tested sample data, the smaller the gap between the model test result and the true label value, the easier the to-be-tested sample data is to learn (equivalent to the to-be-tested sample data being easier to learn during model training). Therefore, the calculation of the data score of the to-be-tested sample data can include the following steps:

[0100] (1) For the to-be-tested sample data, calculate a maximum value (max), a minimum value (min), a median value (median), an average value (avg), and a standard deviation (std) in the M model test results; in addition, obtain a true label value (label) of the to-be-tested sample data. Here, the true label value is used to indicate whether the to-be-tested sample data belongs to positive sample data or negative sample data.

[0101] (2) Calculate an absolute value of a difference between the average value (avg) and the true label value (label), denoted as a first difference value; calculate an absolute value of a difference between the median value (median) and the true label value (label), denoted as a second difference value; and calculate a difference between the maximum value (max) and the minimum value (min), denoted as a third difference value.

[0102] (3) the first difference value, the second difference value, the third difference value and the standard deviation (std) are respectively weighted and summed, and finally the data score of the to-be-tested sample data is obtained. Here, the weight of each of the first difference value, the second difference value, the third difference value and the standard deviation is determined according to the actual application scenario, and the embodiment of the present application is not limited here.

[0103] The smaller the gap between the model test result and the true label value, the easier the sample data is to learn, and the smaller the data score, indicating that the sample data is easy to learn, while the larger the data score, the more difficult the sample data is to learn.

[0104] In this way, through the above processing steps, the data score of each sample data is determined.

[0105] S302: based on the data score of each of the plurality of sample data, determine the training difficulty category to which the plurality of sample data belongs.

[0106] It should be noted that according to the data score of the sample data, the training difficulty category of each sample data can be further determined. Here, the higher the data score of the sample data, the greater the training difficulty of the sample data, and the specific corresponding rule between the data score and the training difficulty category can be determined according to the actual use environment, and the embodiment of the present application is not limited here.

[0107] In a specific embodiment, the training difficulty category is divided into three categories, namely simple category, regular category and difficult category. The determination of the training difficulty category of each of the plurality of sample data based on the data score of each of the plurality of sample data can include:

[0108] determine a first score threshold and a second score threshold; wherein the first score threshold is less than the second score threshold;

[0109] if the data score of one of the sample data is less than the first score threshold, the training difficulty category of the one of the sample data is determined as the simple category;

[0110] if the data score of one of the sample data is greater than or equal to the first score threshold and less than the second score threshold, the training difficulty category of the one of the sample data is determined as the regular category;

[0111] if the data score of one of the sample data is greater than or equal to the second score threshold, the training difficulty category of the one of the sample data is determined as the difficult category.

[0112] It should be noted that by using the first score threshold and the second score threshold, the sample data can be divided into the simple category, the regular category and the difficult category.

[0113] Specifically, if the data score of the sample data is less than the first score threshold, it is determined that the training difficulty category of the sample data is a simple category; if the data score of the sample data is greater than or equal to the first score threshold but less than the second score threshold, it is determined that the training difficulty category of the sample data is a regular category; and if the data score of the sample data is greater than or equal to the second score threshold, it is determined that the training difficulty category of the sample data is a difficult category.

[0114] Here, the first score threshold and the second score threshold are preset according to actual application scenarios. For example, the data scores of all sample data can be sorted from small to large, and in the sorted data scores, the data score at the 30% position is determined as the first score threshold, and the data score at the 70% position is determined as the second score threshold. In this way, all sample data can be conveniently classified into simple categories (30%), regular categories (50%), and difficult categories (30%) according to the proportion, that is, the sample data is stratified.

[0115] In this way, the sample data and the respective training difficulty categories of the sample data are obtained, so that subsequent training is performed according to the respective training difficulty categories of the sample data.

[0116] S303: stratified sampling of the plurality of sample data based on the training difficulty categories to which the plurality of sample data belong, to determine the at least two sample data sets.

[0117] It should be noted that in the related art, whether the training difficulty of the sample data is easy or not, the status of all sample data in training is the same, which leads to poor effect of model training. Therefore, in the embodiments of the present application, the at least two sample data sets are formed by stratified sampling according to the training difficulty categories of the sample data, so that the difficulty proportions of different sample data sets are different, and subsequent training can be performed according to the characteristics of the sample data.

[0118] In this way, at least two sample data sets are determined from the plurality of sample data by stratified sampling, and the sample data is optimized by stratified sampling, thereby improving the effect of model training.

[0119] S202: training at least two preset models by using the at least two sample data sets respectively, to obtain at least two trained models.

[0120] It should be noted that the at least two preset models are trained respectively based on the at least two sample data sets obtained in the foregoing, so that two training models are obtained correspondingly. Here, the number of sample data sets and the number of preset models are one-to-one corresponding, and one sample data set is used to train one preset model. In addition, the at least two preset models can be models of the same architecture or models of different architectures, and the embodiments of the present application do not limit this.

[0121] S203: Model fusion is performed on the at least two training models to obtain a target model.

[0122] It should be noted that the at least two training models are fused to obtain a target model. Here, the method of model fusion can refer to the existing model fusion methods, such as average fusion, weighted average fusion, supervised model fusion (such as blending and stacking), and the like.

[0123] It should also be noted that the embodiments of the present application can fuse any number of training models. In a specific embodiment, two different sample data sets can be determined, two different training models can be trained, and a target model can be obtained by fusion.

[0124] In this case, the at least two sample data sets include a first sample data set and a second sample data set; and the stratified sampling of the plurality of sample data based on the training difficulty categories to which the plurality of sample data belong to, to determine the at least two sample data sets, can include:

[0125] The simple sample data set, the regular sample data set, and the difficult sample data set are sampled, the first preset proportion of the simple sample data set, the second preset proportion of the regular sample data set, and the third preset proportion of the difficult sample data set obtained by sampling are determined as the first sample data set, and the fourth preset proportion of the simple sample data set, the fifth preset proportion of the regular sample data set, and the sixth proportion of the difficult sample data set obtained by sampling are determined as the second sample data set.

[0126] The simple sample data set includes sample data of all training difficulty categories of simple categories in the plurality of sample data, the regular sample data set includes sample data of all training difficulty categories of regular categories in the plurality of sample data, and the difficult sample data set includes sample data of all training difficulty categories of difficult categories in the plurality of sample data.

[0127] It should be noted that first, the plurality of sample data is divided into a simple sample data set, a regular sample data set and a difficult sample data set. Specifically, the simple sample data set includes sample data of all training difficulty categories in the plurality of sample data Simple category, the regular sample data set includes sample data of all training difficulty categories in the plurality of sample data Regular category, and the difficult sample data set includes sample data of all training difficulty categories in the plurality of sample data Difficult category. This process is actually a sample stratification process, and subsequent sample data of a certain proportion is taken from the simple sample data set, the regular sample data set and the difficult sample data set to form at least two sample data sets, thereby realizing stratified sampling

[0128] Secondly, the simple sample data set, the regular sample data set and the difficult sample data set are adopted, and the simple sample data set of the first preset proportion value, the regular sample data set of the second preset proportion value and the difficult sample data set of the third preset proportion value are determined as the first sample data set; the simple sample data set of the fourth preset proportion value, the regular sample data set of the fifth preset proportion value and the difficult sample data set of the sixth preset proportion value are determined as the second sample data set.

[0129] The first preset proportion value, the second preset proportion value, the third preset proportion value, the fourth preset proportion value, the fifth preset proportion value and the sixth preset proportion value can be determined according to actual application scenarios. In a specific embodiment, the first preset proportion value is 100%, the second preset proportion value is 80%, the third preset proportion value is 20%, the fourth preset proportion value is 20%, the fifth preset proportion value is 80%, and the sixth preset proportion value is 100%. At this time, the first sample data set includes all simple sample data, 80% of the regular sample data and 20% of the difficult sample data; the second sample data set includes 20% of the simple sample data, 80% of the regular sample data and all of the difficult sample data.

[0130] It should be noted that in order to further implement the idea of stratified sampling, the specific sample data set (simple sample data set / regular sample data set / difficult sample data set) can be divided into multiple levels when sampling, and a certain proportion of sample data is sampled at each level. The sampling proportions of the multiple levels are added to the preset proportion value of the sample data set.

[0131] For example, when 80% of the regular dataset needs to be sampled, the regular dataset is divided into three levels according to the data score size, 24% of the data is sampled in the first level, 32% of the data is sampled in the second level, and 24% of the data is sampled in the third level; when 20% of the difficult dataset needs to be sampled, the difficult dataset is divided into three levels according to the data score size, 6% of the data is sampled in the first level, 8% of the data is sampled in the second level, and 6% of the data is sampled in the third level.

[0132] In this way, for the first sample dataset and the second sample dataset, the first sample dataset is more biased towards simple sample data, and the second sample dataset is more biased towards difficult sample data,

[0133] When the at least two sample datasets include the first sample dataset and the second sample dataset, correspondingly, the at least two preset models can include a first preset model and a second preset model, and the at least two training models can include a first training model and a second training model; the training of the at least two preset models by using the at least two sample datasets respectively to obtain the at least two training models can include:

[0134] training the first preset model by using the first sample dataset to obtain the first training model;

[0135] training the second preset model by using the second sample dataset to obtain the second training model.

[0136] It should be noted that the first preset model is trained by using the first sample dataset to obtain the first training model, and the second preset model is trained by using the second sample dataset to obtain the second training model. Here, the first preset model, the second preset model and the aforementioned preset score model can be models of the same architecture, or can be different from each other, and the embodiments of the present application are not limited.

[0137] It should also be noted that in this case, the first training model is more biased towards feature extraction of simple sample data, and the second training model is more biased towards feature extraction of difficult sample data. Correspondingly, the model fusion of the at least two training models to obtain the target model can include:

[0138] determining the first training model as a reference model and the second training model as a model to be fused;

[0139] fusing the model to be fused and the reference model based on a preset fusion algorithm to obtain the target model.

[0140] It should be noted that, according to the foregoing, the first training model is more focused on the features of the easy sample, and the second training model is more focused on the features of the hard sample, so the first training model can be used as a benchmark model, and the second training model can be used as a to-be-fused model, and then the to-be-fused model and the benchmark model are fused according to a preset fusion algorithm (such as a blending algorithm or a stacking algorithm).

[0141] In this way, the target model is fused from the first training model and the second training model, which can better fuse the features of the simple sample data and the features of the difficult sample data, thereby improving the prediction effect and generalization ability of the model.

[0142] In addition, 5 different training data sets can also be determined according to 5 sample data sets, so as to train 5 different training models, and then the 5 different training models are fused to obtain the final target model. The number of sample data sets is not limited in the present application.

[0143] Further, in some embodiments, the method further comprises:

[0144] Obtaining user data of each of a plurality of to-be-predicted users;

[0145] Inputting the user data of each of the plurality of to-be-predicted users into the target model to obtain a prediction value of each of the plurality of to-be-predicted users;

[0146] Based on the prediction value of each of the plurality of to-be-predicted users, the plurality of to-be-predicted users are sorted, and at least one target user is determined from the plurality of to-be-predicted users according to the sorting result.

[0147] It should be noted that, taking the target model as an example, the model is used to find target users similar to the seed population, at this time, after the target model is trained according to the sample data, the user data of the plurality of to-be-predicted users is calculated respectively by using the target model to obtain the prediction value of each of the plurality of to-be-predicted users; the plurality of to-be-predicted users are sorted in descending order according to the prediction value, and the first K to-be-predicted users in the sorting are determined as target users, so as to determine target users similar to the seed population, so as to facilitate subsequent product recommendation for the target users, K is a positive integer, and the value of K is determined according to the actual application scenario.

[0148] Generally, the related art tends to improve the effect of the Look-alike model by more meticulous data preprocessing, more feature mining, more complex model structure, and more model fusion. The general process of more meticulous data preprocessing is to perform data preprocessing through more meticulous exploratory data analysis (EDA), to handle the missing measurement values and abnormal problems in the sample data, and then to perform more feature engineering in combination with some conclusions obtained through data analysis. Although this can improve the model effect to a certain extent, it often needs to spend a lot of time and effort, and the final code logic can be particularly complex, and more features can also increase a lot of calculation amount. In addition, more complex model structure and more model fusion are to improve the overall effect by model stacking, which increases a lot of unnecessary calculation amount, and thus reduces the performance of the model. In the embodiments of the present application, the sample is optimized from the perspective of sample data, which can better improve the prediction effect and generalization ability of the model.

[0149] In summary, in the real target population expansion business scenario, the sample quality provided by the customer is often uneven, and the information amount contributed by different quality samples also has a large difference. In the related technical solution, all samples are treated equally, and the model trained in this way often cannot achieve the optimal effect. The embodiments of the present application improve the overall effect of the model system from the perspective of sample optimization and model hierarchical training fusion. In the embodiments of the present application, all samples are layered, but in the actual business scenario, most of the samples are negative samples, so the sample optimization can also be called negative sample layering. Therefore, the embodiments of the present application include the following two points: (1) improving the effect of the model according to the negative sample layering optimization strategy; and (2) improving the generalization of the model through the multi-model hierarchical training and fusion strategy.

[0150] The embodiments of the present application provide a model training method. At least two sample data sets are obtained. Each of the at least two sample data sets includes sample data with different difficulty ratios. At least two preset models are trained through the at least two sample data sets respectively, to obtain at least two training models. The at least two training models are fused to obtain a target model. In this way, by using the at least two sample data sets with different difficulty ratios, the sample data is layered, so as to improve the prediction effect of the target model obtained finally. Moreover, the at least two sample data sets are used for model training respectively, and then the target model is obtained through model fusion, which also improves the generalization ability of the target model.

[0151] In another embodiment of the present application, referring to Figure 4Fig. 4 shows a flowchart of another model training method provided by the embodiments of the present application. As shown in Fig. 4, the method can include the following steps. Figure 4

[0152] S401: Determine the training difficulty category of the sample data.

[0153] It should be noted that in the related art, the focus of model optimization is generally placed on feature engineering and model stacking, such as digging more features, using more complex models, or using multiple models for multi-layer fusion to improve the final model effect. The embodiments of the present application focus on the sample level and the model layered training optimization strategy based on sampling optimization, which can better improve the model effect without making the model too complex.

[0154] Specifically, in the process of model training, some sample data is easy for the model to learn, i.e., the trained model can easily make accurate predictions on these sample data, which is referred to as easy sample (equivalent to the training difficulty category of the sample data being simple); however, some samples may be difficult for the model, i.e., the trained model cannot accurately predict these sample data, which is referred to as hard sample (equivalent to the training difficulty category of the sample data being regular); the sample data between easy sample and hard sample is referred to as medium sample (equivalent to the training difficulty category of the sample data being difficult). In addition, there are two extreme cases: if some samples mix in label information, it is like "peeking at the correct answer" for the model, and such samples will be easily "learned" by the model but have no meaning, which is referred to as easy bad sample. On the contrary, another case is that the features of the sample data are completely random signals, and the model cannot learn any useful information, which is referred to as hard bad sample. For easy bad sample, if the label leakage problem cannot be solved, these samples will affect the training effect of the model and need to be discarded. For hard bad sample, more exploratory data analysis and feature mining are needed to make them available samples, otherwise these samples have no distinguishing degree for the model and are random noise. For easy bad sample and hard bad sample, it is best to distinguish and process them in the data preprocessing step, and it is temporarily assumed that the sample data in the embodiments of the present application does not contain the two cases.

[0155] ​The final purpose of model training is to accurately predict in real business scenarios. Therefore, during the model training stage, the model needs to learn from samples in real business scenarios. If we treat different difficulty levels of samples equally during the training stage, the model is likely to "shirk work" and only learn from easy samples. In this way, the model trained in this way is likely to accurately predict easy samples, but in real scenarios, there are medium samples and hard samples, and the model may not perform well. The industry usually uses feature stacking and model fusion to improve the overall effect of the model, which can solve the problem to some extent, but it will inevitably increase the amount of calculation. The embodiments of the present application optimize from the sample perspective, which can effectively improve the model effect, and can be further combined with feature and model optimization to obtain better results.

[0156] Based on this idea, the specific method of the embodiments of the present application can be divided into two steps: sample definition of different levels, hierarchical sampling combination and model training and fusion.

[0157] Therefore, for sample data used for model training, it is necessary to determine the training difficulty category (also referred to as difficulty level) of the sample data. The training difficulty category represents the difficulty of the sample data in the model training process, and also represents the information contribution of the sample data in the model training process, and also represents the accuracy of the trained model in predicting the sample data. Specifically, the training difficulty category includes three categories of simple, regular and difficult, so the sample data can be divided into three categories of easy sample, medium sample and hard sample.

[0158] It should be noted that the embodiments of the present application obtain the data score of the sample data by using two nested 10-fold modeling schemes, and then determine the training difficulty category of the sample data according to the data score.

[0159] First, the data score of the sample data is obtained by using two nested 10-fold modeling schemes. Referring to Figure 5 , a working process schematic diagram of a model training method provided by an embodiment of the present application is shown. As Figure 5 indicated, the model for determining the data score is divided into an outer layer 10-fold cross-validation model and an inner layer 10-fold cross-validation model.

[0160] For the outer ten-fold cross-validation model, first, the original sample data set (or called training set) is randomly divided into 10 parts (the proportion of positive and negative samples in each part is approximately); second, one of the 10 parts (ten-fold) is selected as the test set (equivalent to the first data set to be calculated) and the remaining 9 parts are used as the training set (equivalent to the second data set to be calculated); then, the training set is used to train the preset scoring model, and the trained model (i.e. the target scoring model) is used to make predictions in the test set; finally, for the test sample data in the test set, the data score of the test sample data can be obtained according to the true label value and the model preset value of the test sample data (the score value of the binary classification model is distributed between 0 and 1). In this way, each part of the ten-fold is used as the test set in turn, so that the model score corresponding to each sample data is obtained.

[0161] Specifically, when the training set of 9 parts is used to train the preset model, an inner ten-fold cross-validation model is used. Specifically, the training set (i.e. the outer 9 parts) is further divided into an inner 10-fold sample (randomly stratified, and the proportion of positive and negative samples in each part is close), and each part of the sample is used as a training set to train the preset model, so that 10 trained target scoring sub-models are obtained; each target scoring sub-model is tested in the corresponding test set (i.e. the outer 1 part), i.e. each sample data in the test set is predicted 10 times, so that 10 data scores corresponding to each sample data in the test set are obtained, and finally a prediction matrix P is obtained. m×n where m represents the number of sample data, and n represents the data score of each sample data, which is 10 here. The value can be increased to make the result more reliable, but the cost is an increase in the amount of calculation.

[0162] As shown in Figure 5 , the sample data set is randomly divided into an outer ten-fold, and the first to ninth parts are used as the training set (train) and the tenth part is used as the test set. Taking the tenth part as the test set as an example, when the training set (train) is used for training, the training set (train) is further divided into an inner ten-fold, and each part of the training set is used as a separate training set to train the preset model, so that ten trained preset models are obtained. Then, the ten trained preset models are used to test the test set 10 (test-10-pred) respectively, and 10 data scores are obtained, i.e. the test set 10 scores 1-10 (pred-1-pred-10). According to the above, each sample data will obtain its corresponding 10 data scores, thereby forming a prediction matrix P m×n .

[0163] After obtaining the prediction matrix P m×nAfter that, according to the prediction matrix P m×n , the data score [p i1 …p in ] corresponding to one of the sample data is calculated Some statistical indicators: mean avg, median, standard deviation std, maximum max and minimum min, and then calculate the data score S of each sample data according to formula (1):

[0164] S = a · |label-avg| + b · |label-median| + g · std + d · (max-min) … … (1)

[0165] Where a, b, g and d are hyperparameters, and label is the true label value of the sample data, which can be adjusted according to different experiments.

[0166] Here, the smaller the value of the data score S of the sample data, the easier it is for the model to learn. The reason for this is that if the difference between the true label value (label) and the mean (avg) and median (median) of the predicted value is larger, the model's prediction deviation is larger, indicating that the model is not good enough to learn the sample data, so the sample weight should be increased during the later model training. If the standard deviation (std) of the predicted value and the difference between the maximum and minimum values (max-min) are larger, it means that different sample data has a greater impact on the model's prediction results.

[0167] In practice, it is found that in different business scenarios, the main indicators that affect sample definition will be different, so the weights are dynamically adjusted according to the hyperparameters a, b, g, and d. For example, in some business scenarios, the values are: a = 0.35, b = 0.15, g = 0.35, d = 0.15, or a = 0.25, b = 0.25, g = 0.15, d = 0.35.

[0168] Finally, calculate the 30th percentile (P1) and 70th percentile (P2) of the score S of all samples, and then define the training difficulty category of the sample data according to the following rules: if the sample data S is less than P1, it is recorded as easy sample; if the sample data S is greater than or equal to P1 and less than P2, it is recorded as medium sample; if the sample data S is greater than or equal to P2, it is recorded as hard sample, as shown in formula (2).

[0169]

[0170] In this way, through the above calculation, the training difficulty category of the sample data is determined.

[0171] S402: stratified sampling the sample data according to the training difficulty categories of the sample data to obtain a first sample data set and a second sample data set.

[0172] It should be noted that after determining the training difficulty categories of the sample data, the sample data can be stratified sampled according to the training difficulty categories of the sample data. Referring to Figure 6 , which shows a working process schematic diagram of another model training method provided by the embodiments of the present application. As shown in Figure 6 , first, sample all easy samples, 80% of medium samples and 20% of hard samples (the sampling ratio can be adjusted, and sampling is performed through quantile values) to form a first sample data set, and then sample all hard samples, 80% of medium samples and 20% of easy samples to form a second sample data set.

[0173] As shown in Figure 6 , for a regular data set, it is divided into three levels again according to the data score of each sample data, and then 24%, 32% and 24% of each level are sequentially divided, so as to obtain 80% of the medium sample. For a difficult data set, it is divided into three levels again according to the data score of each sample data, and then 6%, 8% and 6% of each level are sequentially divided, so as to obtain 20% of the hard sample. For a simple data set, it is divided into three levels again according to the data score of each sample data, and then 6%, 8% and 6% of each level are sequentially divided, so as to obtain 20% of the easy sample.

[0174] In this way, stratified sampling optimization can be better achieved, thereby improving the prediction effect of the target model.

[0175] S403: model training and model fusion are performed using the first sample data set and the second sample data set to obtain a target model.

[0176] It should be noted that, as shown in Figure 6 , model training using the first sample data set can obtain a benchmark training model Y1 (equivalent to the first training model described above), which will be more fully trained for easy samples and medium samples, ensuring that the prediction result will not deviate greatly. In addition, model training using the second sample data set can obtain a training model Y2 (equivalent to the second training model described above) that pays more attention to hard samples.

[0177] After obtaining the training model Y1 and the training model Y2, a supervised model fusion technology such as a blending and stacking model fusion technology is used to fuse the training model Y1 and the training model Y2 to obtain a final result. Here, the supervised fusion effect is better, and the simple weighted average fusion result deviation is slightly large, but is still better than the existing model training method. In a specific embodiment, the original sample data set can be used as supervision in the model fusion process, or another collected sample data set can be used for model fusion.

[0178] In addition, in order to further improve the optimization effect, more different proportion sampling and model fusion can be performed. Here, only as an illustration, in general, training the two models can achieve good results.

[0179] In summary, in the related art, the optimization of the data sampling processing method is mainly concentrated in feature selection, data balancing, typical sample sampling, and random sampling, and many of them are still limited to the optimization of sampling classification effect under laboratory conditions. For actual application systems, the correlation between the distribution of data and the final prediction target result is more complex, and the data sampling processing method in the traditional Look-alike prediction problem cannot necessarily achieve the optimal result when applied to an actual system.

[0180] The technical key point that the embodiments of the present application hope to protect is a Look-alike model scheme based on sample hierarchical sampling optimization and model hierarchical training. In a real business scenario, the quality of sample data is uneven, and the amount of information that can be contributed also has a large difference. The model training method in the related art treats all samples equally, and the model trained in this way often cannot achieve the optimal effect. The present model trains the model through two nested ten-fold cross-validation, obtains 10 prediction results for each sample through different sample combinations, calculates the index S for positioning the level of the sample through statistical analysis of the prediction results, and trains and fuses the hierarchical sampling model according to the index. The target model is obtained, and the prediction result of the business problem is obtained by using the target model, which can improve the prediction effect.

[0181] The embodiments of the present application provide a model training method. Through the detailed description of the foregoing embodiments by the embodiments, it can be seen that at least two sample data sets with different difficulty ratios are used to realize hierarchical sampling of sample data, thereby improving the prediction effect of the target model obtained finally; and the at least two sample data sets are used to train models respectively, and then the target model is obtained through model fusion, thereby improving the generalization ability of the target model.

[0182] In another embodiment of the present application, referring to Figure 7It shows a component structure schematic diagram of a model training apparatus 50 provided by an embodiment of the present application. As shown in Figure 7 The model training apparatus 50 comprises an obtaining unit 501, a training unit 502 and a fusion unit 503, wherein

[0183] The obtaining unit 501 is configured to obtain at least two sample data sets; wherein the at least two sample data sets each comprise sample data with different difficulty ratios;

[0184] The training unit 502 is configured to train at least two preset models respectively through the at least two sample data sets to obtain at least two training models;

[0185] The fusion unit 503 is configured to fuse the at least two training models to obtain a target model.

[0186] In some embodiments, the obtaining unit 501 is specifically configured to obtain a plurality of sample data, calculate data scores of the plurality of sample data respectively, determine training difficulty categories to which the plurality of sample data belong based on the data scores of the plurality of sample data respectively, and perform stratified sampling on the plurality of sample data based on the training difficulty categories to which the plurality of sample data belong to determine the at least two sample data sets.

[0187] In some embodiments, the obtaining unit 501 is further configured to group the plurality of sample data to obtain N to-be-calculated data sets; wherein N is an integer greater than or equal to 2; determine a first to-be-calculated data set and a second to-be-calculated data set from the N to-be-calculated data sets; wherein the first to-be-calculated data set refers to any one to-be-calculated data set of the N to-be-calculated data sets, and the second to-be-calculated data set comprises all to-be-calculated data sets except the first to-be-calculated data set in the N to-be-calculated data sets; train the preset score model by using the second to-be-calculated data set to obtain a target score model; perform model testing on the target score model by using the first to-be-calculated data set to determine data scores of each sample data in the first to-be-calculated data set respectively; and obtain data scores of the plurality of sample data respectively after determining the data scores of each sample data in the N to-be-calculated data sets.

[0188] In some embodiments, the target score model comprises M target score sub-models; the obtaining unit 501 is further configured to group the second to-be-calculated data set to obtain M to-be-calculated data subsets; wherein M is an integer greater than or equal to 1; and train the preset score model by using the M to-be-calculated data subsets to obtain the M target score sub-models.

[0189] In some embodiments, the obtaining unit 501 is further configured to input the to-be-tested sample data into the M target scoring models, and output M model test results; wherein the to-be-tested sample data refers to any one sample data in the first to-be-calculated data set; based on the M model test results, determine a data score of the to-be-tested sample data.

[0190] In some embodiments, the obtaining unit 501 is further configured to determine a maximum value, a minimum value, a median value, an average value, and a standard deviation from the M model test results, and determine a real label value of the to-be-tested sample data; calculate an absolute value of a difference between the average value and the real label value to obtain a first difference value; calculate an absolute value of a difference between the median value and the real label value to obtain a second difference value; calculate a difference between the maximum value and the minimum value to obtain a third difference value; and perform weighted summation calculation on the first difference value, the second difference value, the third difference value, and the standard deviation to obtain a data score of the to-be-tested sample data.

[0191] In some embodiments, the training difficulty categories include a simple category, a regular category, and a difficult category; the obtaining unit 501 is further configured to determine a first score threshold and a second score threshold; wherein the first score threshold is less than the second score threshold; if a data score of one of the sample data is less than the first score threshold, the training difficulty category of the one of the sample data is determined as the simple category; if a data score of one of the sample data is greater than or equal to the first score threshold and less than the second score threshold, the training difficulty category of the one of the sample data is determined as the regular category; and if a data score of one of the sample data is greater than or equal to the second score threshold, the training difficulty category of the one of the sample data is determined as the difficult category.

[0192] In some embodiments, the at least two sample data sets include a first sample data set and a second sample data set; the obtaining unit 501 is further configured to stratify the plurality of sample data based on training difficulty categories to which the plurality of sample data belong, to obtain a simple sample data set, a regular sample data set, and a difficult sample data set; sample the simple sample data set, the regular sample data set, and the difficult sample data set, and determine the simple sample data set of a first preset proportion, the regular sample data set of a second preset proportion, and the difficult sample data set of a third preset proportion obtained by sampling as the first sample data set, and determine the simple sample data set of a fourth preset proportion, the regular sample data set of a fifth preset proportion, and the difficult sample data set of a sixth proportion obtained by sampling as the second sample data set; wherein the simple sample data set includes sample data of all training difficulty categories of simple categories in the plurality of sample data, the regular sample data set includes sample data of all training difficulty categories of regular categories in the plurality of sample data, and the difficult sample data set includes sample data of all training difficulty categories of difficult categories in the plurality of sample data.

[0193] In some embodiments, the first preset proportion is 100%, the second preset proportion is 80%, the third preset proportion is 20%, the fourth preset proportion is 20%, the fifth preset proportion is 80%, and the sixth preset proportion is 100%.

[0194] In some embodiments, the at least two preset models include a first preset model and a second preset model, and the at least two training models include a first training model and a second training model; the training unit 502 is specifically configured to train the first preset model using the first sample data set to obtain the first training model, and train the second preset model using the second sample data set to obtain the second training model.

[0195] In some embodiments, the fusion unit 503 is specifically configured to determine the first training model as a reference model and the second training model as a to-be-fused model; fuse the to-be-fused model and the reference model based on a preset fusion algorithm to obtain the target model.

[0196] In some embodiments, as Figure 8As shown, the model training apparatus 50 further includes a prediction unit 504 configured to obtain user data of each of a plurality of to-be-predicted users; input the user data of each of the plurality of to-be-predicted users into the target model to obtain a prediction value of each of the plurality of to-be-predicted users; and sort the plurality of to-be-predicted users according to the prediction values of the plurality of to-be-predicted users, and determine at least one target user from the plurality of to-be-predicted users according to a sorting result.

[0197] It can be understood that, in this embodiment, the "unit" can be a partial circuit, a partial processor, a partial program or software, and of course can also be a module, and can also be non-modular. Moreover, the components in this embodiment can be integrated in one processing unit, or can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.

[0198] When the integrated unit is realized in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiment can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in the embodiment. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0199] Therefore, the embodiment provides a computer storage medium, which stores a model training program. The model training program is executed by at least one processor to implement the steps of the method in any one of the foregoing embodiments.

[0200] Based on the foregoing composition of the model training apparatus 50 and the computer storage medium, refer to Figure 9 which shows a specific hardware structure schematic diagram of the model training apparatus 50 provided by the embodiment of the application. As shown in Figure 9As shown, the model training apparatus 50 can include a communication interface 601, a memory 602 and a processor 603; each component is coupled together through a bus device 604. It can be understood that the bus device 604 is used to realize the connection communication between the components. In addition to including a data bus, the bus device 604 also includes a power bus, a control bus and a status signal bus. However, in order to clearly illustrate, all the buses are marked as the bus device 604 in the Figure 9 The communication interface 601 is used for receiving and sending signals in the process of transmitting information with other external network elements;

[0201] The memory 602 is used for storing a computer program capable of running on the processor 603;

[0202] The processor 603 is used for executing the following when running the computer program:

[0203] Obtaining at least two sample data sets; wherein each of the at least two sample data sets includes sample data with different difficulty ratios;

[0204] Training at least two preset models through the at least two sample data sets respectively to obtain at least two training models;

[0205] Model fusion is performed on the at least two training models to obtain a target model.

[0206] It is to be understood that the memory 602 in embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. In one embodiment, nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), used as external cache memory. By way of example, and not limitation, many forms of RAM are available, for example, static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), SynchBurst DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory 602 of the devices and methods described herein are intended to include, without being limited to, these and any other suitable types of memory.

[0207] The processor 603 can be an integrated circuit chip including a processing unit that is configured to process signals. In implementation, the steps of the above-described method can be completed by the integrated logic circuit of the processor 603 or by an instruction in a form of software. The processor 603 described above can be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods, steps and logical block diagrams disclosed in the embodiments of the present application can be implemented or executed by the processor 603. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code executed by the processor or a combination of hardware and software modules in the processor. The software module can reside in a storage medium such as a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an electrically programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), registers, or a storage device. The storage medium is located in the memory 602, and the processor 603 reads information in the memory 602 and combines the hardware to complete the steps of the above-described method.

[0208] It can be understood that the embodiments described in the present application can be implemented in hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing the functions described in the present application or a combination thereof.

[0209] For software implementation, the technologies described in the present application can be implemented by modules (for example, procedures, functions, and so on) for performing the functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0210] Optionally, as another embodiment, the processor 603 is further configured to execute the steps of the method of any one of the preceding embodiments when running the computer program.

[0211] Based on the above model training device 50 composition and hardware structure diagram. Referring to Figure 10 , which shows a model training device 70 provided by the embodiment of the application. As Figure 10 shown, the model training device 70 at least includes the model training device 50 of any one of the preceding embodiments.

[0212] For the model training device 70, by using at least two sample data sets with different difficulty ratios, hierarchical sampling of sample data is realized, thereby improving the prediction effect of the finally obtained target model; moreover, at least two sample data sets are used for model training respectively, and then the target model is obtained through model fusion, which also improves the generalization ability of the target model.

[0213] The above is only a preferred embodiment of the present application, and is not intended to limit the protection scope of the present application.

[0214] It should be noted that in the present application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0215] The above sequence number of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments.

[0216] The methods disclosed in the several method embodiments of the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0217] The features disclosed in the several product embodiments of the present application can be combined arbitrarily without conflict to obtain new product embodiments.

[0218] The features disclosed in the several method or device embodiments of the present application can be combined arbitrarily without conflict to obtain new method or device embodiments.

[0219] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, The method includes: Obtain at least two sample datasets; wherein the proportion of easy and difficult data in each of the at least two sample datasets is different; the sample datasets include user data of multiple users, and the user data includes user behavior data regarding products; At least two preset models are trained using the at least two sample datasets respectively to obtain at least two trained models; the preset models are preset similar population expansion models. The at least two trained models are fused to obtain the target model; the target model is an extended model of a target similar population. Acquire user data for each of the multiple users to be predicted, including user behavior data related to products; The user data of each of the multiple users to be predicted is input into the target model to obtain the predicted value of each of the multiple users to be predicted; Based on the predicted values ​​of the plurality of users to be predicted, the plurality of users to be predicted are sorted, and at least one target user is determined from the plurality of users to be predicted according to the sorting results, so as to recommend the product to the at least one target user; The acquisition of at least two sample datasets includes: Acquire multiple sample data and calculate the data score for each of the multiple sample data; Based on the data scores of the multiple sample data, the training difficulty category to which the multiple sample data belong is determined; Based on the training difficulty category to which the multiple sample data belong, stratified sampling is performed on the multiple sample data to determine the at least two sample datasets.

2. The model training method according to claim 1, characterized in that, The calculation of the data scores for each of the multiple sample data includes: The multiple sample data are grouped to obtain N datasets to be calculated; where N is an integer greater than or equal to 2. From the N datasets to be computed, a first dataset and a second dataset are determined; wherein, the first dataset refers to any one of the N datasets to be computed, and the second dataset includes all datasets to be computed except the first dataset. The preset scoring model is trained using the second dataset to be calculated to obtain the target scoring model; The target scoring model is tested using the first dataset to be calculated, and the data score of each sample data in the first dataset to be calculated is determined. After determining the data score for each sample data in the N datasets to be calculated, the data scores for each of the multiple sample data are obtained.

3. The model training method according to claim 2, characterized in that, The target scoring model includes M target scoring sub-models; the step of training the preset scoring model using the second dataset to be calculated to obtain the target scoring model includes: The second dataset to be calculated is divided into M subsets of data to be calculated; where M is an integer greater than or equal to 1. The preset scoring model is trained using the M subsets of data to be calculated to obtain the M target scoring sub-models.

4. The model training method according to claim 3, characterized in that, The step of testing the target scoring model using the first dataset to be calculated, and determining the data score for each sample in the first dataset to be calculated, includes: The test sample data is input into the M target scoring sub-models, and M model test results are output; wherein, the test sample data refers to any sample data in the first dataset to be calculated; Based on the test results of the M models, the data score of the sample data to be tested is determined.

5. The model training method according to claim 4, characterized in that, The process of determining the data score for the test sample data based on the test results of the M models includes: From the M model test results, determine the maximum value, minimum value, median value, mean value and standard deviation, and determine the true label value of the test sample data; Calculate the absolute value of the difference between the average value and the true label value to obtain the first difference; Calculate the absolute value of the difference between the median value and the true label value to obtain the second difference; Calculate the difference between the maximum value and the minimum value to obtain the third difference; The first difference, the second difference, the third difference, and the standard deviation are weighted and summed to obtain the data score of the test sample data.

6. The model training method according to claim 5, characterized in that, The training difficulty categories include easy, normal, and hard categories; determining the training difficulty category of each of the multiple sample data based on their respective data scores includes: Determine a first scoring threshold and a second scoring threshold; wherein the first scoring threshold is less than the second scoring threshold; If the data score of one of the sample data is less than the first score threshold, then the training difficulty category of the one sample data is determined to be the simple category; If the data score of one of the sample data is greater than or equal to the first score threshold and less than the second score threshold, then the training difficulty category of the one sample data is determined to be the regular category; If the data score of one of the sample data is greater than or equal to the second score threshold, then the training difficulty category of the one sample data is determined to be the difficult category.

7. The model training method according to claim 6, characterized in that, The at least two sample datasets include a first sample dataset and a second sample dataset; the step of stratifying the multiple sample data based on the training difficulty category to which the multiple sample data belong to, to determine the at least two sample datasets, includes: The multiple sample data are stratified based on the training difficulty category to which they belong, resulting in a simple sample dataset, a regular sample dataset, and a difficult sample dataset. The simple sample dataset, the regular sample dataset, and the difficult sample dataset are sampled. The simple sample dataset with a first preset ratio, the regular sample dataset with a second preset ratio, and the difficult sample dataset with a third preset ratio are determined as the first sample dataset. The simple sample dataset with a fourth preset ratio, the regular sample dataset with a fifth preset ratio, and the difficult sample dataset with a sixth preset ratio are determined as the second sample dataset. The simple sample dataset includes all sample data whose training difficulty category is simple, the regular sample dataset includes all sample data whose training difficulty category is regular, and the hard sample dataset includes all sample data whose training difficulty category is hard.

8. The model training method according to claim 7, characterized in that, The first preset ratio is 100%, the second preset ratio is 80%, the third preset ratio is 20%, the fourth preset ratio is 20%, the fifth preset ratio is 80%, and the sixth preset ratio is 100%.

9. The model training method according to claim 7, characterized in that, The at least two preset models include a first preset model and a second preset model, and the at least two training models include a first training model and a second training model; The step of training at least two preset models using the at least two sample datasets to obtain at least two trained models includes: The first preset model is trained using the first sample dataset to obtain the first trained model; The second preset model is trained using the second sample dataset to obtain the second trained model.

10. The model training method according to claim 9, characterized in that, The step of fusing the at least two trained models to obtain the target model includes: The first training model is determined as the baseline model, and the second training model is determined as the model to be fused. The target model is obtained by fusing the model to be fused and the benchmark model based on a preset fusion algorithm.

11. A model training device, characterized in that, The model training device includes an acquisition unit, a training unit, a fusion unit, and a prediction unit, wherein... The acquisition unit is configured to acquire at least two sample datasets; wherein the at least two sample datasets each contain sample data with different difficulty ratios; the sample datasets include user data from multiple users, and the user data includes user behavior data related to products; The training unit is configured to train at least two preset models using the at least two sample datasets respectively, to obtain at least two training models; the preset models are preset similar population expansion models. The fusion unit is configured to fuse the at least two trained models to obtain a target model; the target model is a target similarity group extended model. The prediction unit is configured to acquire user data of multiple users to be predicted, the user data including user behavior data related to products; input the user data of the multiple users to be predicted into the target model to obtain the prediction value of each of the multiple users to be predicted; sort the multiple users to be predicted based on the prediction value of each of the multiple users to be predicted, and determine at least one target user from the multiple users to be predicted according to the sorting result, so as to recommend the product to the at least one target user; The acquisition unit is specifically configured to acquire multiple sample data, calculate the data score of each of the multiple sample data; determine the training difficulty category to which the multiple sample data belong based on the data score of each of the multiple sample data; and perform stratified sampling on the multiple sample data based on the training difficulty category to which the multiple sample data belong to determine the at least two sample datasets.

12. A model training device, characterized in that, The model training device includes a memory and a processor; wherein... The memory is used to store computer programs that can run on the processor; The processor is configured to perform the steps of the method as described in any one of claims 1 to 10 when running the computer program.

13. A computer storage medium, characterized in that, The computer storage medium stores a model training program, which, when executed by at least one processor, implements the steps of the method as described in any one of claims 1 to 10.

14. A model training device, characterized in that, The model training device includes at least the model training apparatus as described in claim 11 or 12.

Citation Information

Patent Citations

  • Difficult sample mining method and device, electronic equipment and computer readable storage medium

    CN110956255A

  • Target detection model training method and device and storage medium

    CN111753870A