Data classification model training method, classification method, device, equipment and medium
By undersampling and oversampling the historical data, a balanced training set is formed and the classification model is iteratively trained, which solves the low accuracy problem caused by data imbalance in the existing technology, and improves the training effect and accuracy of the classification model.
Patent Information
- Application Number
- CN202210248165.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-03-14
AI Technical Summary
The prior art training results are poor due to data imbalance when training classification models and low classification accuracy. Especially in the telephone customer service scenario, there are very few complaint phones and many consultation phones, which leads to the model misidentifying the complaint phone as a consultation phone.
By dividing historical data samples into a collection of minority classes and a collection of majority classes, undersampling and oversampling are performed to form a balanced training set, and the classification model is iteratively trained until the preset accuracy conditions are met.
Through data balance processing, the training effect and classification accuracy of the classification model are improved, and the low accuracy problem caused by data imbalance in the prior art is overcome.
Smart Images

Figure CN114662580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a training method for a data classification model, a data classification method, an apparatus, a computer device and a storage medium. Background Art
[0002] The data classification problem is one of the most common problems in the field of machine learning. Existing commonly used classification models include, for example, logistic regression algorithm model, k-nearest neighbor algorithm model, decision tree algorithm model, and support vector machine algorithm model. As machine learning algorithms are applied in more and more application scenarios, some problems have arisen in the application of classification models. Among them, the poor training effect of the classification model due to unbalanced data leads to low classification accuracy of the trained classification model, and the impact of unbalanced data distribution on the classification effect is particularly significant. It is very difficult to obtain balanced distribution data in some specific application scenarios. For example, in the telephone customer service scenario, there are very few complaint calls and very many consultation calls, and the number of the two types of calls differs by hundreds or even thousands of times, which brings great difficulties to the training of customer complaint classification models. In the prior art, the classification model is directly trained using historical data. Since no processing is performed on the historical data used for training, the training effect is poor. The trained classification model will misidentify most of the complaint calls as consultation calls, and the classification accuracy is low. Therefore, how to overcome the problem of poor training effect and low classification accuracy of the trained classification model caused by unbalanced training data when training the classification model is a technical problem to be solved at present. Summary of the invention
[0003] Based on this, it is necessary to provide a data classification model training method, data classification method, device, computer equipment and storage medium to address the problems of poor training effect and low classification accuracy of the trained classification model caused by historical data imbalance when training the classification model.
[0004] A training method for a data classification model, comprising:
[0005] Dividing the multiple historical data samples acquired in advance into a minority class sample set and a majority class sample set;
[0006] Undersampling the majority class sample set to obtain an undersampled set;
[0007] Performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition;
[0008] Detecting whether the classification model that satisfies the first preset condition satisfies the second preset condition;
[0009] If the second preset condition is not met, oversampling the minority class sample set based on the classification model that meets the first preset condition, and adding the oversampled data samples to the training set;
[0010] A second iterative training is performed on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0011] A data classification method, comprising:
[0012] Obtain data to be classified;
[0013] The steps of the above-mentioned data classification model training method; and,
[0014] The data to be classified is classified using the data classification model that meets the second preset condition.
[0015] A training device for a data classification model, comprising:
[0016] A partitioning module, used for partitioning a plurality of pre-acquired historical data samples into a minority class sample set and a majority class sample set;
[0017] An undersampling module, configured to obtain an undersampling set by undersampling the majority class sample set;
[0018] A first iterative training module, used to perform a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set, to obtain a classification model that meets a first preset condition;
[0019] A detection module, used to detect whether the classification model that meets the first preset condition meets the second preset condition;
[0020] An oversampling module, configured to oversample the minority class sample set based on the classification model that satisfies the first preset condition if the second preset condition is not met, and add the oversampled data samples to the training set;
[0021] The second iterative training module is used to perform second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0022] A data classification device, comprising:
[0023] A module for acquiring data to be classified, used for acquiring data to be classified;
[0024] A training device for the above-mentioned data classification model; and,
[0025] The classification module is used to classify the data to be classified using the classification model that reaches the preset training stop condition.
[0026] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the above-mentioned data classification model training method and / or the steps of the above-mentioned data classification method.
[0027] A storage medium storing computer-readable instructions, which, when executed by one or more processors, causes the one or more processors to execute the steps of the above-mentioned data classification model training method and / or the steps of the above-mentioned data classification method.
[0028] The training method, device, computer equipment and storage medium of the data classification model divide the pre-acquired multiple historical data samples into a minority class sample set and a majority class sample set, undersample the majority class sample set to obtain an undersampled set, perform a first iterative training on the preset classification model based on the training set composed of the minority class sample set and the undersampled set to obtain a classification model that meets the first preset condition, detect whether the classification model that meets the first preset condition meets the second preset condition, and if it does not meet the second preset condition, oversample the minority class sample set based on the classification model that meets the first preset condition, add the oversampled data samples to the training set, perform a second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition; because the undersampled data and the oversampled data are used when training the classification model, the data used for training the classification model is well balanced, the training effect on the classification model is good, and the classification accuracy of the trained classification model is high, which overcomes the problems of poor training effect and low classification accuracy of the trained classification model caused by the imbalance of training data used when training the classification model in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0030] Figure 1 A diagram of an application environment of a training method for a data classification model provided in one embodiment;
[0031] Figure 2is a flowchart of a method for training a data classification model in one embodiment;
[0032] Figure 3 A flowchart of a method for training a data classification model for a specific example;
[0033] Figure 4 A structural block diagram of a training device for a data classification model provided in one embodiment;
[0034] Figure 5 FIG. 4 is a block diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0036] It is understood that the terms "first", "second", "third", etc. are only used to distinguish descriptions and should not be understood to indicate or imply relative importance. It should also be understood that although the terms "first", "second", "third", etc. are used in the text to describe various elements in some embodiments of the present application, these elements should not be limited by these terms. These terms are only used to distinguish various elements.
[0037] refer to Figure 1 As shown, the training method of the data classification model provided in the embodiment of the present application can be applied to Figure 1 In an application environment, the client can communicate with the server through a network. The server can divide the multiple historical data samples obtained from the client into a minority class sample set and a majority class sample set, undersample the majority class sample set to obtain an undersampled set, perform a first iterative training on the preset classification model based on the training set composed of the minority class sample set and the undersampled set, and obtain a classification model that meets the first preset condition, and then detect whether the classification model that meets the first preset condition meets the second preset condition. If the second preset condition is not met, the minority class sample set is oversampled based on the classification model that meets the first preset condition, and the data samples obtained by oversampling are added to the training set, and the classification model that meets the first preset condition is subjected to a second iterative training based on the updated training set to obtain a data classification model that meets the second preset condition. The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers.
[0038] Oversampling and undersampling are two common methods for processing unbalanced data. When training a classification model, the oversampling method repeats the minority data samples that account for a very small proportion many times to increase the number of data samples of this type, while the undersampling method randomly samples the majority data samples that account for a very large proportion to reduce the number of data samples of this type. Both methods can adjust the number of data samples so that data of different categories tend to be balanced. However, the inventors found that the traditional oversampling method randomly selects a number of minority data samples from the data set, copies them, and adds them to the data set, which easily causes the classification model to overfit these data samples, which is not conducive to the generalization of the classification model; the traditional undersampling method randomly discards some majority data samples, and these discarded data samples may contain important information. If the classification model loses this information, it cannot accurately identify the category.
[0039] refer to Figure 2 As shown, in one embodiment, a method for training a data classification model is proposed, which may include steps S10 to S60:
[0040] S10, dividing the plurality of historical data samples acquired in advance into a minority class sample set and a majority class sample set.
[0041] In some implementations, the plurality of pre-acquired historical data samples include two types of data samples; step S10 may include:
[0042] Respectively counting the number of the two types of data samples in the multiple historical data samples;
[0043] The quantities of the two data samples are compared, and the data samples with a smaller quantity are used to form the minority class sample set, and the data samples with a larger quantity are used to form the majority class sample set.
[0044] For example, the multiple data samples may include positive data samples and negative data samples, and each data sample belonging to the positive data sample is labeled with a first label, and each data sample belonging to the negative data sample is labeled with a second label. By counting the number of first labels and second labels, the minority class data samples and the majority class data samples can be determined. For example, the first label can be set to 0 and the second label can be set to 1. Assuming that the number of label 0 is a, the number of label 1 is b, and a is less than b, the above positive data samples are minority class data samples, and the negative data samples are majority class data samples.
[0045] Taking the telephone customer service scenario as an example, there are very few complaint calls and a lot of consultation calls. The number of these two types of calls differs by a hundred or even a thousand times. The multiple pre-acquired telephone customer service historical data samples are divided into a minority sample set and a majority sample set. The minority sample set is a set of complaint call data samples, and the majority sample set is a set of consultation call data samples. The complaint call data samples can be marked with label 0, and the consultation call data samples can be marked with label 1. By counting the number of label 0 and label 1, the number of complaint call data samples and the number of consultation call data samples can be determined.
[0046] S20: Under-sampling the majority class sample set to obtain an under-sampling set.
[0047] In some embodiments, step S20 may include:
[0048] A first number of majority class data samples are randomly undersampled from the majority class sample set to form an undersampled set; wherein the absolute value of the difference between the first number and the number of data samples in the minority class sample set is less than a preset threshold.
[0049] refer to Figure 3 As shown, in a specific example, let the majority class sample set be N, the minority class sample set be P, the under-sampling set be N0, and the preset under-sampling iteration number threshold be m under , the preset oversampling iteration threshold is m over .
[0050] In this specific example, undersampling from the majority class sample set to obtain an undersampled set may include:
[0051] A first number of majority class data samples are randomly undersampled from N to form a set N0, wherein the absolute value of the difference between the first number and the number of data samples in P is less than a preset threshold.
[0052] Randomly sample multiple majority class data samples from N that are close to the number of samples in P to form a set N0, where And |P|≈|N0|.
[0053] S30, performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition.
[0054] In some embodiments, the preset classification model may adopt a classification model of the prior art. The first preset condition is reaching a first preset training times threshold or reaching a first preset accuracy threshold; each iteration training in the first iteration training includes:
[0055] Training the current classification model using a training set consisting of the minority class sample set and the under-sampling set;
[0056] Determine whether the current training reaches a first preset training times threshold;
[0057] If the first preset training times threshold is not reached, the classification model after this training is used to perform classification prediction on the remaining data samples in the majority class sample set;
[0058] Determining whether the classification prediction result reaches a first preset accuracy threshold;
[0059] If the first preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the under-sampling set to obtain an updated under-sampling set; the updated under-sampling set is used for the next iterative training in the first iterative training.
[0060] In some embodiments, the method of using the trained classification model to perform classification prediction on the remaining data samples in the majority class sample set includes:
[0061] Using the classification model after the current training, predict the probability value of each remaining data sample in the majority class sample set belonging to the minority class sample set and the probability value of each remaining data sample in the majority class sample set;
[0062] The data sample with classification prediction error is a data sample whose probability value belonging to the minority class sample set is greater than the probability value belonging to the majority class sample set.
[0063] In the foregoing specific example, performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition may include:
[0064] Establish a misclassified sample set E N ; Among them, the initial misclassified sample set E N is an empty set;
[0065] Use P and N0 to train the preset classification model to obtain the trained classification model;
[0066] The trained classification model is used to predict the probability distribution of each data sample in the set N-N0 in different categories, and all the probability values of the minority class data sample categories greater than the preset probability threshold t N The data samples are added to the misclassified sample set E N ;
[0067] If the misclassified sample set Then stop undersampling; otherwise, merge E N and N0, and update N0 using the merged set; where the merge E N and N0 to obtain N0∪E N , and then use N0∪E N Update N0, that is, N0 = N0 ∪ E N ;
[0068] Determine whether the current undersampling times reaches the preset undersampling iteration times threshold m under If m is not reached under , then repeat the above training steps to continue training until the current number of undersampling reaches m under Stop training when
[0069] In this embodiment, the majority class data samples with a similar number of minority class data samples are randomly undersampled to form a class-balanced training set, and the preset classification model is trained using the training set. Then, data samples that are incorrectly predicted by the classification model are gradually added to the training set, and the majority class data samples that are difficult to classify are added to the training set. Therefore, the undersampling method tends to retain the majority class data samples that are difficult to classify. These data samples that are difficult to classify often carry important category information, and retaining these data samples that are difficult to classify is conducive to the classification model correctly predicting the majority class data samples.
[0070] S40: Detect whether the classification model that meets the first preset condition meets the second preset condition.
[0071] In some embodiments, the second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; step S40 includes:
[0072] Using the classification model that meets the first preset condition to perform classification prediction on the minority class sample set, to obtain a classification prediction result;
[0073] Compare the obtained classification prediction result with the second preset accuracy threshold to determine whether the classification prediction result reaches the second preset accuracy threshold;
[0074] If the second preset accuracy threshold is reached, it is determined whether the number of training times reaches the second preset training times threshold.
[0075] S50: If the second preset condition is not met, oversampling the minority class sample set based on the classification model that meets the first preset condition, and adding the oversampled data samples to the training set.
[0076] In some embodiments, oversampling the minority class sample set based on the classification model that meets the first preset condition includes: using the classification model that meets the first preset condition to perform classification prediction on the minority class sample set, and using data samples with classification prediction errors as oversampled data samples based on the classification prediction results.
[0077] S60. Perform a second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0078] In some embodiments, each iteration of the second iteration of the training includes:
[0079] Train the current classification model using the updated training set;
[0080] Determine whether the current training reaches a second preset training times threshold;
[0081] If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set;
[0082] Determining whether the classification prediction result reaches a second preset accuracy threshold;
[0083] If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
[0084] The second preset accuracy threshold may be, for example, 100%, or other accuracy values, and may be set according to actual needs.
[0085] In some embodiments, determining whether the classification prediction result reaches a second preset accuracy threshold includes:
[0086] According to the number of misclassified minority class data samples in the classification prediction result, it is determined whether the classification prediction result reaches a second preset accuracy threshold.
[0087] In the aforementioned example, performing a second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition may include:
[0088] Establish a minority class sample set P0 and initialize P0 with P, that is, P0 = P;
[0089] Establish a misclassified sample set E P ; Among them, the initial misclassified sample set EP is an empty set;
[0090] The classification model trained by P0 and N0 is used to predict each data sample in the set P, and all the probability values of the majority class data samples are greater than the threshold t P The data samples are added to the misclassified sample set E P ;
[0091] like Then stop oversampling; otherwise, E P The data samples in are added to P0;
[0092] Determine whether the current oversampling times reaches the preset oversampling iteration times threshold m over ; If the current oversampling times does not reach m over , then repeat the above steps until the current oversampling times reaches m over Stop when.
[0093] In this embodiment, all minority class data samples are predicted using a classification model that meets the first preset condition, and the data samples that are predicted incorrectly are repeatedly added to the training set, and then the classification model is continuously trained using the updated training set, and all minority class data samples are continuously predicted, and the process is repeated until all minority class data samples are correctly predicted. Therefore, unlike the random oversampling in the prior art, the oversampling in this embodiment tends to enhance the minority class data samples that are difficult to classify, and is a biased oversampling that can ensure the enhancement of the degree of difficulty in classification, so as to improve the training effect of the classification model and obtain a classification model with higher classification accuracy.
[0094] In the method of this embodiment, since under-sampled data and over-sampled data are used to train the classification model, the data used to train the classification model is better balanced, the training effect of the classification model is good, and the classification accuracy of the trained classification model is high, which overcomes the problem of poor training effect and low classification accuracy of the trained classification model caused by the imbalance of training data used in training the classification model in the prior art.
[0095] In one embodiment, a data classification method is proposed, comprising:
[0096] S00. Obtain data to be classified.
[0097] Taking the telephone customer service scenario as an example, the data to be classified may be the telephone data received by the customer service, and these telephone data need to be classified into complaint calls and consultation calls.
[0098] The steps of the data classification model training method of any of the above embodiments; and
[0099] S70: Classify the data to be classified using the data classification model that meets the second preset condition.
[0100] Taking the telephone customer service scenario as an example, the data to be classified is input into a data classification model that meets the second preset condition for processing to obtain a classification result.
[0101] refer to Figure 4 As shown, in one embodiment, a training device for a data classification model is proposed, comprising:
[0102] A partitioning module, used for partitioning a plurality of pre-acquired historical data samples into a minority class sample set and a majority class sample set;
[0103] An undersampling module, configured to obtain an undersampling set by undersampling the majority class sample set;
[0104] A first iterative training module, used to perform a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set, to obtain a classification model that meets a first preset condition;
[0105] A detection module, used to detect whether the classification model that meets the first preset condition meets the second preset condition;
[0106] An oversampling module, configured to oversample the minority class sample set based on the classification model that satisfies the first preset condition if the second preset condition is not met, and add the oversampled data samples to the training set;
[0107] The second iterative training module is used to perform second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0108] In some implementations, the plurality of pre-acquired historical data samples include two types of data samples; and the division module is further specifically configured to:
[0109] Respectively counting the number of the two types of data samples in the multiple historical data samples;
[0110] The quantities of the two data samples are compared, and the data samples with a smaller quantity are used to form the minority class sample set, and the data samples with a larger quantity are used to form the majority class sample set.
[0111] In some embodiments, the first preset condition is reaching a first preset training times threshold or reaching a first preset accuracy threshold; each iteration training in the first iteration training includes:
[0112] Training the current classification model using a training set consisting of the minority class sample set and the under-sampling set;
[0113] Determine whether the current training reaches a first preset training times threshold;
[0114] If the first preset training times threshold is not reached, the classification model after this training is used to perform classification prediction on the remaining data samples in the majority class sample set;
[0115] Determining whether the classification prediction result reaches a first preset accuracy threshold;
[0116] If the first preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the under-sampling set to obtain an updated under-sampling set; the updated under-sampling set is used for the next iterative training in the first iterative training.
[0117] In some embodiments, the method of using the trained classification model to perform classification prediction on the remaining data samples in the majority class sample set includes:
[0118] Using the classification model after the current training, predict the probability value of each remaining data sample in the majority class sample set belonging to the minority class sample set and the probability value of each remaining data sample in the majority class sample set;
[0119] The data sample with classification prediction error is a data sample whose probability value belonging to the minority class sample set is greater than the probability value belonging to the majority class sample set.
[0120] In some embodiments, the second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; each iteration training in the second iteration training includes:
[0121] Train the current classification model using the updated training set;
[0122] Determine whether the current training reaches a second preset training times threshold;
[0123] If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set;
[0124] Determining whether the classification prediction result reaches a second preset accuracy threshold;
[0125] If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
[0126] In some embodiments, determining whether the classification prediction result reaches a second preset accuracy threshold includes:
[0127] According to the number of misclassified minority class data samples in the classification prediction result, it is determined whether the classification prediction result reaches a second preset accuracy threshold.
[0128] In some embodiments, the undersampling module is specifically configured to:
[0129] A first number of majority class data samples are randomly undersampled from the majority class sample set to form an undersampled set; wherein the absolute value of the difference between the first number and the number of data samples in the minority class sample set is less than a preset threshold.
[0130] In one embodiment, a data classification device is provided, comprising:
[0131] A module for acquiring data to be classified, used for acquiring data to be classified;
[0132] A training device for a data classification model as described in any of the above embodiments; and
[0133] The classification module is used to classify the data to be classified using the classification model that reaches the preset training stop condition.
[0134] like Figure 5 As shown, in one embodiment, a computer device is provided, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program:
[0135] Dividing the multiple historical data samples acquired in advance into a minority class sample set and a majority class sample set;
[0136] Undersampling the majority class sample set to obtain an undersampled set;
[0137] Performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition;
[0138] Detecting whether the classification model that satisfies the first preset condition satisfies the second preset condition;
[0139] If the second preset condition is not met, oversampling the minority class sample set based on the classification model that meets the first preset condition, and adding the oversampled data samples to the training set;
[0140] A second iterative training is performed on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0141] In some embodiments, the first preset condition is reaching a first preset training times threshold or reaching a first preset accuracy threshold; each iteration of the first iteration training performed by the processor includes:
[0142] Training the current classification model using a training set consisting of the minority class sample set and the under-sampling set;
[0143] Determine whether the current training reaches a first preset training times threshold;
[0144] If the first preset training times threshold is not reached, the classification model after this training is used to perform classification prediction on the remaining data samples in the majority class sample set;
[0145] Determining whether the classification prediction result reaches a first preset accuracy threshold;
[0146] If the first preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the under-sampling set to obtain an updated under-sampling set; the updated under-sampling set is used for the next iterative training in the first iterative training.
[0147] In one embodiment, the processor uses the trained classification model to perform classification prediction on the remaining data samples in the majority class sample set, including:
[0148] Using the classification model after the current training, predict the probability value of each remaining data sample in the majority class sample set belonging to the minority class sample set and the probability value of each remaining data sample in the majority class sample set;
[0149] The data sample with classification prediction error is a data sample whose probability value belonging to the minority class sample set is greater than the probability value belonging to the majority class sample set.
[0150] In some embodiments, the second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; each iteration of the second iteration training performed by the processor includes:
[0151] Train the current classification model using the updated training set;
[0152] Determine whether the current training reaches a second preset training times threshold;
[0153] If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set;
[0154] Determining whether the classification prediction result reaches a second preset accuracy threshold;
[0155] If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
[0156] In one embodiment, the determining whether the classification prediction result reaches a second preset accuracy threshold performed by the processor includes:
[0157] According to the number of misclassified minority class data samples in the classification prediction result, it is determined whether the classification prediction result reaches a second preset accuracy threshold.
[0158] In one embodiment, a computer device is provided, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0159] Obtain data to be classified;
[0160] The steps of the data classification model training method described in any of the above embodiments; and
[0161] The data to be classified is classified using the data classification model that meets the second preset condition.
[0162] In one embodiment, a storage medium storing computer-readable instructions is provided. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the following steps:
[0163] Dividing the multiple historical data samples acquired in advance into a minority class sample set and a majority class sample set;
[0164] Undersampling the majority class sample set to obtain an undersampled set;
[0165] Performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition;
[0166] Detecting whether the classification model that satisfies the first preset condition satisfies the second preset condition;
[0167] If the second preset condition is not met, oversampling the minority class sample set based on the classification model that meets the first preset condition, and adding the oversampled data samples to the training set;
[0168] A second iterative training is performed on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition.
[0169] In some embodiments, the first preset condition is reaching a first preset training times threshold or reaching a first preset accuracy threshold; each iteration of the first iteration training performed by the processor includes:
[0170] Training the current classification model using a training set consisting of the minority class sample set and the under-sampling set;
[0171] Determine whether the current training reaches a first preset training times threshold;
[0172] If the first preset training times threshold is not reached, the classification model after this training is used to perform classification prediction on the remaining data samples in the majority class sample set;
[0173] Determining whether the classification prediction result reaches a first preset accuracy threshold;
[0174] If the first preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the under-sampling set to obtain an updated under-sampling set; the updated under-sampling set is used for the next iterative training in the first iterative training.
[0175] In one embodiment, the processor uses the trained classification model to perform classification prediction on the remaining data samples in the majority class sample set, including:
[0176] Using the classification model after the current training, predict the probability value of each remaining data sample in the majority class sample set belonging to the minority class sample set and the probability value of each remaining data sample in the majority class sample set;
[0177] The data sample with classification prediction error is a data sample whose probability value belonging to the minority class sample set is greater than the probability value belonging to the majority class sample set.
[0178] In some embodiments, the second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; each iteration of the second iteration training performed by the processor includes:
[0179] Train the current classification model using the updated training set;
[0180] Determine whether the current training reaches a second preset training times threshold;
[0181] If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set;
[0182] Determining whether the classification prediction result reaches a second preset accuracy threshold;
[0183] If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
[0184] In one embodiment, the determining whether the classification prediction result reaches a second preset accuracy threshold performed by the processor includes:
[0185] According to the number of misclassified minority class data samples in the classification prediction result, it is determined whether the classification prediction result reaches a second preset accuracy threshold.
[0186] In one embodiment, a storage medium storing computer-readable instructions is provided. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the following steps:
[0187] Obtain data to be classified;
[0188] The steps of the data classification model training method described in any of the above embodiments; and
[0189] The data to be classified is classified using the data classification model that meets the second preset condition.
[0190] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0191] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0192] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for training a data classification model, characterized in that: The data classification model is applied to telephone classification in a telephone customer service scenario, and the training method includes: Divide the plurality of historical data samples acquired in advance into a minority sample set and a majority sample set; the minority sample set is a set of complaint telephone data samples, and the majority sample set is a set of consultation telephone data samples; Undersampling the majority class sample set to obtain an undersampled set; Performing a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set to obtain a classification model that meets a first preset condition; Detecting whether the classification model that satisfies the first preset condition satisfies the second preset condition; If the second preset condition is not met, oversampling the minority class sample set based on the classification model that meets the first preset condition, and adding the oversampled data samples to the training set; Performing a second iterative training on the classification model that meets the first preset condition based on the updated training set to obtain a data classification model that meets the second preset condition; The plurality of pre-acquired historical data samples include two types of data samples; and the plurality of pre-acquired historical data samples are divided into a minority class sample set and a majority class sample set, including: Respectively counting the number of the two types of data samples in the multiple historical data samples; Comparing the numbers of the two data samples, using the data sample with a smaller number to form the minority class sample set, and using the data sample with a larger number to form the majority class sample set; The second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; each iteration training in the second iteration training includes: Train the current classification model using the updated training set; Determine whether the current training reaches a second preset training times threshold; If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set; Determining whether the classification prediction result reaches a second preset accuracy threshold; If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
2. The method according to claim 1, characterized in that The first preset condition is reaching a first preset training times threshold or reaching a first preset accuracy threshold; Each iteration training in the first iteration training includes: Training the current classification model using a training set consisting of the minority class sample set and the under-sampling set; Determine whether the current training reaches a first preset training times threshold; If the first preset training times threshold is not reached, using the classification model after this training to perform classification prediction on the remaining data samples in the majority class sample set; Determining whether the classification prediction result reaches a first preset accuracy threshold; If the first preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the under-sampling set to obtain an updated under-sampling set; the updated under-sampling set is used for the next iterative training in the first iterative training.
3. The method according to claim 2, characterized in that The method of using the trained classification model to perform classification prediction on the remaining data samples in the majority class sample set includes: Using the classification model after the current training, predict the probability value of each remaining data sample in the majority class sample set belonging to the minority class sample set and the probability value of each remaining data sample in the majority class sample set; The data sample with classification prediction error is a data sample whose probability value belonging to the minority class sample set is greater than the probability value belonging to the majority class sample set.
4. The method according to claim 1, characterized in that The determining whether the classification prediction result reaches a second preset accuracy threshold comprises: According to the number of misclassified minority class data samples in the classification prediction result, it is determined whether the classification prediction result reaches a second preset accuracy threshold.
5. A data classification method, characterized in that: include: Obtain data to be classified; The steps of the method according to any one of claims 1 to 4; as well as, The data to be classified is classified using the data classification model that meets the second preset condition.
6. A training device for a data classification model, characterized in that: The data classification model is applied to telephone classification in a telephone customer service scenario, and the training device includes: A division module, used to divide the pre-acquired multiple historical data samples into a minority class sample set and a majority class sample set; the minority class sample set is a set of complaint telephone data samples, and the majority class sample set is a set of consultation telephone data samples; An undersampling module, configured to obtain an undersampling set by undersampling the majority class sample set; A first iterative training module, used to perform a first iterative training on a preset classification model based on a training set consisting of the minority class sample set and the under-sampling set, to obtain a classification model that meets a first preset condition; A detection module, used to detect whether the classification model that meets the first preset condition meets the second preset condition; An oversampling module, configured to oversample the minority class sample set based on the classification model that satisfies the first preset condition if the second preset condition is not met, and add the oversampled data samples to the training set; A second iterative training module, used to perform a second iterative training on the classification model that meets the first preset condition based on the updated training set, to obtain a data classification model that meets the second preset condition; The plurality of pre-acquired historical data samples include two types of data samples; the division module is further used for: Respectively counting the number of the two types of data samples in the multiple historical data samples; Comparing the numbers of the two data samples, using the data sample with a smaller number to form the minority class sample set, and using the data sample with a larger number to form the majority class sample set; The second preset condition is reaching a second preset training times threshold or reaching a second preset accuracy threshold; each iteration training in the second iteration training includes: Train the current classification model using the updated training set; Determine whether the current training reaches a second preset training times threshold; If the second preset training times threshold is not reached, the classification model after the current training is used to perform classification prediction on the minority class sample set; Determining whether the classification prediction result reaches a second preset accuracy threshold; If the second preset accuracy threshold is not reached, the data samples with classification prediction errors are added to the minority class sample set to obtain an updated minority class sample set; the updated minority class sample set is used as the updated training set for the next iterative training in the second iterative training.
7. A data classification device, characterized in that: include: A module for acquiring data to be classified, used for acquiring data to be classified; The training device according to claim 6; as well as, The classification module is used to classify the data to be classified using the classification model that reaches the preset training stop condition.
8. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 4 and / or the steps of the method according to claim 5.
9. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the method according to any one of claims 1 to 4 and / or the steps of the method according to claim 5.
Citation Information
Patent Citations
Neural network model training method and device and data classification method and device
CN111881948A