Small Sample Data Augmentation Method and System for Bank Customer Complaint Label Classification
Data enhancement of bank customer complaint labels through deep neural network model and reverse translation technology is solved, and the classification accuracy of small sample data is achieved, higher classification accuracy and generalization capabilities are achieved, and the efficiency of bank complaint management is improved.
Patent Information
- Application Number
- CN202211455144.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-21
AI Technical Summary
In bank complaint management, small sample data leads to limited label classification accuracy, the sample word vectors generated by existing data enhancement methods are highly similar, failing to effectively capture complex data distribution, and insufficient generalization ability of small sample classification.
Tag classification and text sample grouping are performed through deep neural network models, adding noise to complaint text features with high probability of misclassification, data augmentation is used using reverse translation technology, and combining confidence learning and manual verification to generate new understandable samples.
It improves the classification accuracy and generalization ability of small sample data, enhances the ability to identify various types of label samples, and improves the efficiency and accuracy of bank customer complaint handling.
Smart Images

Figure CN116049730B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a small sample data augmentation method and system for bank customer complaint label classification. Background Art
[0002] In bank consumer protection work and product service management, customer complaint management occupies an important position. In recent years, with the increasingly strict regulatory requirements put forward by the People's Bank of China, the China Banking and Insurance Regulatory Commission, etc., and the increasing diversification of banking services and products, the complexity of customer complaint problems has increased day by day. Efficient handling of customer complaints can effectively improve the bank's service level, enhance customer relationships and stimulate product innovation capabilities. On the contrary, it is easy to trigger the bank's public opinion risk and lead to customer loss.
[0003] In the bank complaint management system, usually, a multi-dimensional standard complaint text label system is constructed, and with the help of artificial intelligence automatic classification technology, the workload of customer service personnel for classification and disposal diversion is reduced, and the work efficiency and quality are improved. Based on a deep neural network classifier, a high accuracy can be achieved on the premise that the training samples of complaint label text data are sufficient. However, in the actual business scenario, there are many banking business channels and product categories, and the amount of text data of certain specific complaint labels obtained through real complaint channels is very limited, resulting in insufficient sample sizes for a considerable proportion of label categories in the label system, bringing the problem of sample imbalance, with limited classification accuracy for small samples, so that the corresponding complaints cannot be correctly diverted and disposed of or require a lot of manual intervention.
[0004] Patent document CN114118273A (application number: CN202111425938.4) discloses a method for augmenting data for extreme multi-label classification based on label and text block attention mechanisms, including: selecting an original data set; learning the high-level semantic representation of each word in the text through BERT; splitting the text into several text blocks of equal length, and obtaining the representation of the entire text block by averaging the high-level semantic representations of each word in the text block; calculating the correlation between the representation of each text block and the vector representation of the label through an attention mechanism, fusing the representations of all text blocks, obtaining a complete label-text block relationship model after training, and then performing data augmentation according to the correlation, and finally outputting an augmented new data set.
[0005] Data augmentation is an important way to solve the problem of training with small sample data. Existing text data augmentation techniques usually use methods such as synonym replacement, random insertion, random swapping, and random deletion. However, the text augmented by these methods has very similar word vectors, and the improvement effect on the classification results is limited. The reason is that small sample augmentation not only needs to solve the problem of unbalanced classification sample quantities, but more importantly, it needs to solve the imbalance of samples with different classification difficulties. The deficiencies of existing data augmentation methods are as follows: 1. The generated sample word vectors are too similar; 2. They do not capture complex data distributions; 3. The classification generalization ability of small-class samples is not high. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a small sample data augmentation method and system for bank customer complaint label classification.
[0007] The small sample data augmentation method for bank customer complaint label classification provided by the present invention includes:
[0008] Step 1: Establish a complaint text label system and augment complaint samples;
[0009] Step 2: Through a deep neural network model, perform label classification and group text sample data;
[0010] Step 3: Add noise to the complaint text features with a misclassification probability higher than a preset threshold;
[0011] Step 4: Use back translation technology to augment the misclassified complaint samples;
[0012] Step 5: Perform automatic verification and auxiliary verification after sample augmentation.
[0013] Preferably, according to the complaint text label system, count the sample quantity of each complaint category in the training data, classify the categories with a sample quantity less than the threshold as small categories, classify the 5 categories with the largest sample quantity as large categories, and do not process the remaining samples;
[0014] Segment the small-class complaints, select the most important N words in each complaint content based on the TF-IDF technology, sort them based on the Word2vec distance from the class label and then select the most important N words, fix these 2N words as the keywords unchanged, and randomly select other words to be replaced with the content from large-class complaint samples;
[0015] The replacement rule is as follows: the number of words after word segmentation of the minor-category complaint is M, and a×M non-keyword words are randomly selected as the words to be replaced; for the major-category complaint, word segmentation is performed, and the Word2vec distance between the major-category word segmentation and the minor-category words to be replaced is calculated. The major-category word segmentation with the closest distance is selected to replace the minor-category words to be replaced, generating a new complaint marked as the minor category. Among them, the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect, and a < 0.5.
[0016] Preferably, a deep neural network model and a dataset are constructed, a complaint text dataset is obtained, and it is divided into a training set, a validation set, and a test set;
[0017] The deep neural network model is trained with the training set, the deep neural network model is evaluated with the validation set to find the best parameters, and the trained deep neural network model is used to perform label classification on the complaint text test set. The classification results are compared with the true values to obtain a confusion matrix. Among them, each column of the confusion matrix represents the predicted category, and each row of the confusion matrix represents the true belonging category of the data. The number of observed values misclassified and correctly classified by the classification model is statistically counted through the confusion matrix. The complaints with classification errors are grouped and extracted according to the true value and the predicted value, and are organized into a text file in the form of true value - predicted value - complaint content.
[0018] Preferably, according to the analysis of the classification results on the test set, if the number of samples with the true value of the complaint label being category I misclassified as category II is higher than the preset threshold, then noise from the samples with the true label of category II is added to the complaint samples with the label of category I and correctly classified in the training set. The deep neural network is used as the text classification model to train on the training set after adding noise to enhance the ability of the deep neural network text classification model to handle category II noise; for the complaints with the category I label correctly classified, word segmentation is performed, and based on the TF-IDF technology, the most important N words in each complaint content are selected. Based on the Word2vec distance from the class label, the most important N words are selected again. These 2N words are fixed as keywords and remain unchanged, and other words are randomly selected and replaced with the content from the complaint samples with the true label of category II; the replacement rule is as follows: the number of words after word segmentation of the category I complaint is M, and a×M non-keyword words are randomly selected as the words to be replaced; for the complaint text of the true category II label, word segmentation is performed, and the Word2vec distance between the category II word segmentation and the category I words to be replaced is calculated. The category II word segmentation with the closest distance is selected to replace the category I words to be replaced, generating a new complaint sample marked as category I. Among them, the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect. To avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a is less than 0.3.
[0019] Preferably, the complaint samples with misclassified categories are translated into other languages and then reverse-translated to generate multiple enhanced complaint samples;
[0020] Obtain the confidence ranking of the enhanced samples through confidence learning, train the deep neural network model for the enhanced samples with a confidence ranking higher than the threshold, and determine whether to train the deep neural network model for the enhanced samples with a confidence ranking lower than the threshold with the assistance of customer service staff.
[0021] According to the small sample data augmentation system for bank customer complaint label classification provided by the present invention, it includes:
[0022] Module M1: Establish a complaint text label system and enhance the complaint samples;
[0023] Module M2: Through the deep neural network model, perform label classification and text sample data grouping;
[0024] Module M3: Add noise to the complaint text features with a misclassification probability higher than the preset threshold;
[0025] Module M4: Use back translation technology to perform data augmentation on the misclassified complaint samples;
[0026] Module M5: Perform automatic verification and assisted verification after sample augmentation.
[0027] Preferably, according to the complaint text label system, count the sample size of each complaint category in the training data, classify the categories with a sample quantity less than the threshold as small categories, classify the 5 categories with the largest sample quantity as large categories, and do not process the remaining samples;
[0028] Segment the small category complaints, select the most important N words of each complaint content based on the TF-IDF technology, sort them based on the Word2vec distance from the class label and then select the most important N words, fix these 2N words as the keywords unchanged, and randomly select other words to be replaced with the content from the large category complaint samples.
[0029] The replacement rule is: the number of words after segmenting the small category complaints is M, and randomly select a×M non-keyword words as the words to be replaced; segment the large category complaints, calculate the Word2vec distance between the large category segmentation and the small category words to be replaced, and select the large category segmentation with the closest distance to replace the small category words to be replaced, generating new complaints marked as small categories, where the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect, and a < 0.5.
[0030] Preferably, construct a deep neural network model and a data set, obtain the complaint text data set, and divide it into a training set, a validation set, and a test set;
[0031] Train a deep neural network model using a training set, evaluate the deep neural network model through a validation set to find the best parameters, use the trained deep neural network model to perform label classification on a complaint text test set, compare the classification results with the true values to obtain a confusion matrix. Among them, each column of the confusion matrix represents the predicted category, and each row of the confusion matrix represents the true belonging category of the data. Count the number of observed values that are misclassified and correctly classified by the classification model through the confusion matrix, group and extract the misclassified complaints according to the true value and predicted value based on the confusion matrix, and organize them into a text file in the form of true value - predicted value - complaint content.
[0032] Preferably, according to the analysis of the classification results on the test set, if the number of samples with the true value of the complaint label being category I misclassified as category II is higher than the preset threshold, then add noise from the samples with the true label of category II to the complaint samples with the label of category I and correctly classified in the training set, and use the deep neural network as a text classification model to train on the training set after adding noise to enhance the ability of the deep neural network text classification model to handle category II noise; segment the complaints with the correct classification of category I label, select the most important N words for each complaint content based on the TF-IDF technology, select the most important N words again based on the Word2vec distance from the class label, and fix these 2N words as the keywords unchanged, and randomly select other words to be replaced with the content from the complaint samples with the true label of category II; the replacement rule is: the number of words after segmenting the category I complaint is M, and randomly select a×M number of non-keywords as the words to be replaced; segment the complaint text with the true category II label, calculate the Word2vec distance between the category II segments and the replaced words of category I, and select the category II segment with the closest distance to replace the replaced words of category I to generate new complaint samples marked as category I, where the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect. To avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a is less than 0.3.
[0033] Preferably, translate the misclassified complaint samples into other languages and then perform reverse translation to generate multiple enhanced complaint samples;
[0034] Obtain the confidence ranking of the enhanced samples through confidence learning, train the deep neural network model for the enhanced samples with a confidence ranking higher than the threshold, and determine whether to train the deep neural network model for the enhanced samples with a confidence ranking lower than the threshold with the assistance of customer service staff.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] (1) The present invention uses complaint label texts containing diverse information to replace some words of the original samples to generate new understandable samples. The new small-class samples incorporate the characteristics of the large-class labels, maintain the core semantics unchanged, and improve the quality of sample generation;
[0037] (2) On the one hand, the present invention adds noise features of common misclassified categories to correctly classified samples, and on the other hand, expands misclassified samples by means of back translation. At the same time, the effectiveness of enhanced samples is determined based on confidence interval judgment and manual assisted verification, improving the generalization ability of subsequent classification algorithms for various labeled samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0039] Figure 1 is a flowchart of a specific embodiment of the present invention;
[0040] Figure 2 is a flowchart of enhancing complaint samples of minor class labels in the present invention;
[0041] Figure 3 is a flowchart of selecting keywords in the present invention;
[0042] Figure 4 is an example diagram of a classification result confusion matrix. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0044] Embodiment 1:
[0045] As Figure 1 , the present invention proposes a small sample data enhancement method for the intelligent label classification scenario of bank customer complaints. By adopting statistical methods TF-IDF and Word2Vec, and through techniques of adding noise and back translation according to the results of automatic label classification, continuous text data enhancement is carried out to achieve the balance of text samples under different complaint labels and improve the classification accuracy of various complaint labels. The specific steps are as follows:
[0046] Step 1: Enhance complaint samples of minor class labels.
[0047] As Figure 2 and Figure 3, according to the complaint text tag system, count the sample size of each complaint category in the training data. Classify the categories with sample numbers less than the threshold as small categories, and classify the 5 categories with the largest sample sizes as large categories. Do not process the remaining samples. Segment the small-category complaints, select the most important N words for each complaint content based on the TF-IDF method, then select the most important N words based on the word2vec distance from the class label, and fix these 2N words as the keywords. Randomly select other words and replace them with the content from the large-category complaint samples. The replacement rule is: the number of words after segmenting the small-category complaints is M, and randomly select a×M non-keyword words as the words to be replaced. Segment the large-category complaints, calculate the word2vec distance between the large-category segmentation and the replaced words of the small category, and select the large-category segmentation with the closest distance to replace the replaced words of the small category to generate new complaints marked as small categories. Among them, the quantities M, N, and the coefficient a can be dynamically adjusted according to the actual effect, and a < 0.5.
[0048] Step 2: Automatic label classification and text sample data grouping.
[0049] Divide the data set into a training set, a validation set, and a test set. Use the small-sample data augmentation method mentioned in Step 1 above to expand the training set samples for the training set. Select a text classification model, such as BERT and LSTM, etc. The input is the bank customer complaint text, and the output is different category labels. Train the model on the training set and evaluate the model on the validation set to find the best parameters. Use the trained deep neural network model to perform automatic label classification on the complaint text test set. Compare the classification results with the true values to obtain a confusion matrix, such as Figure 4 , the confusion matrix is a special two-dimensional (actual and predicted) table, and there is the same set of categories in both dimensions. Each column represents the predicted category, and each row represents the true belonging category of the data. The number of observed values that are misclassified and correctly classified by the classification model can be counted separately. Group and extract the complaints with classification errors according to the true values and predicted values based on the confusion matrix, and organize them into text files according to "true value - predicted value - complaint content".
[0050] The training of the deep neural network model includes: applying the method in Step 1 to enhance the complaint samples with small-category labels in the training set, selecting a deep neural network text classification model, such as BERT and LSTM, etc., and inputting the enhanced training set into the deep neural network text classification model for training. Stop training when the training error of the model no longer decreases. The input of the model is the complaint text, which is data of Chinese character string type; the output is the probability of different complaint labels.
[0051] Step 3: Enhancement of correctly classified complaint samples: Add noise using the complaint text features with a relatively high probability of misclassification.
[0052] According to the analysis of the classification results on the test set, if the true value of the complaint label is class I but is misclassified as class II with a relatively large number of samples, then add the noise from the samples with the true label of class II to the complaint samples with the label of class I that are correctly classified in the training set. Use a deep neural network as the text classification model to train on the training set after adding noise to enhance the ability of the deep neural network text classification model to handle class II noise; segment the complaints with the correct classification of class I labels, select the most important N words for each complaint content based on the TF-IDF method, and then select the most important N words based on the word2vec distance from the class label. Fix these 2N words as the keywords and randomly select other words to be replaced with the content from the complaint samples with the true label of class II; the replacement rule is: the number of words after segmenting the class I complaints is M, and randomly select a×M non-keyword words as the words to be replaced; segment the complaint text with the true class II label, calculate the word2vec distance between the class II segments and the replaced words in class I, and select the class II segment with the closest distance to replace the replaced words in class I to generate new complaint samples marked as class I, where the quantities M, N, and the coefficient a can be dynamically adjusted according to the actual effect, a < 0.5, and to avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a should be less than 0.3.
[0053] The establishment of the classification model includes: applying the method in step 1 to enhance the complaint samples with small class labels in the training set, inputting the enhanced training set into the deep neural network text classification model for training, testing the trained model on the test set, and adding noise to the training set samples according to the method in step three based on the test results, and then inputting the training set after adding noise into the deep neural network text classification model for training again. The input of the classification model is the complaint text, which is data of Chinese string type; the output is the probabilities of different complaint labels.
[0054] Step 4: Enhancement of misclassified complaint samples: Use back translation technology for data enhancement.
[0055] The samples with misclassified complaint labels in the training set contain more noise and fewer keywords. Translate the misclassified complaint samples into other languages and then perform back translation to generate multiple enhanced complaint samples.
[0056] Step 5: Automatic verification and assisted verification after small sample enhancement.
[0057] Obtain the confidence ranking of the enhanced samples through confidence learning. The enhanced samples with a confidence ranking higher than the threshold can participate in model training, and the enhanced samples with a confidence ranking lower than the threshold are manually assisted by customer service staff to determine whether to participate in the training of the deep neural network text classification model.
[0058] Example 2:
[0059] The present invention also provides a small sample data augmentation system for bank customer complaint label classification. The small sample data augmentation system for bank customer complaint label classification can be implemented by executing the process steps of the small sample data augmentation method for bank customer complaint label classification. That is, those skilled in the art can understand the small sample data augmentation method for bank customer complaint label classification as the preferred implementation manner of the small sample data augmentation system for bank customer complaint label classification.
[0060] The small sample data augmentation system for bank customer complaint label classification provided by the present invention includes: Module M1: Establish a complaint text label system to augment complaint samples; Module M2: Perform label classification and text sample data grouping through a deep neural network model; Module M3: Add noise to the complaint text features with a misclassification probability higher than a preset threshold; Module M4: Use back translation technology to augment the misclassified complaint samples; Module M5: Perform automatic verification and auxiliary verification after sample augmentation.
[0061] According to the complaint text label system, count the sample size of each complaint category in the training data. Classify the categories with a sample quantity less than the threshold as small categories, classify the 5 categories with the largest sample quantity as large categories, and do not process the remaining samples; Segment the small category complaints, select the most important N words of each complaint content based on the TF-IDF technology, sort them based on the Word2vec distance from the class label and then select the most important N words, and fix these 2N words as the keywords unchanged. Randomly select other words and replace them with the content from the large category complaint samples; The replacement rule is: The number of words after segmenting the small category complaints is M, and randomly select a×M non-keyword words as the words to be replaced; Segment the large category complaints, calculate the Word2vec distance between the large category segments and the small category words to be replaced, and select the large category segment with the closest distance to replace the small category word to be replaced to generate a new complaint marked as a small category, where the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect, and a < 0.5.
[0062] Construct a deep neural network model and a data set, obtain a complaint text data set, and divide it into a training set, a validation set, and a test set; Train the deep neural network model through the training set, evaluate the deep neural network model through the validation set to find the best parameters, use the trained deep neural network model to perform label classification on the complaint text test set, compare the classification result with the true value to obtain a confusion matrix. Among them, each column of the confusion matrix represents the predicted category, and each row of the confusion matrix represents the true attribution category of the data. Count the number of observed values of misclassified and correctly classified categories in the classification model through the confusion matrix, group and extract the misclassified complaints according to the true value and the predicted value, and organize them into a text file in the form of true value - predicted value - complaint content.
[0063] According to the analysis of the classification results on the test set, if the number of samples with the true value of the complaint label being Class I but misclassified as Class II is higher than the preset threshold, then noise from the samples with the true label of Class II is added to the complaint samples with the label of Class I and correctly classified in the training set. The deep neural network is used as the text classification model to train on the training set after adding noise, strengthening the ability of the deep neural network text classification model to handle Class II noise; the complaint texts correctly classified as Class I are segmented, and based on the TF-IDF technology, the most important N words of each complaint content are selected. Then, based on the Word2vec distance from the class label, the most important N words are selected again. These 2N words are fixed as the keywords and remain unchanged, and other words are randomly selected and replaced with the content from the complaint samples with the true label of Class II; the replacement rule is: the number of words after segmenting the Class I complaint is M, and a×M non-keyword words are randomly selected as the words to be replaced; the complaint texts with the true Class II label are segmented, the Word2vec distance between the Class II segments and the replaced words of Class I is calculated, and the Class II segments with the closest distance are selected to replace the replaced words of Class I, generating new complaint samples marked as Class I, where the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect. To avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a is less than 0.3.
[0064] The complaint samples with misclassified categories are translated into other languages and then reverse-translated to generate multiple enhanced complaint samples; through confidence learning, the confidence ranking of the enhanced samples is obtained. For the enhanced samples with a confidence ranking higher than the threshold, the deep neural network model is trained, and for the enhanced samples with a confidence ranking lower than the threshold, it is determined by customer service staff whether to train the deep neural network model manually.
[0065] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structures within the hardware component.
[0066] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A small sample data augmentation method for bank customer complaint label classification, characterized in that, Including: Step 1: Establish a complaint text label system and enhance the complaint samples; Step 2: Through a deep neural network model, perform label classification and text sample data grouping; Step 3: Add noise to the complaint text features with a misclassification probability higher than the preset threshold; Step 4: Use back-translation technology to perform data enhancement on the misclassified complaint samples; Step 5: Perform automatic verification and auxiliary verification after sample enhancement; According to the complaint text label system, count the sample size of each complaint category in the training data, classify the categories with a sample quantity less than the threshold as small categories, classify the 5 categories with the largest sample quantity as large categories, and do not process the remaining samples; Segment the small-category complaints, select the most important N words for each complaint content based on the TF-IDF technology, select the most important N words again based on the Word2vec distance from the class label, fix these 2N words as the keywords unchanged, and randomly select other words to replace them with the content from the large-category complaint samples; The replacement rule is: the number of words after segmenting the small-category complaints is M, and randomly select a×M non-keyword words as the words to be replaced; segment the large-category complaints, calculate the Word2vec distance between the large-category segments and the small-category words to be replaced, and select the large-category segment with the closest distance to replace the small-category words to be replaced, generating new complaints marked as small categories, where the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect, and a < 0.
5.
2. The small sample data enhancement method for bank customer complaint label classification according to claim 1, wherein Build a deep neural network model and a dataset, obtain the complaint text dataset, and divide it into a training set, a validation set, and a test set; Train the deep neural network model through the training set, evaluate the deep neural network model through the validation set to find the best parameters, use the trained deep neural network model to perform label classification on the complaint text test set, compare the classification results with the true values to obtain a confusion matrix. Among them, each column of the confusion matrix represents the predicted category, and each row of the confusion matrix represents the true belonging category of the data. Count the number of observed values misclassified and correctly classified by the classification model through the confusion matrix, group and extract the misclassified complaints according to the true value and the predicted value, and organize them into a text file in the form of true value - predicted value - complaint content.
3. The small sample data enhancement method for bank customer complaint label classification according to claim 2, characterized in that, According to the analysis of the classification results on the test set, if the number of samples with the true value of the complaint label being category I and being misclassified as category II is higher than the preset threshold, then add the noise from the samples with the true label of category II to the complaint samples with the label of category I and being correctly classified in the training set, and use the deep neural network as the text classification model to train on the training set after adding noise to strengthen the ability of the deep neural network text classification model to handle category II noise; Complaint word segmentation for correct classification of type I labels. Based on the TF-IDF technique, select the most important N words from each complaint content. Then, based on the Word2vec distance from the class label, sort and select the most important N words again. Fix these 2N words as the keywords and randomly select other words to be replaced with the content from the real label type II complaint samples. The replacement rule is as follows: The number of words after word segmentation of type I complaints is M. Randomly select a×M non-keyword words as the words to be replaced. Perform word segmentation on the real type II label complaint text, calculate the Word2vec distance between the type II word segmentation and the type I words to be replaced, and select the type II word segmentation with the closest distance to replace the type I words to be replaced, generating a new complaint labeled as type I. Among them, the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect. To avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a is less than 0.
3.
4. The small sample data augmentation method for bank customer complaint label classification according to claim 2, characterized in that Translate the complaint samples with misclassified categories into other languages and then perform reverse translation to generate multiple enhanced complaint samples. Obtain the confidence ranking of the enhanced samples through confidence learning. For the enhanced samples with a confidence ranking higher than the threshold, perform deep neural network model training. For the enhanced samples with a confidence ranking lower than the threshold, it is determined by customer service staff whether to perform deep neural network model training manually.
5. A small-sample data augmentation system for bank customer complaint label classification, characterized in that, Including: Module M1: Establish a complaint text label system and enhance the complaint samples. Module M2: Through the deep neural network model, perform label classification and text sample data grouping. Module M3: Add noise to the complaint text features with a misclassification probability higher than the preset threshold. Module M4: Use reverse translation technology to perform data enhancement on the misclassified complaint samples. Module M5: Perform automatic verification and assisted verification after sample enhancement. According to the complaint text label system, count the sample size of each complaint category in the training data. Classify the categories with a sample quantity less than the threshold as small categories, and classify the 5 categories with the largest sample quantity as large categories. The remaining samples are not processed. Perform word segmentation on the small-category complaints. Based on the TF-IDF technique, select the most important N words from each complaint content. Then, based on the Word2vec distance from the class label, sort and select the most important N words again. Fix these 2N words as the keywords and randomly select other words to be replaced with the content from the large-category complaint samples. The replacement rule is as follows: The number of words after word segmentation of small-category complaints is M. Randomly select a×M non-keyword words as the words to be replaced. Perform word segmentation on the large-category complaints, calculate the Word2vec distance between the large-category word segmentation and the small-category words to be replaced, and select the large-category word segmentation with the closest distance to replace the small-category words to be replaced, generating a new complaint labeled as small category. Among them, the quantities M, N, and the coefficient a are dynamically adjusted according to the actual effect, and a < 0.
5.
6. The small sample data augmentation system for bank customer complaint label classification according to claim 5, characterized in that Construct a deep neural network model and a dataset, obtain the complaint text dataset, and divide it into a training set, a validation set, and a test set. Train a deep neural network model with a training set, evaluate the deep neural network model with a validation set to find the best parameters, use the trained deep neural network model to perform label classification on the complaint text test set, compare the classification results with the true values to obtain a confusion matrix. Each column of the confusion matrix represents the predicted category, and each row represents the true belonging category of the data. Count the number of observed values misclassified and correctly classified by the classification model through the confusion matrix, group and extract the misclassified complaints according to the true values and predicted values, and organize them into a text file in the form of true value - predicted value - complaint content.
7. The small sample data augmentation system for bank customer complaint label classification according to claim 6, characterized in that, According to the analysis of the classification results on the test set, if the number of samples with the true value of the complaint label being class I misclassified as class II is higher than the preset threshold, add noise from the samples with the true label of class II to the complaint samples with the label of class I and correctly classified in the training set, and use the deep neural network as the text classification model to train on the training set after adding noise to enhance the ability of the deep neural network text classification model to handle class II noise; Segment the complaints with the correct classification of class I label. Based on the TF-IDF technology, select the most important N words for each complaint content. Then, select the most important N words again based on the Word2vec distance from the class label. Fix these 2N words as the keywords and randomly select other words to be replaced with the content from the complaint samples with the true label of class II. The replacement rule is: the number of words after segmenting the class I complaints is M, and randomly select a×M non-keyword words as the words to be replaced. Segment the complaint text with the true class II label, calculate the Word2vec distance between the class II segments and the replaced words in class I, and select the class II segments with the closest distance to replace the replaced words in class I to generate new complaint samples marked as class I. The values of M, N, and the coefficient a are dynamically adjusted according to the actual effect. To avoid adding too much noise to the correctly classified samples and affecting the subsequent classification effect, the value of a is less than 0.
3.
8. The small sample data enhancement system for bank customer complaint label classification according to claim 6, characterized in that Translate the misclassified complaint samples into other languages and then perform reverse translation to generate multiple enhanced complaint samples; Obtain the confidence ranking of the enhanced samples through confidence learning, train the deep neural network model for the enhanced samples with a confidence ranking higher than the threshold, and for the enhanced samples with a confidence ranking lower than the threshold, it is determined by customer service staff whether to train the deep neural network model manually.
Citation Information
Patent Citations
Extreme multi-label classification data enhancement method based on label and text block attention mechanism
CN114118273A
An extreme multi-label classification data enhancement method based on label and text block attention mechanism
CN114118273B
Method for automatically classifying large-scale customer complaint data
CN108710651A
Intelligent work order distribution method and system
CN112528031A