Complaint work order classification method and device based on ensemble learning
By adopting an integrated learning method in the processing of complaint work ticket data, combining the gradient enhancement method and business expert model, the problems of low efficiency and insufficient accuracy of complaint work ticket classification in the existing technology are solved, and rapid and accurate internal control compliance problems are achieved, and the quality and efficiency of internal control management is improved.
Patent Information
- Application Number
- CN202411968179.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is inefficient when handling massive complaint work order data, with limited calculation accuracy and failure to effectively combine the experience of business experts, resulting in insufficient classification accuracy.
Using an integrated learning-based method, the acquired complaint work ticket data is preprocessed, key fields are extracted and text data is normalized; the preprocessed data is input to the pretrained integrated model, which is optimized and adjusted by combining the gradient enhancement method and the business expert model to generate the final classification result.
It has achieved rapid and accurate discovery of work orders involving internal control and compliance issues from massive complaint work order data, improving the effective resolution of complaints and the quality and efficiency of internal control management.
Smart Images

Figure CN120067316A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing and can also be used in the financial field. Specifically, it relates to a method and device for classifying complaint work orders based on ensemble learning. Background Art
[0002] Currently, the methods for processing complaint work orders are mainly divided into two types: manual processing and algorithmic automated processing. Manual processing relies on the professional knowledge and experience of business personnel to analyze and label the content of complaint work orders. For example, business personnel need to classify and mark the work orders according to the complaint details and handling opinions of the work orders (such as whether it involves internal control compliance issues or the specific field of the complaint). Although it has flexibility and accuracy, when faced with a large amount of work order data, manual processing is inefficient, time-consuming, and laborious, and is easily affected by subjective human factors, making it difficult to meet the high-efficiency requirements of modern data processing.
[0003] In contrast, algorithmic automated processing aims to improve efficiency and consistency, and classifies work order data through a machine learning model. Its main processes include data preprocessing and model classification. Data preprocessing is used to clean the work order data and structure it for use by the machine learning model; model classification classifies the preprocessed data through traditional algorithms such as Naive Bayes and decision trees.
[0004] However, the existing automated processing methods still face many problems in practical applications. On the one hand, models based on traditional algorithms such as Naive Bayes and decision trees have limited computational accuracy when processing complex text data, especially when faced with work orders with long texts or complex semantics, the model performance is insufficient. On the other hand, the existing models only rely on the results of machine learning training and do not incorporate the experience and judgment of business experts. The in-depth understanding of specific domain problems by business experts is the key to improving classification accuracy, but the existing methods have not achieved an organic combination of expert experience and model capabilities. In addition, the low efficiency of work order data preprocessing is also a shortcoming of the existing technology. For example, when processing work order data containing a large amount of agent conversations or repetitive clichés, common preprocessing methods are difficult to effectively clean up worthless content, resulting in serious data redundancy problems, which further affect the model performance.
[0005] This section aims to provide background or context for the embodiments of the present application stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section. Summary of the Invention
[0006] In view of the problems in the prior art, this application provides a method and device for classifying complaint work orders based on ensemble learning, which can quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, help effectively solve customer complaint problems, and improve the quality and efficiency of internal control management.
[0007] To solve at least one of the above problems, the present application provides the following technical solutions:
[0008] According to the first aspect of the embodiments of the present application, the present application provides a complaint work order classification method based on ensemble learning, including:
[0009] Preprocess the obtained complaint work order data, extract key fields and normalize the text data;
[0010] Input the preprocessed complaint work order data into a pre-trained ensemble model, where the ensemble model is obtained by optimizing the parameters of a machine learning model based on the gradient boosting method and combining a business expert model to adjust and fuse the classification results;
[0011] Classify the complaint work order data according to the output result of the ensemble model to obtain the internal control defect classification result of the complaint work order data;
[0012] Review the internal control defect classification result and iteratively optimize the ensemble model by backpropagating data.
[0013] According to any implementation manner of the present application, the pre-training process of the ensemble model includes:
[0014] Obtain a complaint work order data set from a multi-source data platform, and each historical complaint work order data in the complaint work order data set includes a complaint details field, a handling opinion field, an organization number, and a region field;
[0015] Preprocess the complaint work order data set, and the preprocessing includes data cleaning and data tokenization;
[0016] Input the preprocessed complaint work order data set into a pre-obtained word vector model to obtain corresponding word vectors, and generate corresponding sentence vectors based on the word vectors according to the average value or matrix splicing strategy;
[0017] Use the sentence vectors as training samples, construct a machine learning model by the gradient boosting method, and optimize the parameters of the machine learning model through grid search, and optimize the classification results in combination with a business expert model to generate the ensemble model.
[0018] According to any implementation manner of the present application, the preprocessing of the complaint work order data set includes:
[0019] Remove worthless text in the complaint work order data set through regular expressions, and extract core information including complaint details and handling opinions;
[0020] Perform word segmentation on the complaint work order dataset using the part-of-speech tagging mode of the Jieba word segmentation tool, and retain nouns, verbs, and pronouns in the text;
[0021] Perform word segmentation on preset domain terms based on a custom industry dictionary, and filter out words without semantic value and preset information by constructing a stop word list.
[0022] According to any embodiment of the present application, inputting the preprocessed complaint work order dataset into a pre-acquired word vector model to obtain corresponding word vectors, and generating corresponding sentence vectors based on the average value or matrix concatenation strategy according to the word vectors, including:
[0023] Input the word-segmented text data into a pre-trained word vector model, and perform vectorization processing on each word in the text through the word vector model to generate word vectors containing semantic information;
[0024] Process the word vectors in an average value or matrix concatenation manner to obtain sentence vectors that meet the input conditions of the classification model.
[0025] According to any embodiment of the present application, using the sentence vectors as training samples, constructing a machine learning model through the gradient boosting method, and optimizing the parameters of the machine learning model through grid search, and tuning the classification results in combination with a business expert model to generate the integrated model, including:
[0026] Generate synthetic samples of a small number of categories for the sentence vectors using oversampling technology, and streamline the data of a large number of categories in combination with undersampling technology;
[0027] Divide the sampled data into a training set and a test set to construct a machine learning model based on the gradient boosting method, and set the learning rate, number of iterations, maximum depth, and subsample ratio through grid search to adjust the model parameters;
[0028] Input the classification results of the gradient boosting model into the business expert model, and adjust and optimize the classification results by combining industry rules and experience to generate an integrated model including the machine learning model and the business expert model.
[0029] According to any embodiment of the present application, review the classification results of the internal control defects, and perform iterative optimization on the integrated model through backpropagated data, including:
[0030] Identify errors in the classification results based on the annotations during the review process, and generate new annotation data according to the review results;
[0031] Merge the backpropagated annotation data with the existing dataset, update the training set of the gradient boosting model and retrain it, and adjust the rules and parameters of the business expert model to make the integrated model adapt to the characteristics of the new annotation data;
[0032] Classify the unlabeled data using the updated integrated model, and deliver the classification results for review again for iterative optimization.
[0033] According to any implementation manner of the present application, it further includes:
[0034] Collect the historical complaint work order data of each user respectively, extract the complaint occurrence time, frequency and complaint theme of different users, and perform standardization processing on the complaint work order data;
[0035] Use a time series analysis model to construct a complaint prediction model based on the historical complaint work order data of each user, and the complaint prediction model is used to identify the repeated complaint behavior patterns of each user within a preset time period.
[0036] According to the second aspect of the embodiments of the present application, the present application provides a complaint work order classification device based on ensemble learning, including:
[0037] A preprocessing module for preprocessing the obtained complaint work order data, extracting key fields and normalizing the text data;
[0038] A model input module for inputting the preprocessed complaint work order data into a pre-trained integrated model, wherein the integrated model is obtained by optimizing the parameters of a machine learning model based on the gradient boosting method and combining a business expert model to adjust and fuse the classification results;
[0039] A data classification module for classifying the complaint work order data according to the output result of the integrated model to obtain the internal control defect classification result of the complaint work order data;
[0040] An iterative optimization module for reviewing the internal control defect classification result and iteratively optimizing the integrated model by backpropagating data.
[0041] According to any implementation manner of the present application, the pre-training process of the integrated model includes:
[0042] A data acquisition module for acquiring a complaint work order data set from a multi-source data platform, and each historical complaint work order data in the complaint work order data set includes a complaint details field, a handling opinion field, an organization number and a region field;
[0043] A data processing module for preprocessing the complaint work order data set, and the preprocessing includes data cleaning and data tokenization;
[0044] A vector conversion module, configured to input the preprocessed complaint work order dataset into a pre-obtained word vector model to obtain corresponding word vectors, and generate corresponding sentence vectors based on the word vectors according to an average value or matrix concatenation strategy;
[0045] A model integration module, configured to use the sentence vectors as training samples, construct a machine learning model through the gradient boosting method, optimize the parameters of the machine learning model through grid search, and optimize the classification results in combination with a business expert model to generate the integrated model.
[0046] According to any implementation manner of the present application, the data processing module includes:
[0047] An information screening unit, configured to remove worthless text in the complaint work order dataset through regular expressions and extract core information including complaint details and handling opinions;
[0048] A word segmentation unit, configured to perform word segmentation processing on the complaint work order dataset by using the part-of-speech tagging mode of the Jieba word segmentation tool, and retain nouns, verbs, and pronouns in the text;
[0049] A value filtering unit, configured to perform word segmentation on preset domain terms based on a custom industry dictionary, and filter out meaningless value words and preset information by constructing a stop word list.
[0050] According to any implementation manner of the present application, the vector conversion module includes:
[0051] A word vector conversion unit, configured to input the segmented text data into a pre-trained word vector model, and perform vectorization processing on each word in the text through the word vector model to generate word vectors containing semantic information;
[0052] A sentence vector conversion unit, configured to process the word vectors by using an average value or matrix concatenation method to obtain sentence vectors that meet the input conditions of the classification model.
[0053] According to any implementation manner of the present application, the model integration module includes:
[0054] A vector sampling unit, configured to generate synthetic samples of a small number of categories for the sentence vectors by using an oversampling technique, and streamline the data of a large number of categories in combination with an undersampling technique;
[0055] A parameter optimization unit, configured to divide the sampled data into a training set and a test set to construct a machine learning model based on the gradient boosting method, and set the learning rate, number of iterations, maximum depth, and subsample ratio through grid search to adjust the model parameters;
[0056] A model integration unit, which is used to input the classification results of the gradient boosting model into the business expert model, adjust and optimize the classification results by combining industry rules and experience, and generate an integrated model that includes a machine learning model and a business expert model.
[0057] According to any implementation manner of the present application, the iterative optimization module includes:
[0058] An error review unit, which is used to identify errors in the classification results based on the annotations during the review process, and generate new annotation data according to the review results;
[0059] A data update unit, which is used to merge the back-transmitted annotation data with the existing data set, update the training set of the gradient boosting model and retrain it, and adjust the rules and parameters of the business expert model to make the integrated model adapt to the characteristics of the new annotation data;
[0060] An iterative optimization unit, which is used to classify unlabeled data by using the updated integrated model, and deliver the classification results for review again for iterative optimization.
[0061] According to any implementation manner of the present application, it further includes a complaint prediction module, including:
[0062] A data induction unit, which is used to collect the historical complaint work order data of each user respectively, extract the complaint occurrence time, frequency and complaint theme of different users, and perform standardization processing on the complaint work order data;
[0063] A prediction model construction unit, which is used to use a time series analysis model to construct a complaint prediction model based on the historical complaint work order data of each user, and the complaint prediction model is used to identify the repeated complaint behavior patterns of each user within a preset time period.
[0064] According to the third aspect of the embodiments of the present application, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of the complaint work order classification method based on integrated learning.
[0065] According to the fourth aspect of the embodiments of the present application, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the complaint work order classification method based on integrated learning.
[0066] According to the fifth aspect of the embodiments of the present application, the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps of the complaint work order classification method based on integrated learning.
[0067] As can be seen from the above technical solutions, the present application provides a method and device for classifying complaint work orders based on ensemble learning. By preprocessing the obtained complaint work order data, key fields are extracted and the text data is normalized; the preprocessed complaint work order data is input into a pre-trained ensemble model, where the ensemble model is obtained by optimizing the parameters of a machine learning model based on the gradient boosting method and combining a business expert model to adjust and fuse the classification results; the complaint work order data is classified according to the output result of the ensemble model to obtain the internal control defect classification result of the complaint work order data; the internal control defect classification result is reviewed, and the ensemble model is iteratively optimized by backpropagating data; it is possible to quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, contribute to the effective resolution of customer complaint problems, and improve the quality and efficiency of internal control management. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0069] Figure 1 FIG. 1 is one of the flow diagrams of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0070] Figure 2 FIG. 2 is another flow diagram of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0071] Figure 3 FIG. 3 is yet another flow diagram of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0072] Figure 4 FIG. 4 is still another flow diagram of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0073] Figure 5 FIG. 5 is a schematic diagram of text vectorization in the embodiments of the present application;
[0074] Figure 6 FIG. 6 is one of the flow diagrams of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0075] Figure 7 FIG. 7 is another flow diagram of the method for classifying complaint work orders based on ensemble learning in the embodiments of the present application;
[0076] Figure 8This is the seventh flowchart diagram of the complaint work order classification method based on ensemble learning in the embodiments of the present application;
[0077] Figure 9 This is one of the structural diagrams of the complaint work order classification device based on ensemble learning in the embodiments of the present application;
[0078] Figure 10 This is the structural diagram of the electronic device in the embodiments of the present application. Detailed implementation manners
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0080] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0081] In the technical solutions of the present application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0082] It should be noted that in the embodiments of the present application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0083] Considering the problems existing in the current complaint work order processing method, the present application provides a complaint work order classification method and device based on ensemble learning, which can quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, help effectively solve customer complaint problems, and improve the quality and efficiency of internal control management.
[0084] In order to be able to quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, help effectively solve customer complaint problems, and improve the quality and efficiency of internal control management, the present application provides an embodiment of a complaint work order classification method based on ensemble learning. Refer to Figure 1 , the complaint work order classification method based on ensemble learning specifically includes the following content:
[0085] Step S11: Preprocess the obtained complaint work order data, extract key fields, and normalize the text data.
[0086] The preprocessing of complaint work order data is the key starting point of the classification process. First, the original complaint work order data can be obtained from a multi-source data platform. This data may contain multiple fields such as complaint details, handling opinions, organization numbers, and regions. Some of these fields (such as complaint details and handling opinions) are the main basis for classification.
[0087] During the preprocessing, regular expressions are used to clean the data and remove meaningless text content, such as agent scripts and repetitive clichés. At the same time, key fields with classification value are extracted and different weights are assigned to reduce data redundancy. In addition, to further improve data quality, Chinese word segmentation is performed on the text content. The part-of-speech tagging mode of the Jieba word segmentation tool is adopted, and only nouns, verbs, and pronouns are retained. By constructing a custom dictionary, it is ensured that industry-specific terms (such as "composite financial platinum card") are correctly segmented and recognized. At the same time, worthless words are filtered through a stop word list, and preset information (such as names and ID numbers) is removed to ensure privacy security.
[0088] After these preprocessing steps, the work order data is transformed into structured and standardized high-quality text data, which meets the input requirements of the subsequent classification model.
[0089] Step S12: Input the preprocessed complaint work order data into a pre-trained integrated model. Among them, the integrated model is obtained by optimizing the parameters of a machine learning model based on the gradient boosting method and combining a business expert model to adjust and fuse the classification results.
[0090] Among them, the integrated model consists of an Xgboost machine learning model based on the gradient boosting method and a business expert model. The former is responsible for efficiently processing complex work order text features, and the latter adjusts the classification results according to domain rules. The machine learning model automatically extracts and classifies multi-dimensional features of the work order data and has powerful computing capabilities; while the business expert model embeds domain internal control rules, such as supplementing the judgment logic and adjusting the boundary conditions for internal control defects in different fields.
[0091] Exemplarily, during the construction of the integrated model, the Xgboost algorithm uses the SMOTE oversampling technique to balance the data distribution, and optimizes parameters such as the learning rate, number of iterations, and maximum depth through grid search to achieve the best classification effect. The construction of the business expert model is based on the summary of historical data experience and is continuously adjusted and updated in combination with the rule set. Through the combination of machine learning and the expert model, the integrated model can not only capture complex text features but also optimize the classification boundary in terms of business.
[0092] Step S13: Classify the complaint work order data according to the output result of the integrated model to obtain the internal control defect classification result of the complaint work order data.
[0093] The output result of the integrated model is used to classify the complaint work order data. The output classification result of the model includes two parts of information: one is to judge whether the work order involves internal control defects, and the other is to specifically classify it into the field or category to which the internal control defect belongs. Through the fast computing ability of the machine learning model, the model can extract patterns from a large number of text features and classify them. For example, it can identify unreasonable operations or service quality problems hidden in specific complaint details. At the same time, the business expert model optimizes the classification result at the business level. For example, by setting rules to adjust the classification result of boundary cases, it ensures that the classification logic is consistent with the actual business requirements. The finally generated classification result can accurately label the internal control problems of the work order.
[0094] Step S14: Review the classification result of the internal control defect and iteratively optimize the integrated model through the feedback data.
[0095] After the classification is completed, review the classification result of the internal control defect to further improve the reliability of the classification. The review process is manually or automatically verified by business personnel in combination with the actual scenario and experience to discover possible misjudged or missed cases in the classification model and relabel these cases.
[0096] Exemplarily, a rule engine can be defined based on domain knowledge, including keyword matching, logical constraints, and anomaly detection rules. For example, it can detect whether high-frequency words such as "delay" and "timeout" are included in the complaint details to mark "service problems". Subsequently, compare the classification result output by the integrated model with the rule engine. If the classification result does not conform to the rule (such as high-frequency unreasonable words are not marked as internal control problems), it is marked as "possibly misclassified". At the same time, use a time series analysis model to detect the trend and anomalies of the classification result. For example, a warning is triggered when the number of a certain classification surges abnormally in a short period of time. For the abnormally classified cases marked by the rule or trend, automatically re-enter them into the model and perform multiple iterative classifications by adjusting the model weights to correct the classification result. Finally, a verification report including the correction rate and anomaly cases is generated, and special work orders are pushed to business personnel for review to ensure the accuracy and consistency of the classification result.
[0097] After the review is completed, the newly added marked data will be used as feedback data and fed back to the integrated model for update. By merging these newly marked data with the original training data set, the Xgboost model is retrained, and at the same time, the rules of the business expert model are adjusted to enhance the adaptability of the model to complex scenarios. The review and feedback optimization process is an iterative cycle, continuously improving the model performance to make its classification result more accurate and stable, so as to meet the dynamic needs of the actual business.
[0098] As can be seen from the above description, the complaint work order classification method based on ensemble learning provided by the embodiments of the present application can quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, help effectively solve customer complaint problems, and improve the quality and efficiency of internal control management.
[0099] In an embodiment of the complaint work order classification method based on ensemble learning of the present application, refer to Figure 2 , the pre-training process of the ensemble model includes:
[0100] Step S01: Obtain a complaint work order data set from a multi-source data platform. Each historical complaint work order data in the complaint work order data set includes a complaint details field, a handling opinion field, an organization number, and a region field.
[0101] In the pre-training process of the ensemble model, first obtain a complaint work order data set from a multi-source data platform to ensure the diversity and representativeness of the data set. Each historical complaint work order data includes multiple key fields, such as a complaint details field, a handling opinion field, an organization number, and a region field. Among them, the complaint details field and the handling opinion field contain the core information of the work order and are the main basis for classification; the organization number and the region field are used as auxiliary fields to provide information on the business source and geographical distribution of the work order, enhancing the context background of classification. During the collection process, work order data within the province can be extracted from the zipper table of the unified customer service physical subsystem, and at the same time, combined with the labeled data of the out-of-province branches to expand the extensiveness of the data, laying a foundation for subsequent model training.
[0102] Exemplarily, the customer complaint work order data of the present application comes from the WFS_REL_TMPL_INFO table in the unified customer service physical subsystem, channel service - employee channel logic subsystem. It contains 27 fields, and a zipper table process is performed on it according to the work order number to obtain the complaint details and handling opinions of the work order as the main basis for classification.
[0103] The customer complaint data comes from an excel file. After analyzing the table header, the complaint details and handling opinion fields in it are intercepted as the main basis for classification.
[0104] Step S02: Preprocess the complaint work order data set. The preprocessing includes data cleaning and data tokenization.
[0105] The obtained dataset usually contains a large amount of redundant information and formatting problems, so it needs to be cleaned and normalized through preprocessing. Data cleaning first uses regular expressions to remove meaningless text content in work orders, such as agent scripts and repetitive clichés, and at the same time eliminates information related to preset information (such as names and ID numbers). Subsequently, the complaint details and handling opinions fields are extracted as the main analysis objects, and weights are assigned to different fields to highlight key information. The cleaned data is tokenized through a Chinese word segmentation tool, using the part-of-speech tagging mode of Jieba Segmentation, only retaining nouns, verbs, and pronouns, and combining a custom dictionary to handle industry-specific terms (such as "UnionPay Composite Wealth Platinum Card") to ensure the accuracy and business relevance of the tokenization results. Finally, words with no semantic value are filtered through a stop word list to further optimize the text quality.
[0106] Step S03: Input the preprocessed complaint work order dataset into a pre-obtained word vector model to obtain corresponding word vectors, and generate corresponding sentence vectors based on the average value or matrix concatenation strategy according to the word vectors.
[0107] The preprocessed text data is input into a pre-trained Word2vec word vector model for vectorization processing. The Word2vec model learns the context semantic relationships of a large-scale corpus and converts each word in the text into a low-dimensional numerical vector, thereby retaining the semantic information of the text. Subsequently, based on the generated word vectors, an average value or matrix concatenation strategy is used to aggregate multiple word vectors to generate a sentence vector containing overall semantic information. The average value strategy obtains the sentence vector by calculating the mean of all word vectors in the text, while the matrix concatenation strategy expresses the sentence semantics more comprehensively by maintaining the order relationship of the word vectors. The finally generated sentence vector can adapt to the input requirements of machine learning models and provide a structured data representation for the classification process.
[0108] Step S04: Use the sentence vector as a training sample, construct a machine learning model through the gradient boosting method, optimize the parameters of the machine learning model through grid search, and tune the classification results in combination with the business expert model to generate the integrated model.
[0109] During the pre-training process, a machine learning model is constructed based on the gradient boosting method (Xgboost algorithm) to efficiently process the preprocessed sentence vector data. The Xgboost model effectively captures complex patterns in the complaint work orders by iteratively generating a series of weak classifiers (decision trees) and combining them into a strong classifier. Subsequently, a grid search method is used to optimize the model parameters, including adjusting the learning rate, tree depth, number of iterations, etc., to improve the model performance. The optimized machine learning model is combined with the business expert model to construct the final integrated model. The business expert model adjusts the classification results of the machine learning model by incorporating domain knowledge, such as setting specific classification rules or optimizing the classification results of boundary cases. The finally generated integrated model not only has the powerful computing ability of machine learning but also combines the profound experience and judgment of business experts to support the classification task.
[0110] In an embodiment of the complaint work order classification method based on ensemble learning of the present application, refer to Figure 3 , the preprocessing of the complaint work order data set includes:
[0111] Step S02A: Remove the worthless text in the complaint work order data set through regular expressions, and extract the core information including the complaint details and handling opinions.
[0112] The complaint data comes from the front line of the business, with characteristics such as non-standard formats, containing worthless content such as agent jargon, and containing some preset information. It is necessary to standardize the text data through data cleaning. The present application first analyzes the text data, uses regular expressions to remove worthless text such as agent jargon, and at the same time intercepts five parts of it as keyword fields to reduce the text length and improve the text accuracy.
[0113] Step S02B: Perform word segmentation on the complaint work order data set using the part-of-speech tagging mode of the Jieba word segmentation tool, and retain the nouns, verbs, and pronouns in the text.
[0114] Word segmentation is one of the basic steps in natural language processing. The processing methods for word segmentation of Chinese and English texts are quite different. For English texts, spaces can be used as the basis for word segmentation; for Chinese texts, since there are no obvious word boundaries in sentences, the fuzzy and diverse word segmentation standards have always been a major difficulty in word segmentation, and different professional fields have different dictionary structures, and different word segmentation results will be obtained according to different word segmentation rules. Therefore, when processing Chinese texts, the word segmentation method will directly affect the effect of subsequent calculations. Chinese word segmentation algorithms mainly include string matching-based, statistics-based, and machine learning-based word segmentation methods and simulating human understanding of sentences. However, due to the complex Chinese semantics, this kind of word segmentation system needs to be developed. From the perspective of word segmentation granularity, the currently popular Chinese word segmentation methods are character-level word segmentation and Jieba word segmentation. The present application uses the Jieba word segmentation method.
[0115] Step S02C: Perform word segmentation on terms in a preset field based on a custom industry dictionary, and filter out words with no semantic value and preset information by constructing a stop word list.
[0116] To improve the accuracy of the word segmentation results, this application adds a custom dictionary on the basis of Jieba word segmentation to ensure that some words with industry characteristics, such as "UnionPay Composite Financial Platinum Card", "LongPay", "credit limit", etc. are correctly segmented. At the same time, a stop word list is added to filter out some words with low semantic value.
[0117] To further improve the accuracy of the text after word segmentation, this application uses the part-of-speech tagging mode of Jieba word segmentation, only retaining nouns, verbs, and pronouns, while further reducing the sample length and filtering out preset information such as names, ID numbers, and accounts.
[0118] In an embodiment of the complaint work order classification method based on ensemble learning in this application, refer to Figure 4 wherein, inputting the preprocessed complaint work order data set into a pre-obtained word vector model to obtain corresponding word vectors, and generating corresponding sentence vectors based on the average value or matrix splicing strategy according to the word vectors includes:
[0119] Step S03A: Input the word-segmented text data into a pre-trained word vector model, and perform vectorization processing on each word in the text through the word vector model to generate word vectors containing semantic information;
[0120] Step S03B: Process the word vectors in an average value or matrix splicing manner to obtain sentence vectors that meet the input conditions of the classification model.
[0121] Human language is highly ambiguous, and a sentence may have multiple meanings or metaphors, while computers currently cannot truly understand the meaning of language or words. Therefore, represent text information as vectors that can express the semantics of the text, and then analyze the vectors or use machine learning to build models. Currently, common text vectorization methods include one-hot encoding, Bag of Words Model, TF-IDF, N-Gram, Word2vec, etc. Through experimental result comparison, this application selects the relatively better-performing Word2vec as the tool for vectorizing work order data represented by natural language.
[0122] Implementing text vectorization through Word2vec includes three steps: obtaining the Word2vec model, generating word vectors, and generating sentence vectors.
[0123] (1) Obtain the Word2vec model
[0124] There are two ways to obtain the Word2vec model: referring to a pre-trained model and training a model by oneself. The pre-trained model is trained based on a large-scale public dataset and has the characteristics of rich corpus features and comprehensive coverage of text features; the self-trained model is trained based on the customer complaint dataset of this application and has limited prediction features but is more targeted for this application. Through experimental comparison, this application selects the pre-trained Word2vec model with better performance.
[0125] In an optional embodiment, the corpus used for model training is financial news from multiple websites. The size of the corpus is 6.2GB, the size of the vocabulary is 2.785 million, the dimension of the word vector is 300, and the window size is 5.
[0126] (2) Generate word vectors
[0127] Take the segmented text data as input, and based on the obtained Word2vec model, perform the conversion from natural language to numerical vectors in units of words.
[0128] (3) Generate sentence vectors
[0129] There are various strategies for generating sentence vectors based on word vectors, such as taking the average value and matrix splicing. The dimension and length of the sentence vector need to meet the input conditions required by the subsequent classification model. The final obtained sentence vector is a numerical vector containing the semantic information of the original data.
[0130] Figure 5 Fig. shows a schematic diagram of text vectorization provided by this application, including the selection of text vectorization methods, the acquisition of word vector models, and the specific steps of generating word vectors and sentence vectors. This application selects common text vectorization methods for experimental comparison, including one-hot encoding, Bag of Words Model, TF-IDF, N-Gram, Bert, and Word2vec. Through comparison of experimental results, Word2vec with better performance is finally selected as the vectorization tool, which can effectively convert natural language text into high-quality numerical expressions.
[0131] This application captures the hidden emotion information in complaint work orders through keyword extraction, deep data cleaning, and text vectorization, combined with a sentiment analysis model of context semantics, enhances the model's ability to express long texts and complex semantics, and improves the accuracy and adaptability of the classification model.
[0132] In an embodiment of the complaint work order classification method based on ensemble learning of this application, see Figure 6, taking the sentence vector as a training sample, constructing a machine learning model by gradient boosting method, optimizing the parameters of the machine learning model through grid search, and tuning the classification result in combination with a business expert model to generate the integrated model, including:
[0133] Step S04A: Generate synthetic samples of a small number of categories for the sentence vector by oversampling technology, and streamline the data of a large number of categories in combination with undersampling technology.
[0134] Exemplarily, this application uses the holdout method to divide the dataset. Take an 8-2 ratio for splitting, directly randomly divide the data into a training set and a test set, use the training set to generate a model, and use the test set to test the accuracy and error of the model to verify the effectiveness of the model.
[0135] The data used in this application includes two major categories: valuable complaints and worthless complaints, with a distribution ratio of approximately 1-30, belonging to an imbalanced dataset. Due to the uneven distribution of features in the dataset, the sample features of the category corresponding to the small amount of data are insufficient, and the sample features of the category corresponding to the large amount of data are redundant. Using an imbalanced dataset for training a classification model is likely to cause the model to be overfitted or underfitted.
[0136] Through the resampling method, resample the divided dataset again to reduce the negative impact of data imbalance on the training of the classification model to a certain extent. The resampling method includes oversampling and undersampling. Oversampling is to expand the data of the category with less data volume, and undersampling is to reduce the data of the category with more data volume. Among them, the oversampling method includes naive random oversampling, interpolation sampling, etc., and the undersampling method includes naive random undersampling, nearest neighbor algorithm, etc.
[0137] Through result comparison, this application selects the combination of SMOTE (Synthetic Minority Over-sampling Technique) oversampling and Condensed-nn (Condensed Nearest Neighbors) undersampling for resampling.
[0138] Step S04B: Divide the sampled data into a training set and a test set to construct a machine learning model based on the gradient boosting method, and set the learning rate, number of iterations, maximum depth, and subsample ratio through grid search to adjust the model parameters.
[0139] Traditional machine learning algorithms require relatively less computing resources during training and relatively shorter training time; deep learning algorithms have stronger computing capabilities for long texts and samples with complex features, and self-training is prone to underfitting or overfitting when the sample size is insufficient. Since the labeling of the work order data used in this application requires a solid business foundation and high professionalism, the first two batches of data are all manually labeled, so the data volume is limited.
[0140] Taking into account the computing resources used during the training process and the time cost required for iterative optimization, this application initially used the first two batches of provincial datasets to conduct a comparative analysis of multiple classification algorithms and found that SVM performed well and had a lower time cost during the training process. After adding datasets later, a comparative analysis of multiple algorithms was conducted, and it was found that Xgboost performed well.
[0141] Xgboost is short for eXtreme Gradient Boosting, that is, Extreme Gradient Boosting Tree, which is a boosting tree model that integrates many tree models to form a stronger classifier. The tree model used therein is the CART regression tree model. Xgboost is one of the Boosting algorithms. The idea of the Boosting algorithm is to integrate many weak classifiers to form a strong classifier (a serialization method where there is a strong dependence between individual learners and they must be generated serially).
[0142] According to the grid search method for parameter tuning, the final Xgboost parameters are set as follows:
[0143] learning_rate = 0.1
[0144] The learning rate (also known as the step size) controls the update amplitude of the model weights in each iteration. A lower learning rate can improve the generalization ability of the model.
[0145] n_estimators = 400
[0146] The number of base learners, that is, the total number of decision trees included in the model. More trees may increase the fitting ability of the model.
[0147] subsample = 1
[0148] The subsampling ratio is used to control the proportion of samples randomly selected when building each tree. A value of 1 means using all samples.
[0149] max_depth = 20
[0150] The maximum depth of each tree controls the complexity of the tree. A larger depth allows the model to fit more complex patterns.
[0151] init = None
[0152] The initial model is usually used to set a base model to guide the initial training of Xgboost. Setting it to None means that no initial model is provided.
[0153] random_state = None
[0154] Random seed, used to control the initial state of random number generation, ensuring the repeatability of model training results. Setting it to None means the random seed is not fixed, and different results may be produced each time it runs.
[0155] max_features = None
[0156] Controls the number of features available when building each tree. None means using all features, reducing the limitation of feature selection, thus allowing the model to learn from all features.
[0157] max_leaf_nodes = None
[0158] Determines the maximum number of leaf nodes allowed in each tree, used to control the complexity of the tree. None means not restricting the number of leaf nodes and letting the model decide automatically.
[0159] seed = 0
[0160] Random seed, used to set the basis for random number generation. Similar to random_state, but used for global control of the random process. Fixed at 0 to ensure the randomness of training remains consistent across different runs.
[0161] Step S04C: Input the classification results of the gradient boosting model into the business expert model, and adjust and optimize the classification results by combining industry rules and experience to generate an integrated model that includes both the machine learning model and the business expert model.
[0162] This application forms an integrated model by combining the content matching expert model with the machine learning classification algorithm to further improve the classification efficiency and accuracy. This combination takes advantage of the strengths of both methods: the machine learning classification algorithm is responsible for efficiently processing large-scale complaint work order data to generate preliminary classification results, while the content matching expert model uses domain knowledge to optimize the classification results.
[0163] For example, the Xgboost model can quickly determine whether a certain complaint involves the issue of "service delay" by extracting and learning the text features of the work order. If the algorithm finds that keywords such as "timeout" and "not arriving on time" frequently appear in a certain work order during model training, then the work order is preliminarily classified as a "service problem". However, some work orders may contain ambiguous or abnormal expressions, such as "delayed but the customer has been notified" mentioned in the description, which may not belong to the regular service problem.
[0164] At this time, the content matching expert model comes into play. Based on predefined rules, such as "if the work order handling opinion contains words like 'has been notified' or 'the customer has agreed', it is not considered a valuable complaint", the expert model makes a secondary judgment on the preliminary classification result. Through rule optimization, this work order may ultimately be reclassified as a "worthless complaint".
[0165] This combination method is particularly important when dealing with complex scenarios. For example, a complaint mentions "the product arrived late and the customer service did not reply". The machine learning algorithm may only focus on the keyword "late" and classify it as a logistics problem, while the expert model combines the context to supplement the judgment that "the customer service did not reply" may be a customer service quality problem, and finally classifies the work order as a "comprehensive service problem". Through the collaborative work of machine learning and the expert model, the integrated model not only retains the powerful computing ability of machine learning but also integrates the accuracy of expert experience, thus more efficiently and accurately determining whether the complaint data is a valuable complaint.
[0166] In an embodiment of the complaint work order classification method based on ensemble learning of the present application, refer to Figure 7 , the review of the internal control defect classification result and the iterative optimization of the integrated model through the feedback data include:
[0167] Step S14A: Identify the errors in the classification result based on the annotations during the review process, and generate new annotation data according to the review result.
[0168] In the review stage, the internal control defect classification result output by the integrated model is reviewed, and the classification errors are identified manually or automatically. For example, a complaint work order is classified as a "service problem", but after review, it is found that it actually belongs to a "product quality problem", and this misclassification will be re-annotated. The review process can combine domain rules and actual scenarios, and ensure the accuracy of the annotation by carefully analyzing the complaint details field and the handling opinion field. After the review is completed, the newly generated annotation data is recorded and stored for subsequent model update and optimization.
[0169] Step S14B: Merge the feedback annotation data with the existing data set, update the training set of the gradient boosting model and retrain it, and adjust the rules and parameters of the business expert model to make the integrated model adapt to the characteristics of the new annotation data.
[0170] After merging the labeled data generated during the review process with the original dataset, reconstruct the training set for the gradient boosting model (such as Xgboost), and retrain the model with the newly added data as a supplement to optimize the classification ability of the model. At the same time, adjust the rules and parameters of the business expert model according to the characteristics of the newly added labeled data. For example, for some frequently misclassified cases in the new labeled data, add or modify the judgment rules of the expert model (such as introducing specific keywords or context conditions) to ensure that the integrated model better adapts to the characteristics of the new data in the next round of classification. Through this process, the model continuously learns and absorbs new information, thereby improving the overall classification performance.
[0171] Step S14C: Classify the unlabeled data using the updated integrated model, and deliver the classification results for review again for iterative optimization.
[0172] After completing the model update, use the optimized integrated model to perform new classification predictions on the unlabeled data. The new classification results will be delivered for review again. By combining with domain experts or automated verification mechanisms, further annotation and correction of possible classification errors will be carried out. During the review process, the newly discovered errors and the newly added labeled data will be passed back again to update the training set and adjust the model rules, thereby completing the next round of iterative optimization. Through continuous review, feedback, and update, the integrated model gradually enhances its classification ability, adapts to the characteristics of dynamically changing complaint data, and significantly improves the classification accuracy and efficiency.
[0173] In an embodiment of the complaint work order classification method based on ensemble learning of the present application, refer to Figure 8 and may specifically include the following content:
[0174] Step S15: Collect the historical complaint work order data of each user respectively, extract the complaint occurrence time, frequency, and complaint theme of different users, and perform standardization processing on the complaint work order data.
[0175] Among them, all historical complaint work order data related to users can be obtained from a multi-source data platform and sorted in chronological order. Each complaint work order includes key information such as complaint details, complaint themes, and complaint handling opinions, as well as a timestamp (complaint occurrence time) and frequency (the number of complaints of a user within a certain period). These information are extracted to form a dataset reflecting the user behavior pattern. Subsequently, standardization processing is performed on these data, such as unifying the time format and normalizing the complaint frequency index.
[0176] Step S16: Use a time series analysis model to construct a complaint prediction model based on the historical complaint work order data of each user, and the complaint prediction model is used to identify the repeated complaint behavior patterns of each user within a preset time period.
[0177] After data standardization, a time series analysis model is used to model the historical complaint data of users. The time series analysis model generates a complaint prediction model for users by capturing the time dependence and trend characteristics in the complaint data.
[0178] Specifically, the model analyzes the pattern of each user's complaint frequency changing over time (such as some users complaining frequently during specific time periods), and identifies the periodicity and outliers in their behavior patterns. For example, through historical data, it can be found that a certain user has complained about product quality in multiple past quarters, and it usually occurs at the end of the quarter. Thus, it can be predicted that the user may have a similar complaint behavior again at the end of the future quarter. The model can also combine the characteristics of the complaint theme (such as a certain user's multiple complaints about a specific product) to further predict possible complaint types or problem areas.
[0179] Through this process, the complaint prediction model can generate a personalized behavior pattern for each user to identify their repeated complaint behavior within a preset future time period. This model provides data support for enterprises to intervene in potential problems in advance and optimize the customer service process, and also provides warning information for internal control management.
[0180] Based on the foregoing embodiments, in view of the characteristics of large volume and complex features of customer complaint work order data, the present application proposes a method for classifying complaint work orders based on natural language processing and machine learning. The present application transforms the analysis of customer complaint work orders into a text classification task, and adopts natural language processing and machine learning algorithms, effectively improving the processing efficiency and accuracy of complaint work orders.
[0181] The overall implementation process of the present application includes three parts: data processing, text vectorization, and classification model. In the data processing link, through steps such as obtaining key fields, data cleaning, Chinese word segmentation, and removing stop words, the original text data is normalized and the feature quality is optimized. In the text vectorization link, a pre-trained Word2vec model is used to generate word vectors, and sentence vectors are generated through strategies such as taking the average value or matrix concatenation, retaining the semantic information of the original data. The classification model link includes the construction and training of a machine learning model based on Xgboost, and the construction and adjustment of rules for the business expert model. By combining the machine learning model with the business expert model, the present application forms an integrated model with high precision, further improving the classification effect of work order data. The final model can automatically determine whether a complaint work order has internal control defects and give efficient and accurate analysis results.
[0182] This application has the following advantages: First, by using a machine learning model to replace manual analysis and processing, the efficiency of complaint work order processing has been significantly improved. Second, when applied to the analysis and processing of complaint work orders, it has no emotional tendency towards data elements such as regions, complainants, and handlers. Third, the machine learning model based on Xgboost is superior to traditional decision tree and random forest models in classification accuracy, especially in the analysis scenario of bank customer complaint data. Fourth, this application focuses on the analysis of bank customer complaint data during the model construction and training process, and is outstanding in the theme of bank customer complaints, effectively improving the analysis ability for bank customer complaint problems. Fifth, this application combines the powerful computing power of the machine learning model with the expert knowledge of business experts, and effectively improves the accuracy and reliability of the analysis results through the integrated model. This method is particularly suitable for the processing of bank customer complaint work orders, helping to quickly locate problems and optimize service processes. In summary, this application provides an efficient and accurate solution for customer complaint data analysis and has broad practical application value.
[0183] In order to quickly and accurately discover work orders involving internal control compliance issues from a large amount of complaint work order data, assist in the effective resolution of customer complaint problems, and improve the quality and efficiency of internal control management, this application provides an embodiment of an integrated learning-based complaint work order classification device for implementing all or part of the above-mentioned integrated learning-based complaint work order classification method. Refer to Figure 9 , since the principle of the device for solving problems is similar to that of the integrated learning-based complaint work order classification method, the implementation of the device can refer to the implementation of the integrated learning-based complaint work order classification method, and the repeated parts will not be elaborated.
[0184] The integrated learning-based complaint work order classification device 002 provided by the embodiment of this application is shown in Figure 9 , and includes a preprocessing module 21, a model input module 22, a data classification module 23, and an iterative optimization module 24, where:
[0185] The preprocessing module is used to preprocess the obtained complaint work order data, extract key fields, and standardize the text data;
[0186] The model input module is used to input the preprocessed complaint work order data into a pre-trained integrated model, where the integrated model is obtained by optimizing the parameters of a machine learning model based on the gradient boosting method and combining the business expert model to adjust and fuse the classification results;
[0187] The data classification module is used to classify the complaint work order data according to the output result of the integrated model to obtain the internal control defect classification result of the complaint work order data;
[0188] An iterative optimization module for reviewing the classification results of the internal control defects and iteratively optimizing the integrated model through the feedback data.
[0189] According to any implementation manner of the present application, the pre-training process of the integrated model includes:
[0190] A data acquisition module for acquiring a complaint work order data set from a multi-source data platform, where each historical complaint work order data in the complaint work order data set includes a complaint details field, a handling opinion field, an organization number, and a region field;
[0191] A data processing module for preprocessing the complaint work order data set, and the preprocessing includes data cleaning and data word segmentation;
[0192] A vector conversion module for inputting the preprocessed complaint work order data set into a pre-acquired word vector model to obtain corresponding word vectors, and generating corresponding sentence vectors based on the word vectors according to an average value or a matrix splicing strategy;
[0193] A model integration module for using the sentence vectors as training samples, constructing a machine learning model through the gradient boosting method, optimizing the parameters of the machine learning model through grid search, and tuning the classification results in combination with a business expert model to generate the integrated model.
[0194] According to any implementation manner of the present application, the data processing module includes:
[0195] An information screening unit for removing worthless text in the complaint work order data set through regular expressions and extracting core information including complaint details and handling opinions;
[0196] A word segmentation unit for performing word segmentation processing on the complaint work order data set in the part-of-speech tagging mode of the Jieba word segmentation tool, and retaining nouns, verbs, and pronouns in the text;
[0197] A value filtering unit for performing word segmentation on preset domain terms based on a custom industry dictionary and filtering out meaningless value words and preset information by constructing a stop word list.
[0198] According to any implementation manner of the present application, the vector conversion module includes:
[0199] A word vector conversion unit for inputting the segmented text data into a pre-trained word vector model and performing vectorization processing on each word in the text through the word vector model to generate word vectors containing semantic information;
[0200] A sentence vector conversion unit for processing the word vectors in an average value or matrix splicing manner to obtain sentence vectors that meet the input conditions of the classification model.
[0201] According to any embodiment of the present application, the model integration module includes:
[0202] A vector sampling unit, configured to generate synthetic samples of a small number of categories for the sentence vectors by using an oversampling technique, and streamline the data of a large number of categories in combination with an undersampling technique;
[0203] A parameter optimization unit, configured to divide the sampled data into a training set and a test set to construct a machine learning model based on the gradient boosting method, and set the learning rate, the number of iterations, the maximum depth, and the subsample ratio through grid search to adjust the model parameters;
[0204] A model integration unit, configured to input the classification result of the gradient boosting model into the business expert model, and adjust and optimize the classification result by combining industry rules and experience to generate an integrated model including the machine learning model and the business expert model.
[0205] According to any embodiment of the present application, the iterative optimization module includes:
[0206] An error review unit, configured to identify errors in the classification result based on the annotations in the review process, and generate new annotation data according to the review result;
[0207] A data update unit, configured to merge the back-transmitted annotation data with the existing data set, update the training set of the gradient boosting model and retrain it, and adjust the rules and parameters of the business expert model to make the integrated model adapt to the characteristics of the new annotation data;
[0208] An iterative optimization unit, configured to classify the unannotated data by using the updated integrated model, and deliver the classification result for review again for iterative optimization.
[0209] According to any embodiment of the present application, it further includes a complaint prediction module, including:
[0210] A data induction unit, configured to collect the historical complaint work order data of each user respectively, extract the complaint occurrence time, frequency and complaint theme of different users, and perform standardization processing on the complaint work order data;
[0211] A prediction model construction unit, configured to use a time series analysis model to construct a complaint prediction model based on the historical complaint work order data of each user, and the complaint prediction model is used to identify the repeated complaint behavior patterns of each user within a preset time period.
[0212] As can be seen from the above description, the complaint work order classification device based on ensemble learning provided by the embodiments of the present application can quickly and accurately discover the work orders involving internal control compliance issues from a large amount of complaint work order data, help effectively solve customer complaint problems, and improve the quality and efficiency of internal control management.
[0213] It should be noted that the complaint work order classification method based on ensemble learning provided in the embodiments of the present application can be used in the financial field and can also be used in any technical field other than the financial field. The embodiments of the present application do not limit the application fields of the bank training system based on virtual reality and its training method.
[0214] Figure 10 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of the present application, as Figure 10 shown. The electronic device 003 includes: a processor 301, a memory 302, and a bus 303.
[0215] Among them, the processor 301 and the memory 302 communicate with each other through the bus 303.
[0216] The processor 301 is used to call the program instructions in the memory 302 to execute the methods provided in the above-mentioned method embodiments.
[0217] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the complaint work order classification method based on ensemble learning described above is implemented.
[0218] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the complaint work order classification method based on ensemble learning described above is implemented.
[0219] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0220] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1One or more processes and / or boxes Figure 1 Apparatus for the functions specified in one or more boxes
[0221] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes
[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 One or more processes and / or boxes Figure 1 Steps for the functions specified in one or more boxes
[0223] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A complaint ticket classification method based on ensemble learning, characterized in that: include: Preprocess the acquired complaint ticket data, extract key fields and normalize text data; Inputting the pre-processed complaint work order data into a pre-trained integrated model, wherein the integrated model is obtained by optimizing parameters of a machine learning model based on a gradient boosting method and adjusting and fusing the classification results in accordance with the rules of a business expert model; Classify the complaint work order data according to the output result of the integrated model to obtain the internal control defect classification result of the complaint work order data; The internal control defect classification results are reviewed, and the integrated model is iteratively optimized through feedback data.
2. The method for classifying complaint tickets based on ensemble learning according to claim 1 is characterized in that: The pre-training process of the integrated model includes: Acquire a complaint work order data set from a multi-source data platform, where each historical complaint work order data in the complaint work order data set includes a complaint details field, a handling opinion field, an institution number, and a region field; Preprocessing the complaint work order data set, wherein the preprocessing includes data cleaning and data segmentation; Input the preprocessed complaint ticket dataset into the pre-acquired word vector model to obtain the corresponding word vector, and generate the corresponding sentence vector based on the word vector based on the average value or matrix concatenation strategy; The sentence vector is used as a training sample, a machine learning model is constructed by a gradient boosting method, and the parameters of the machine learning model are optimized by grid search. The classification results are tuned in combination with a business expert model to generate the integrated model.
3. The method for classifying complaint work orders based on ensemble learning according to claim 2 is characterized in that: The preprocessing of the complaint work order data set includes: Use regular expressions to remove worthless text from the complaint ticket dataset and extract core information including complaint details and handling opinions; The part-of-speech tagging mode of the Jieba word segmentation tool is used to perform word segmentation on the complaint ticket dataset, retaining nouns, verbs, and pronouns in the text; Segment preset domain terms based on custom industry dictionaries, and filter out words without semantic value and preset information by building a stop word list.
4. The method for classifying complaint work orders based on ensemble learning according to claim 2 is characterized in that: The preprocessed complaint ticket dataset is input into the pre-acquired word vector model to obtain the corresponding word vector, and based on the average value or matrix splicing strategy, the corresponding sentence vector is generated according to the word vector, including: The segmented text data is input into a pre-trained word vector model, and each word in the text is vectorized by the word vector model to generate a word vector containing semantic information; The word vectors are processed by averaging or matrix concatenation to obtain sentence vectors that meet the input conditions of the classification model.
5. The method for classifying complaint tickets based on ensemble learning according to claim 2 is characterized in that: The method of using the sentence vector as a training sample, constructing a machine learning model by a gradient boosting method, optimizing parameters of the machine learning model by a grid search, and tuning the classification results in combination with a business expert model to generate the integrated model includes: Using oversampling technology to generate synthetic samples of a small number of categories for the sentence vector, and combining undersampling technology to simplify data of a large number of categories; The sampled data is divided into training set and test set to build a machine learning model based on gradient boosting method. The learning rate, number of iterations, maximum depth, sub-sample ratio and model parameters are adjusted by grid search. The classification results of the gradient boosting model are input into the business expert model, and the classification results are adjusted and optimized by combining industry rules and experience to generate an integrated model that includes the machine learning model and the business expert model.
6. The method for classifying complaint tickets based on ensemble learning according to claim 1 is characterized in that: The reviewing of the internal control defect classification results and iterative optimization of the integrated model through feedback data include: Identify errors in the classification results based on the annotations during the review process, and generate new annotation data based on the review results; Merge the returned annotated data with the existing data set, update the training set of the gradient boosting model and retrain it, adjust the rules and parameters of the business expert model, and adapt the integrated model to the characteristics of the new annotated data; The updated integrated model is used to classify the unlabeled data, and the classification results are submitted for review again for iterative optimization.
7. The method for classifying complaint tickets based on ensemble learning according to claim 1 is characterized in that: Also includes: Collect historical complaint ticket data of each user separately, extract the time, frequency and subject of complaints from different users, and standardize the complaint ticket data; Using a time series analysis model, a complaint prediction model is constructed based on each user's historical complaint ticket data. The complaint prediction model is used to identify each user's repeated complaint behavior patterns within a preset time period.
8. A complaint work order classification device based on ensemble learning, characterized in that: include: The preprocessing module is used to preprocess the acquired complaint ticket data, extract key fields and normalize text data; A model input module, used to input the pre-processed complaint work order data into a pre-trained integrated model, wherein the integrated model is obtained by optimizing the parameters of a machine learning model based on a gradient boosting method, and adjusting and fusing the classification results in accordance with the rules in combination with a business expert model; A data classification module, used to classify the complaint work order data according to the output result of the integrated model to obtain the internal control defect classification result of the complaint work order data; The iterative optimization module is used to review the internal control defect classification results and iteratively optimize the integrated model through feedback data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the complaint ticket classification method based on integrated learning as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the complaint ticket classification method based on integrated learning as described in any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the complaint ticket classification method based on integrated learning as described in any one of claims 1 to 7 are implemented.