Complaint text classification method and device

Through the complaint text classification method that integrates multi-layer neural network layer and multimodal features, the problem of insufficient classification efficiency and accuracy in the existing technology is solved, and adaptive complaint text classification is realized, which improves classification efficiency and accuracy, and adapts to different types of complaint texts.

CN120386868APending Publication Date: 2025-07-29CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510515895.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing complaint text classification methods have shortcomings in classification efficiency and accuracy, especially in dealing with diversity, complexity and real-time, and it is difficult to effectively deal with the problems of language organization confusion, text feature sparsity, and data set classification imbalance.

Method used

The classification model of multi-layer neural network layer is adopted to classify and predict multimodal features layer by layer, judge whether to terminate the prediction through uncertainty probability threshold, combine user information, graph structure features and text features to fusion, and correct bias using classification algorithms and background knowledge to build an adaptive complaint text classification system.

Benefits of technology

It significantly improves the efficiency and accuracy of complaint text classification, can quickly handle simple and complex complaint text, reduce computing resources, improve the reliability and accuracy of classification results, adapt to different needs, improve customer satisfaction and reduce operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386868A_ABST
    Figure CN120386868A_ABST
Patent Text Reader

Abstract

The invention provides a complaint text classification method and device. The method comprises the steps of determining multi-modal features of complaint texts of a user; based on multiple neural network layers of a classification model, performing classification prediction on the multi-modal features layer by layer, and taking the hierarchical prediction result of any neural network layer as a classification result under the condition that the uncertainty probability of the hierarchical prediction result of any neural network layer is smaller than a preset probability threshold value; otherwise, continuing to perform classification prediction on the multi-modal features based on the next neural network layer until a classification result is output. According to the method provided by the invention, the multi-modal features are applied through the classification model, and the uncertainty probability of the hierarchical prediction result of the neural network layer is compared with the preset probability threshold value, so that the classification prediction adaptive to the user complaint text is realized, the efficiency and accuracy of the classification prediction are greatly improved, and the reliability of the classification result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for classifying complaint texts. Background Art

[0002] In the communication industry, classifying complaint texts faces challenges of diversity, complexity, and real-time nature. At the same time, some difficult-to-solve problems are also encountered, such as chaotic language organization, sparsity of text features, and imbalance in dataset classification. Existing technical solutions include "manual complaint text classification based on rules", "complaint text classification based on statistical methods", and "complaint text classification based on a single machine learning or deep learning algorithm".

[0003] However, the classification efficiency of existing technical solutions is limited, and the ability to process deep semantic information of texts is also limited, resulting in the need to improve the accuracy of classification results. Summary of the Invention

[0004] The present invention provides a method and device for classifying complaint texts to solve the defects of low classification efficiency and low classification accuracy in classifying complaint texts in the prior art.

[0005] The present invention provides a method for classifying complaint texts, including: Obtaining a user complaint text; Determining multi-modal features of the user complaint text; Based on the multi-layer neural network layer of the classification model, classifying and predicting the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than a preset probability threshold, taking the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; When the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continuing to classify and predict the multi-modal features based on the next neural network layer until the classification result is output.

[0006] According to the method for classifying complaint texts provided by the present invention, the determining the multi-modal features of the user complaint text includes: Calculating the importance value of the word segmentation based on the word frequency and inverse document frequency of the word segmentation in the user complaint text; Performing classification prediction based on the importance value of the word segmentation to obtain a preliminary classification result; Extracting the basic multi-modal features of the user complaint text, and splicing to obtain the multi-modal features based on the preliminary classification result and the basic multi-modal features.

[0007] A classification method for complaint texts provided by the present invention, which performs classification prediction based on the importance degree value of word segmentation to obtain a preliminary classification result, including: Applying the importance degree value of word segmentation based on a classification algorithm to perform classification prediction to obtain a source classification result; Based on background knowledge, identifying deviations in the source classification result of the user complaint text to obtain a deviation detection result; Based on the true business information of the user complaint text, correcting the deviation detection result to obtain the preliminary classification result; The background knowledge is constructed based on the true business information and prior business information.

[0008] A classification method for complaint texts provided by the present invention, which extracts the basic multi-modal features of the user complaint text, including: Obtaining the user information corresponding to the user complaint text and extracting the user information features of the user information; Extracting the sentence-level text features of the user complaint text; Obtaining the relevant complaint texts of the user complaint text, constructing a graph structure based on the user complaint text and the users to which the relevant complaint texts belong as nodes and the calling and called relationship between the users as connection edges; Extracting the node features of each node in the graph structure; Based on the user information features, the sentence-level text features, and the node features, splicing to obtain the basic multi-modal features.

[0009] A classification method for complaint texts provided by the present invention, which extracts the node features of each node in the graph structure, including: Respectively taking each node as the starting node; Based on a random walk strategy, selecting the next node of the starting node until a random walk sequence of a preset length is generated; Taking the random walk sequences corresponding to each node as the node features of each node; The sampling function of the random walk strategy is determined based on the calling and called frequency between the users corresponding to the nodes.

[0010] A classification method for complaint texts provided by the present invention, the training steps of the classification model include: Obtaining teacher sample complaint texts and an initial classification model, each initial neural network layer of the initial classification model corresponds to a classifier, taking the classifier corresponding to the last initial neural network layer as the initial teacher classifier; taking the classifiers corresponding to the initial neural network layers other than the last initial neural network layer as the initial student classifiers; Based on the teacher sample complaint text, perform parameter iteration on the multi-layer initial neural network layer of the initial classification model and the initial teacher classifier to obtain the multi-layer neural network layer and the teacher classifier; Based on the student sample complaint text and the soft labels, perform parameter iteration on the initial student classifier to obtain the student classifier; The soft labels are obtained by applying the student sample complaint text for prediction based on the multi-layer neural network layer and the teacher classifier.

[0011] The present invention also provides a classification device for complaint text, including: An acquisition unit that acquires the user's complaint text; A feature extraction unit that determines the multi-modal features of the user's complaint text; A first classification prediction unit that, based on the multi-layer neural network layer of the classification model, classifies and predicts the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, use the hierarchical prediction result of the any neural network layer as the classification result of the user's complaint text; A second classification prediction unit that, when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continues to classify and predict the multi-modal features based on the next neural network layer until the classification result is output.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the classification method of the complaint text as described in any one of the above.

[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the classification method of the complaint text as described in any one of the above.

[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the classification method of the complaint text as described in any one of the above.

[0015] The classification method and device for complaint texts provided by the present invention classify and predict multi-modal features layer by layer through the multi-layer neural network layer of the classification model. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, the hierarchical prediction result of any neural network layer is used as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of any neural network layer is not less than the preset probability threshold, the multi-modal features are continuously classified and predicted based on the next neural network layer until the classification result is output, realizing the adaptive classification based on the user complaint text, greatly improving the efficiency and accuracy of the classification prediction, and ensuring the reliability of the classification result. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the implementation of the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 FIG. 9 is one of the flowcharts of the classification method for complaint texts provided by the present invention; Figure 2 FIG. 12 is the flowchart of obtaining the preliminary classification result provided by the present invention; Figure 3 FIG. 15 is the flowchart of the multi-modal feature fusion method provided by the present invention; Figure 4 FIG. 18 is the structural diagram of the initial classification model provided by the present invention; Figure 5 FIG. 21 is the training flowchart of the student classifier provided by the present invention; Figure 6 FIG. 24 is another flowchart of the classification method for complaint texts provided by the present invention; Figure 7 FIG. 27 is the structural diagram of the classification device for complaint texts provided by the present invention; Figure 8 FIG. 30 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0019] It should be noted that for the problem of complaint text classification, the existing technical solutions include: Rule-based method: It usually relies on manually formulated rules or templates to classify complaint texts. For example, by identifying keywords, phrases or sentence structures in the text to determine the category of the complaint. This method is simple and intuitive, and easy to understand, but this method has great limitations. First, it is time-consuming and laborious, and a large number of rules need to be manually written and maintained. Second, the rules often only target known formats and patterns, and may not be accurate enough for newly added complaint text formats and patterns. For example, when the content of the complaint changes, the rules may need to be rewritten or adjusted.

[0020] Statistics-based method: It mainly relies on the statistical information of words in the text. It first constructs a statistical word list, which contains all the words that appear in the text and summarizes their relevant statistical information. This method has good interpretability, but has certain requirements for the size of the dataset to ensure the accuracy and reliability of the statistical information. Moreover, this method has limitations in dealing with the problem of feature sparsity and the deep semantic information of the text.

[0021] Single machine learning or deep learning-based method: This method is a text classification technology with high automation and strong adaptability. It usually encodes the text into vectors, and then uses a single machine learning model, such as Naive Bayes, Support Vector Machine or Logistic Regression, etc., or a deep learning model, such as Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN), etc., to automatically learn the features in the text and classify them. This method can effectively process deep semantic information, but in practical applications, in many cases, it works based on a single modality and a single algorithm model, resulting in poor accuracy. Some deep learning algorithms also have problems such as a large number of parameters and long running time.

[0022] Therefore, in view of the above problems, the present invention provides a method for classifying complaint texts to achieve more accurate, faster and more efficient classification of complaint texts. Figure 1 is one of the flow diagrams of the method for classifying complaint texts provided by the present invention, as Figure 1 shown, this method includes: Step 110, obtaining the user's complaint text; Specifically, user complaint texts can be collected from multiple channels such as customer service platforms, social media, emails, online forms, etc. Or, it can be text obtained by transcribing the phone voice of customer complaint business issues. Or, it can also be descriptive text generated by machine customer service or human customer service.

[0023] Step 120, determine the multi-modal features of the user complaint text; Here, the multi-modal features can include text features extracted based on the user complaint text, numerical features extracted based on the user information corresponding to the user complaint text, and graph features of the graph structure constructed based on the user complaint text. Thus, the multi-modal features here can be used to reflect the text semantic information of the user complaint text and related background information.

[0024] Specifically, the multi-modal features include numerical features, graph features, and text features. Among them, for the numerical information in the user information, such as age, income, etc., the numerical information can be directly used as numerical features; for the text information in the user information, such as occupation, region, etc., it can be converted into numerical features through one-hot encoding or label encoding techniques for subsequent model processing. Secondly, for the graph feature extraction part, the Node2Vec algorithm can be used. By simulating the biased random walk process of nodes on the graph to learn the embedding representation of nodes, the learned node embedding vectors can be used as graph features. It can be understood that graph features can capture the structural and relational information of nodes in the graph. Finally, for the text features of the user complaint text, the BERT (Bidirectional Encoder Representations from Transformers) model can be used to perform embedding learning on the user complaint text. Through the Embedding layer of the BERT model, the semantic information in the user complaint text can be captured and converted into vector form as text features.

[0025] Finally, the numerical features, graph features, and text features can be concatenated to obtain a more information-rich feature vector, that is, the multi-modal features of the user complaint text, so as to facilitate subsequent accurate complaint text classification based on the multi-modal features.

[0026] Step 130, based on the multi-layer neural network layer of the classification model, classify and predict the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, use the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; Here, the classification model is constructed based on multiple neural network layers, and each neural network layer corresponds to a classifier for classifying and predicting the features output by that neural network layer. In one embodiment, the classification model (Speed-Transformer model) mainly consists of a backbone and branches. The backbone part mainly consists of three parts: an Embedding layer, an encoder (multiple neural network layers) composed of Transformers, and a backbone classifier. The branch part is to add a branch classifier after each Transformer layer, that is, each neural network layer corresponds to a classifier.

[0027] In addition, the hierarchical prediction result here refers to the probability distribution of the prediction result output by any neural network layer. The uncertainty probability here represents the uncertainty estimate of the hierarchical prediction result of this layer, which is used to reflect the reliability of the hierarchical prediction result of this layer. The smaller the value of the uncertainty probability, the more likely it is to be correctly classified.

[0028] Specifically, the multi-modal features of the user complaint text can be input into the classification model, and through the multiple neural network layers of the classification model, the multi-modal features are classified and predicted layer by layer, and the hierarchical prediction result and the uncertainty probability of this hierarchical prediction result are output.

[0029] Here, the uncertainty probability of the hierarchical prediction result can be calculated through the following formula, as shown in the following formula: In the formula, represents the uncertainty probability of the hierarchical prediction result; N represents the total number of classification categories; represents the hierarchical prediction result.

[0030] Then, by comparing the uncertainty probability of the hierarchical prediction result of each neural network layer with a preset probability threshold. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, it means that the credibility of the hierarchical prediction result of this neural network layer is relatively high, and then the hierarchical prediction result of any neural network layer can be used as the classification result of the user complaint text, thus avoiding the need for further calculation. It should be noted that for relatively simple user complaint texts, high-confidence classification results can be output through the first few neural network layers, greatly reducing the calculation amount.

[0031] Step 140, when the uncertainty probability of the hierarchical prediction result of any neural network layer is not less than the preset probability threshold, continue to classify and predict the multi-modal features based on the next neural network layer until the classification result is output.

[0032] Specifically, when the uncertainty probability of the hierarchical prediction result of the neural network layer is not less than the preset probability threshold, the features output by the previous neural network layer are continuously classified and predicted based on the next neural network layer. If the uncertainty probability of the hierarchical prediction result of the next neural network layer is less than the preset probability threshold, the hierarchical prediction result of the next neural network layer is used as the classification result. If the uncertainty probability of the hierarchical prediction result of the next neural network layer is not less than the preset probability threshold, it is continuously input to the next neural network layer for classification and prediction until the uncertainty probability of the hierarchical prediction result of a certain neural network layer is less than the preset probability threshold, or the hierarchical prediction result of the last neural network layer is used as the classification result. The classification result here can be Home Business-Product Quality-Internet TV-After-sales Service; or Mobile Business-Business Marketing-Tariff Plan-Business Rules.

[0033] It should be noted that the preset probability threshold here is determined by weighing the relationship between speed and performance. When the obtained uncertainty probability is less than the preset probability threshold, it means that the prediction probability is very likely to be the true label value, so it is directly output as the final result. When the uncertainty probability is not less than the preset probability threshold, it means that the prediction probability cannot predict which label it is, so it is used as the input of the next layer of Transformer to continue the prediction. Thus, for the classification of simple user complaint texts, a relatively accurate result can be obtained by using the classification model in the embodiment of the present invention in the bottom classifier, and a relatively accurate classification result can also be obtained from the output of the top classifier for complex complaint texts, so as to achieve the purpose of adaptive reasoning, significantly improve the inference speed of the model, save computing resources, adapt to different requirements, and improve the robustness of the model.

[0034] It should also be noted that, compared with the prior art, for simple or complex user complaint texts, prediction and classification need to be performed through a complete algorithm model by using a fixed algorithm model. The method provided by the embodiment of the present invention can obtain the classification result with less computational effort for simple user complaint texts; for complex user complaint texts, the classification result can also be output in a shorter time. Such an inference process can greatly shorten the running time of the model, improve the classification effect, and ensure the accuracy of the classification result. In other words, the classification model provided by the embodiment of the present invention can realize adaptive classification prediction based on user complaint texts, greatly improving the classification efficiency and accuracy of user complaint texts. Moreover, through this adaptive mechanism, the classification model can effectively process simple samples quickly and output them in advance while keeping the basic structure of the model unchanged, and can achieve fast and accurate classification of complaint texts, helping operators quickly identify and handle various complaint problems, and thus not only improving customer satisfaction and loyalty, but also reducing the operation cost of operators and improving operation efficiency.

[0035] The method provided by the embodiment of the present invention classifies and predicts multimodal features layer by layer through the multi-layer neural network layer of the classification model. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, the hierarchical prediction result of any neural network layer is used as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of any neural network layer is not less than the preset probability threshold, continue to classify and predict the multimodal features based on the next neural network layer until the classification result is output, realizing the adaptive classification based on the user complaint text, greatly improving the efficiency and accuracy of classification prediction, and ensuring the reliability of the classification result.

[0036] Based on any of the above embodiments, step 120 includes: Calculate the importance value of the word segmentation based on the word frequency and inverse document frequency of the word segmentation in the user complaint text; Perform classification prediction based on the importance value of the word segmentation to obtain a preliminary classification result; Extract the basic multimodal features of the user complaint text, and splice them based on the preliminary classification result and the basic multimodal features to obtain the multimodal features.

[0037] Specifically, first, text preprocessing can be performed on the user complaint text. For example, regular expressions are used to remove illegal characters in the user complaint text, including removing HTML tags. Secondly, string processing methods or regular expressions are used to remove punctuation marks, special symbols, etc. in the user complaint text. Then, according to the stop word list, words without practical meaning in the user complaint text, such as "de", "shi", "zai", etc., are removed. Finally, the Jieba word segmentation tool is used to segment the user complaint text into individual word segments.

[0038] Then, the word frequency of each word segment in the user complaint text can be obtained by counting the number of times it appears, denoted as TF. The word frequency here can be calculated by the following formula, as shown below: In the formula, represents the word frequency of the i-th word segment, that is, the number of times the i-th word segment appears in the j-th user complaint text; represents the total number of times all word segments appear in the user complaint text j.

[0039] Secondly, by traversing the sample user complaint text set, all non-repeated words or phrases can be extracted to construct a vocabulary. Then, by traversing the entire sample user complaint text set, the number of occurrences of each segmented word in the user complaint text can be counted in how many sample user complaint texts, and the inverse document frequency of each segmented word can be calculated. Here, the inverse document frequency of the segmented word can be calculated by the following formula, as shown below: In the formula, represents the inverse document frequency of the i-th segmented word; represents the total number of files in the corpus; represents the number of files containing the segmented word , that is, the number of files of .

[0040] Furthermore, by multiplying the word frequency and the inverse document frequency of the segmented words in the user complaint text, the importance value of each segmented word in the user complaint text can be obtained, denoted as the TF-IDF value. Here, the importance value TF-IDF can be calculated by the following formula, as shown below: Next, the importance values of each segmented word in the user complaint text can be converted into a feature matrix, then a feature vector corresponding to the user complaint text can be obtained, and each value in the matrix corresponds to the importance value TF-IDF of a segmented word. Then, through the gradient boosting decision tree algorithm, the feature vector of the user complaint text can be used for classification prediction to obtain a preliminary classification result. The preliminary classification result here can be individual, government and enterprise, family, etc. For example, by using the pre-trained XGBoost classifier and applying the feature vector of the user complaint text for classification prediction, a preliminary classification result can be obtained.

[0041] The training process of the XGBoost classifier here includes: First, use the feature matrix (TF-IDF features) of the training set and the corresponding labels to train the XGBoost classifier. XGBoost is an efficient gradient boosting decision tree algorithm. It is an improvement on the original GBDT, which greatly enhances the model performance. As a forward additive model, the core of XGBoost is to adopt the ensemble idea, integrating multiple weak learners into a strong learner through a certain method. That is, multiple trees jointly make decisions, and the result of each tree is the difference between the target value and the prediction results of all previous trees, and all the results are accumulated to obtain the final result, so as to improve the classification effect of the entire model. In addition, the softmax function can be used as the activation function of the output layer, and the cross-entropy loss is selected as the optimization objective. Then, adjust the parameters of XGBoost to optimize the model performance. Commonly used parameters include the learning rate, the depth of the tree, the subsample ratio, etc. Finally, evaluate the model performance through cross-validation and select the best parameter combination. Among them, the cross-entropy loss can be calculated through the following formula, as shown below: In the formula, represents the cross-entropy loss; n represents the total number of sample user complaint texts; K represents the total number of classification categories; represents the indicator variable. If the i-th sample belongs to the k-th category, then = 1, otherwise = 0; represents the probability that the i-th text predicted by the XGBoost classifier belongs to the k-th category, and this probability is calculated through the softmax function. Specifically, assume that the original score output by the XGBoost classifier is , then the softmax prediction value can be calculated through the following formula, as shown below: In the formula, represents the original score that the i-th text belongs to the j-th category.

[0042] In addition, in the evaluation and application stage of the XGBoost classifier, the trained XGBoost classifier can be evaluated using the test set, and metrics such as classification accuracy, precision, recall, and F1 value are calculated. According to the evaluation results, adjust the model parameters to further optimize the model performance to obtain the final XGBoost classifier.

[0043] Finally, after obtaining the preliminary classification result of the user complaint text, the basic multi-modal features of the user complaint text can be extracted. Specifically, the basic multi-modal features can be constructed based on the text features extracted from the user complaint text, the numerical features extracted from the user information corresponding to the user complaint text, and the graph features of the graph structure constructed based on the user complaint text. Then, the preliminary classification result and the basic multi-modal features can be concatenated to obtain the multi-modal features for subsequent classification prediction.

[0044] It should be noted that adding a pre-classification gating layer before classifying the user complaint text can conduct targeted preliminary screening of the user complaint text at an early stage of the classification process, thereby significantly improving the pertinence and accuracy of the subsequent text classification results, not only improving the efficiency of the entire classification process, but also ensuring the accuracy and reliability of the classification results.

[0045] It should be noted that in order to further improve the accuracy of the preliminary classification result, based on any of the above embodiments, classifying and predicting based on the importance value of the word segmentation to obtain the preliminary classification result, including: Applying the importance value of the word segmentation based on the classification algorithm to perform classification prediction to obtain the source classification result; Based on background knowledge, identifying the deviation of the source classification result of the user complaint text to obtain the deviation detection result; Based on the true business information of the user complaint text, correcting the deviation detection result to obtain the preliminary classification result; The background knowledge is constructed based on the true business information and the prior business information.

[0046] Here, the true business information refers to the personal basic business information, personal rights and interests information, business types already handled, home broadband TV, government and enterprise complaint records, etc. of the user corresponding to the user complaint text. In addition, the prior business information here can be standard business knowledge. Thus, the background knowledge here can be constructed through the true business information and the prior business information.

[0047] Specifically, first, the source classification result can be obtained through classification prediction by applying the importance values of all word segmentations of the user complaint text using a pre-trained XGBoost classifier. Then, through background knowledge, the deviation of the source classification result of the user complaint text can be identified, and the classification result that does not conform to the background knowledge can be identified, that is, the deviation detection result can be obtained. Further, through the true business information of the user complaint text, the deviation detection result can be automatically corrected to achieve automatic adjustment and optimization of the classification, and the preliminary classification result can be obtained.

[0048] In one embodiment, Figure 2It is a schematic flowchart of the process for obtaining the preliminary classification result provided by the present invention. As Figure 2 shown, the method includes: First, obtain the complaint text (user complaint text), and then calculate the TF-IDF values of each word segment in the complaint text. Further, by inputting the TF-IDF values of all word segments of the complaint text into the XGBoost classifier, the source classification result is output by the XGBoost classifier, that is, it belongs to individuals, families, or government and enterprises. Then, perform authenticity rectification on the source classification result to obtain the final preliminary classification result.

[0049] It should be noted that the classification correction based on the authenticity rectification model is the second layer of the pre-classification gating layer, aiming to rectify the source classification result of the first layer of the module. Specifically, mix real data such as the senior professional knowledge and rich experience in this field with the source classification result generated in the foregoing algorithm for in-depth rectification. By real-time tracking and evaluating the deviation between the source classification result and the real result, and using background knowledge, automatic adjustment and optimization of the classification are realized, thereby significantly improving the accuracy and reliability of the preliminary classification, and making the classification method of the complaint text more robust and adaptable.

[0050] The method provided by the embodiment of the present invention can effectively improve the accuracy and reliability of the preliminary classification of user complaint texts by performing authenticity rectification on the source classification result, thereby reducing unnecessary troubles and losses caused by algorithm misjudgment.

[0051] Based on any of the above embodiments, the extraction of the basic multi-modal features of the user complaint text includes: Obtain the user information corresponding to the user complaint text, and extract the user information features of the user information; Extract the sentence-level text features of the user complaint text; Obtain the relevant complaint texts of the user complaint text, and construct a graph structure based on the user complaint text and the users to which the relevant complaint texts belong as nodes and the calling relationship between the users as the connection edges; Extract the node features of each node in the graph structure; Based on the user information features, the sentence-level text features, and the node features, splice them to obtain the basic multi-modal features.

[0052] Specifically, for the user information features, the user information corresponding to the user of the user complaint text can be obtained, such as the user's age, income, occupation, region and other information. Then, the numerical user information can be directly used as the user information features, and the user information features of the non-numerical user information can be extracted through one-hot encoding or label encoding technology.

[0053] For the sentence-level text features of user complaint texts, the sentence-level text features can be extracted from user complaint texts through Embedding in the BERT model. Specifically, any sentence in the user complaint text can be input , with the sentence length being n. The sentence is converted into a vector sequence e through the Embedding layers, and the conversion process is shown as follows: e = Embedding(s) In the formula, e represents the sum of word Embeddings, position Embeddings, and segment Embeddings. Then, layer-by-layer feature extraction can be performed through the Transformer blocks in the encoder, as shown in the following formula: In the formula, (i = -1, 0, 1,..., L - 1) represents the sentence-level text features output by the i-th layer, and L represents the total number of layers of the Transformer;; = e; represents the sentence-level text features output by the (i - 1)-th layer.

[0054] In addition, for the node features of the graph structure, first, the relevant complaint texts of the user complaint text can be sorted as needed. The users to whom the user complaint text and the relevant complaint texts belong are used as nodes, and connection edges are constructed based on the calling and called relationships, acquaintance, and whether there are similar complaints among users. Thus, a graph structure is constructed, where the nodes in the graph represent users and the edges represent the relationships between users.

[0055] Then, the node features of each node in the graph structure can be extracted through a graph neural network or a random walk strategy. Finally, the user information features, sentence-level text features, and node features can be concatenated to obtain the basic multi-modal features.

[0056] It should be noted that different from the previous single-modal complaint text classification, on the basis of user complaint texts, the user personal information features and the node features of the graph structure are fused, which can achieve cross-verification and complementarity between different modal information, enabling the classification model to obtain richer information and thus improving the accuracy of classification.

[0057] In an embodiment, Figure 3 is a schematic flowchart of the multi-modal feature fusion method provided by the present invention, as shown in Figure 3As shown in the figure, the method includes: First, obtain the user personal information features (user information features) extracted from the user's age, gender and other information, the graph features (node features of the graph structure) extracted from the graph structure constructed based on the relationship between the caller and the callee among users, and the complaint text features (sentence-level text features) extracted from the user complaint text information such as business marketing and home broadband. Then, fuse the features obtained above for multi-modal feature fusion to construct a richer feature vector, that is, the basic multi-modal feature. Thereby, the classification model can comprehensively utilize various feature information from different modalities to improve the accuracy and performance of prediction or classification.

[0058] The method provided by the embodiment of the present invention, based on the user complaint text, fuses the user's personal basic information features and the node features of the graph structure, and splices different types of features to form a longer feature vector, making full use of the complementarity between different modalities, enabling the feature vector to more comprehensively reflect all aspects of the user complaint, and further improving the accuracy of text classification.

[0059] Based on any of the above embodiments, the extraction of the node features of each node in the graph structure includes: Respectively use each of the nodes as the starting node; Based on the random walk strategy, select the next node of the starting node until a random walk sequence of a preset length is generated; Use the random walk sequences corresponding to each of the nodes as the node features of each of the nodes; The sampling function of the random walk strategy is determined based on the call frequency between the users corresponding to the nodes.

[0060] Specifically, first, all nodes can be selected as the starting nodes respectively to ensure that all nodes can be taken into account. For any node as the starting node, the next node of the starting node can be selected through the random walk strategy until a random walk sequence of a preset length is generated. In detail, during the random walk process, the next-hop node is selected according to a certain strategy. Among them, two parameters p and q are introduced in the Node2Vec algorithm to adjust the walk strategy. p controls the probability of returning to the previous node, and q controls the probability of walking away from the previous node. By adjusting the values of p and q, the depth and breadth of the walk can be controlled. Finally, repeat the above process until the preset walk length is reached to generate a random walk sequence. It should be noted that considering the information such as the call values of users in business, the Alisa-sample and randomness can be used as the sampling function of the biased random walk to obtain a biased random walk sequence.

[0061] Finally, the random walk sequence corresponding to the node can be used as the node feature of the node, and then the node features of all nodes in the graph structure can be used as the graph feature of the graph structure.

[0062] It should be noted that by learning the embedding representation of the nodes and using the node features of all nodes in the graph structure as the graph feature of the graph structure, the local and global structure information of the nodes can be reflected simultaneously, thereby improving the accuracy of subsequent text classification based on the node features of the graph.

[0063] Based on any of the above embodiments, the training steps of the classification model include: Obtain the teacher sample complaint text and the initial classification model. Each initial neural network layer of the initial classification model corresponds to a classifier, and the classifier corresponding to the last initial neural network layer is used as the initial teacher classifier; the classifiers corresponding to the initial neural network layers other than the last initial neural network layer are used as the initial student classifiers; Based on the teacher sample complaint text, perform parameter iteration on the multi-layer initial neural network layers of the initial classification model and the initial teacher classifier to obtain the multi-layer neural network layers and the teacher classifier; Based on the student sample complaint text and the soft labels, perform parameter iteration on the initial student classifier to obtain the student classifier; The soft labels are obtained by predicting the student sample complaint text based on the multi-layer neural network layers and the teacher classifier.

[0064] Here, Figure 4 is a schematic structural diagram of the initial classification model provided by the present invention. As Figure 4 shown, the initial classification model may include multiple layers of initial neural network layers (Transformer0, Transformer1,..., Transformer10, Transformer11), and each layer of the initial neural network layer includes an initial classifier. In the training stage, the classifier corresponding to the last initial neural network layer can be used as the initial teacher classifier (teacher-Classifier), and the classifiers corresponding to the initial neural network layers other than the last initial neural network layer are used as the initial student classifiers (Student-Classifier0, Student-Classifier1,..., Student-Classifier10).

[0065] Specifically, in the training phase of the classification model, first, obtain the teacher sample complaint texts and the initial classification model. It should be noted that the initial classification model here can be obtained through pre-training. During the pre-training process, through the two tasks of Masked Language Modeling (MLM) and Next Sentence Prediction (NSP), the multi-layer initial neural network of the initial classification model learns rich language knowledge and context representations, laying a solid foundation for subsequent specific tasks.

[0066] Then, in the fine-tuning phase, an initial teacher classifier can be added to the last layer of the multi-layer initial neural network of the initial classification model. Through the pre-labeled teacher sample complaint texts, parameter iteration is performed on the multi-layer initial neural network layers of the initial classification model and the initial teacher classifier to obtain the multi-layer neural network layers and the teacher classifier. Specifically, the teacher sample complaint texts can be input into the multi-layer initial neural network layers and the initial teacher classifier, and the sample teacher prediction classification is output through the initial teacher classifier. Then, the teacher prediction loss can be calculated by using the sample teacher prediction classification and the label corresponding to the sample teacher prediction classification. By minimizing the teacher prediction loss and through the backpropagation algorithm, the parameters of the multi-layer initial neural network layers and the initial teacher classifier are optimized until the performance of the model reaches the preset requirements or the preset number of training rounds.

[0067] Here, the teacher prediction loss can be calculated through the NLLLoss loss function, as shown in the following formula: In the formula, represents the teacher prediction loss; represents the total data volume of the teacher sample complaint texts; represents the classification label of the k-th teacher sample complaint text after one_hot encoding; represents the probability distribution of the teacher prediction classification result. That is, the teacher prediction loss is multiplied by the probability distribution after activation by the log_softmax function, then taking the average value, and finally taking the negative value to obtain.

[0068] Furthermore, after the backbone network (multi-layer initial neural network layers and teacher classifier) is trained, freeze the parameters of the multi-layer initial neural network layers and the teacher classifier in the fine-tuning phase, and the output of the teacher classifier can be used as high-quality soft labels. These soft labels not only contain the original feature embeddings but also contain generalization knowledge. Subsequently, these soft labels are used to train the student classifier.

[0069] Thus, the unlabeled student sample complaint text can be input into the initial student classifier. The soft label of the student sample complaint text is output by the initial student classifier, and the student prediction classification result of the student sample complaint text is output by the initial student classifier. Then, the student prediction loss can be calculated based on the soft label corresponding to the student sample complaint text and the student prediction classification result output by the initial student classifier. The initial student classifier is iteratively optimized through the student prediction loss to obtain the final student classifier.

[0070] Figure 5 is a schematic diagram of the training process of the student classifier provided by the present invention. As Figure 5 shown, the branch structure of the initial classification model includes 12 Transformer layers and 11 student classifiers. The input of the initial classification model is the multi-modal features of the student sample complaint text. First, it passes through the Embedding Layer, and then through Transformer 0. The output features of Transformer 0 are predicted through Student-Classifier 0. Here, it is the prediction probability of the first layer of the student classifier. There is a gap between the current prediction probability and the soft label, and this gap is measured by the loss function KL divergence. Since the student classifiers are independent of each other, the student prediction classification results are respectively compared with the soft label , and the difference is measured by KL divergence. KL divergence, that is, relative entropy, can measure the difference between the two probability distributions of the student classifier and the teacher classifier. The calculation formula for the student prediction loss for any student classifier is as follows: In the formula, represents the student prediction loss; represents the total data volume of the student sample complaint text; represents the student prediction classification result of the i-th student sample complaint text; represents the soft label of the j-th (j = i) student sample complaint text, that is, the teacher prediction classification result of the i-th student sample complaint text output by the teacher classifier.

[0071] At the same time, since there are 11 student classifiers in the initial classification model in total, the losses of each layer of the student classifier are added up to form the final student loss function of the model. The final student loss function is as follows: In the formula, represents the final student loss of the 11-layer student classifier, where N represents the number of layers of the neural network.

[0072] It should be noted that in the figure, the black arrow represents the direction of the data flow in the model, that is, the direction of forward propagation, and the red arrow represents that the student classifier learns feature information from the teacher classifier through the loss function. In the backpropagation stage, the sum of the KL divergence losses of all student classifiers can be calculated as the final loss function of the model. By minimizing the difference between the predicted probability distribution of the student classifier and the predicted probability distribution of the teacher classifier, the parameters of the student classifier are updated, so that the student classifier can learn the knowledge of the teacher classifier. Finally, the student classifier will be able to independently classify new complaint text data and show good performance.

[0073] Figure 6 This is the second flowchart of the classification method for complaint text provided by the present invention. As Figure 6 shown, the method includes: The first stage is the pre-classification gating layer. First, the complaint text (user complaint text) is obtained. Then, the TF-IDF values of each word segment in the complaint text are calculated. Further, by inputting the TF-IDF values of all word segments of the complaint text into the XGBoost classifier, the source classification result is output by the XGBoost classifier, that is, it belongs to an individual, a family, or a government enterprise. Then, the authenticity of the source classification result is corrected to obtain the final preliminary classification result.

[0074] The second stage is the accelerated classification layer of the Speed-Transformer model (classification model) based on multi-modal feature fusion. Specifically, by extracting the numerical features / text features, graph features, and text features of the user information corresponding to the user complaint text, multi-modal feature fusion is performed on the numerical features / text features, graph features, and text features, and the fused basic multi-modal features and the preliminary classification result are input into the Speed-Transformer model together, and the final classification result is output by the Speed-Transformer model.

[0075] It should be noted that first, automatic text classification is performed by using the pre-classification gating layer to quickly identify the subject type (individual, family, government enterprise). Secondly, multi-modal features are integrated, including complaint text features, user personal information features, and graph features, to capture more comprehensive information. Finally, the Speed-Transformer model is adopted, and an adaptive optimization strategy and an early prediction strategy are introduced to significantly improve the model training and inference speed while ensuring performance. At the same time, the model structure is flexible and can adapt to datasets of different scales and complexities, with good scalability. Therefore, the method provided by the embodiments of the present invention can well meet the needs of user complaint text classification in the communication industry.

[0076] It is understandable that, first of all, in order to quickly and accurately preliminarily identify complaint texts, a pre-classification gating layer is introduced. This layer uses the TF-IDF algorithm to extract key features from the text and combines with the XGBoost classification algorithm for preliminary classification. At the same time, it is supplemented by a authenticity correction model to ensure the accuracy of the classification results and reduce potential problems caused by misjudgment.

[0077] Secondly, in order to more comprehensively understand the background and context of the complaint, the method provided by the embodiments of the present invention fuses two different modal information features, namely user personal basic information features and graph features, on the basis of the complaint text features. This multi-modal feature fusion method can make full use of various information sources and improve the accuracy and efficiency of classification. Among them, graph features are extracted through methods such as Node2Vec to capture complex structures and relationships in the text, providing richer context information for complaint text classification.

[0078] Next, the above three types of features are concatenated to construct a more abundant feature vector that combines multiple modal information. This step not only realizes the effective fusion of multi-modal features, but also enables the model to make full use of feature information from different sources, thereby improving the accuracy and performance of text classification.

[0079] Finally, through a novel, speed-adjustable Speed-Transformer model, it has an adaptive inference process. Under different requirements, the inference speed can be flexibly adjusted while avoiding redundant calculations of samples. Specifically, each layer of this classification model is designed with a classifier, and these classifiers are responsible for outputting a prediction probability p for each input sample. The prediction probability p represents the uncertainty estimate of the current layer for the input sample. At the same time, a threshold is set during the prediction process of the Speed-Transformer. If the prediction uncertainty probability of a certain layer is lower than this threshold, then the prediction result of this layer is directly used as the final output, thus avoiding the need for further calculations. On the contrary, if the prediction uncertainty probability is higher than the threshold, then the sample is passed to the classifier of the next layer for further analysis. Through this adaptive mechanism, the Speed-Transformer model can effectively process and output simple samples quickly while keeping the basic structure of the model unchanged, thereby significantly improving the overall inference speed.

[0080] In the design of the classification model provided by the embodiments of the present invention, for simple complaints, users can obtain immediate feedback, and for complex complaints, users can also obtain accurate responses and solutions within a shorter time. This instant feedback and processing mechanism will greatly improve user satisfaction and loyalty. At the same time, it provides a framework for continuous improvement for the operator. As new complaints keep pouring in, the model can continuously learn and optimize to improve the accuracy and efficiency of classification. This will help the operator continuously improve customer service quality and maintain a competitive advantage.

[0081] Based on any of the above embodiments, Figure 7 is a schematic structural diagram of the classification device for complaint texts provided by the present invention, as Figure 7 shown, the device includes: An acquisition unit 710 that acquires user complaint texts; A feature extraction unit 720 that determines multimodal features of the user complaint text; A first classification prediction unit 730 that, based on the multi-layer neural network layer of the classification model, classifies and predicts the multimodal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than a preset probability threshold, the hierarchical prediction result of the any neural network layer is used as the classification result of the user complaint text; A second classification prediction unit 740 that, when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continues to classify and predict the multimodal features based on the next neural network layer until the classification result is output.

[0082] The device provided by the embodiments of the present invention classifies and predicts multimodal features layer by layer through the multi-layer neural network layer of the classification model. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, the hierarchical prediction result of the any neural network layer is used as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of any neural network layer is not less than the preset probability threshold, continue to classify and predict the multimodal features based on the next neural network layer until the classification result is output, realizing adaptive classification based on user complaint texts, greatly improving the efficiency and accuracy of classification prediction, and ensuring the reliability of the classification result.

[0083] Based on any of the above embodiments, the feature extraction unit is specifically used for: Calculating the importance value of the word segmentation based on the word frequency and inverse document frequency of the word segmentation in the user complaint text; Performing classification prediction based on the importance value of the word segmentation to obtain a preliminary classification result; Extract the basic multimodal features of the user complaint text, and based on the preliminary classification result and the basic multimodal features, splice to obtain the multimodal features.

[0084] Based on any of the above embodiments, the feature extraction unit is further specifically configured to: Apply the importance value of the word segmentation based on the classification algorithm to perform classification prediction to obtain the source classification result; Based on the background knowledge, identify the deviation of the source classification result of the user complaint text to obtain the deviation detection result; Based on the true business information of the user complaint text, correct the deviation detection result to obtain the preliminary classification result; The background knowledge is constructed based on the true business information and the prior business information.

[0085] Based on any of the above embodiments, the feature extraction unit is further specifically configured to: Obtain the user information corresponding to the user complaint text, and extract the user information features of the user information; Extract the sentence-level text features of the user complaint text; Obtain the relevant complaint text of the user complaint text, and based on the user complaint text and the users to which the relevant complaint text belongs as nodes, and based on the calling and called relationship between the users as the connection edges, construct a graph structure; Extract the node features of each node in the graph structure; Based on the user information features, the sentence-level text features, and the node features, splice to obtain the basic multimodal features.

[0086] Based on any of the above embodiments, the feature extraction unit is further specifically configured to: Respectively use each of the nodes as the starting node; Based on the random walk strategy, select the next node of the starting node until a random walk sequence of a preset length is generated; Use the random walk sequences corresponding to the respective nodes as the node features of the respective nodes; The sampling function of the random walk strategy is determined based on the calling and called frequency between the users corresponding to the nodes.

[0087] Based on any of the above embodiments, the apparatus further includes a training unit, and the training unit is specifically configured to: Obtain a teacher sample complaint text and an initial classification model. Each initial neural network layer of the initial classification model corresponds to a classifier. Use the classifier corresponding to the last initial neural network layer as the initial teacher classifier; use the classifiers corresponding to the initial neural network layers other than the last initial neural network layer as the initial student classifiers; Based on the teacher sample complaint text, perform parameter iteration on the multi-layer initial neural network layer of the initial classification model and the initial teacher classifier to obtain the multi-layer neural network layer and the teacher classifier; Based on the student sample complaint text and the soft labels, perform parameter iteration on the initial student classifier to obtain the student classifier; The soft labels are obtained by applying the student sample complaint text for prediction based on the multi-layer neural network layer and the teacher classifier.

[0088] Figure 8 An example of a schematic physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the classification method of the complaint text. The method includes: obtaining the user complaint text; determining the multi-modal features of the user complaint text; based on the multi-layer neural network layer of the classification model, performing classification prediction on the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, use the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continue to perform classification prediction on the multi-modal features based on the next neural network layer until the classification result is output.

[0089] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.

[0090] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the classification method of complaint texts provided by the above-mentioned various methods. The method includes: obtaining a user complaint text; determining multimodal features of the user complaint text; based on the multi-layer neural network layer of the classification model, classifying and predicting the multimodal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than a preset probability threshold, using the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continuing to classify and predict the multimodal features based on the next neural network layer until the classification result is output.

[0091] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the classification method of complaint texts provided by the above-mentioned various methods. The method includes: obtaining a user complaint text; determining multimodal features of the user complaint text; based on the multi-layer neural network layer of the classification model, classifying and predicting the multimodal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than a preset probability threshold, using the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continuing to classify and predict the multimodal features based on the next neural network layer until the classification result is output.

[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A classification method for complaint texts, characterized in that, Including: Obtain the user complaint text; Determine the multi-modal features of the user complaint text; Based on the multi-layer neural network layer of the classification model, classify and predict the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, use the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; When the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continue to classify and predict the multi-modal features based on the next neural network layer until the classification result is output.

2. The classification method of the complaint text according to claim 1, characterized in that, The determining the multi-modal features of the user complaint text includes: Based on the word frequency and inverse document frequency of the word segmentation in the user complaint text, calculate the importance value of the word segmentation; Based on the importance value of the word segmentation, perform classification prediction to obtain a preliminary classification result; Extract the basic multi-modal features of the user complaint text, and based on the preliminary classification result and the basic multi-modal features, splice to obtain the multi-modal features.

3. The classification method of the complaint text according to claim 2, characterized in that, The performing classification prediction based on the importance value of the word segmentation to obtain a preliminary classification result includes: Apply the importance value of the word segmentation based on a classification algorithm to perform classification prediction to obtain a source classification result; Based on background knowledge, identify the deviation of the source classification result of the user complaint text to obtain a deviation detection result; Based on the true business information of the user complaint text, correct the deviation detection result to obtain the preliminary classification result; The background knowledge is constructed based on the true business information and prior business information.

4. The classification method of complaint texts according to claim 2, characterized in that The extracting the basic multi-modal features of the user complaint text includes: Obtain the user information corresponding to the user complaint text, and extract the user information features of the user information; Extract the sentence-level text features of the user complaint text; Obtain the relevant complaint text of the user complaint text, construct a graph structure based on the user complaint text and the users to which the relevant complaint text belongs as nodes and the calling and called relationship between the users as connection edges; Extract the node features of each node in the graph structure; Based on the user information features, the sentence-level text features, and the node features, splice to obtain the basic multi-modal features.

5. The classification method of the complaint text according to claim 4, characterized in that, The extracting the node features of each node in the graph structure includes: Respectively use each node as the starting node; Based on the random walk strategy, select the next node of the starting node until a random walk sequence of a preset length is generated; Use the random walk sequences corresponding to each node as the node features of each node; The sampling function of the random walk strategy is determined based on the calling frequency between the users corresponding to the nodes.

6. The classification method of complaint texts according to any one of claims 1 to 5, characterized in that, The training steps of the classification model include: Obtain the teacher sample complaint text and the initial classification model. Each initial neural network layer of the initial classification model corresponds to a classifier. Use the classifier corresponding to the last initial neural network layer as the initial teacher classifier; use the classifiers corresponding to the initial neural network layers other than the last initial neural network layer as the initial student classifiers; Based on the teacher sample complaint text, perform parameter iteration on the multi-layer initial neural network layer of the initial classification model and the initial teacher classifier to obtain the multi-layer neural network layer and the teacher classifier; Based on the student sample complaint text and the soft labels, perform parameter iteration on the initial student classifier to obtain the student classifier; The soft labels are obtained by predicting using the student sample complaint text based on the multi-layer neural network layer and the teacher classifier.

7. A classification device for complaint texts, characterized in that, It includes: An acquisition unit that acquires the user complaint text; A feature extraction unit that determines the multi-modal features of the user complaint text; A first classification prediction unit that, based on the multi-layer neural network layer of the classification model, performs classification prediction on the multi-modal features layer by layer. When the uncertainty probability of the hierarchical prediction result of any neural network layer is less than the preset probability threshold, use the hierarchical prediction result of the any neural network layer as the classification result of the user complaint text; A second classification prediction unit that, when the uncertainty probability of the hierarchical prediction result of the any neural network layer is not less than the preset probability threshold, continues to perform classification prediction on the multi-modal features based on the next neural network layer until the classification result is output.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the classification method of the complaint text according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the classification method of the complaint text according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the classification method of the complaint text according to any one of claims 1 to 6.