Elevator fault description text classification method, system and computer-readable storage medium
Through deep learning methods, the dual classification model is trained, and the fastText and XGBoost models are used to classify elevator fault description text, solving the problems of cognitive differences and inconsistencies in the manual classification in the existing technology, and achieving efficient and accurate fault information management.
Patent Information
- Application Number
- CN202410436851.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-04-12
AI Technical Summary
In the prior art, elevator fault classification depends on manual, and there are problems such as cognitive differences and inconsistent classification standards, resulting in inconvenient fault information management and analysis.
Deep learning method is adopted to train a dual classification model through the expanded and preprocessed text training set, and the fastText word embedding model and XGBoost model are used to automatically classify elevator fault description text.
It realizes efficient and accurate classification of elevator fault description text, improves the accuracy and efficiency of fault information management, and reduces the dependence of manual classification.
Smart Images

Figure CN118035457B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of elevators, and in particular to a method, system and computer-readable storage medium for classifying elevator fault description texts. Background Art
[0002] At present, the elevator maintenance industry is gradually entering the process of informatization, and the elevator maintenance information carrier has also been migrated from the original paper records to various informatization platforms. Although the informatization of elevator maintenance records can promote the standardization and regularization of elevator information management and improve data traceability, it also puts higher requirements on the quality of maintenance personnel's filling.
[0003] Traditionally, elevator fault classification mainly relies on manual classification by maintenance personnel. Maintenance personnel arrive at the elevator fault site and survey the situation. After inspecting and repairing the elevator, they fill in the elevator fault text and fault category. The traditional classification method has the following problems: First, maintenance personnel have different understandings of faults, and the same fault may be classified into several different fault categories; in addition, the description of the same category of faults will also be different due to different personnel; finally, the fault classification standards vary from maintenance company to maintenance company, and the same fault belongs to different categories in different companies, which causes inconvenience in the management and analysis of elevator fault information.
[0004] In the prior art, there are studies on the classification of elevator fault texts, such as the "A Method and System for Elevator Customer Service Text Classification" disclosed in Chinese patent CN109947941A, which uses BERT fused capsule network to encode text and CapsNet network to classify text, thereby classifying the user's elevator fault complaint text. However, the vector encoding of BERT takes into account the context, and the model structure is large, which is not conducive to the classification of large amounts of text. Summary of the invention
[0005] The purpose of the present invention is to propose a method, system and computer-readable storage medium for classifying elevator fault description texts. By expanding and preprocessing the text training set, adopting a deep learning method, and training a dual classification model, it is possible to efficiently and accurately realize the automatic classification of batches of elevator fault description texts, thereby improving the accuracy and efficiency of elevator fault information management.
[0006] To achieve this object, the present invention adopts the following technical solutions:
[0007] A method for classifying elevator fault description texts comprises the following steps:
[0008] S1: Obtain the original text set of elevator fault description text, expand and preprocess the original text set to obtain a text training set;
[0009] S2: Input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set, input the word embedding matrices into the XGBoost model for training, and obtain the target XGBoost model;
[0010] S3: inputting the text training set into the fastText model for training to obtain a target fastText model;
[0011] S4: inputting the elevator fault description text to be classified into the fastText word embedding model for encoding processing, obtaining the word embedding matrix of the elevator fault description text to be classified, and inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result;
[0012] S5: inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result;
[0013] S6: Obtain a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result.
[0014] Furthermore, the S1 comprises the following sub-steps:
[0015] S11: obtaining an original text set of elevator fault description texts, training a MOSS model with the original text set, and obtaining a fine-tuned MOSS model and elevator fault description texts;
[0016] S12: obtaining rare categories that need to be expanded according to the classification in the original text set, inputting the rare categories into the fine-tuned MOSS model to obtain supplementary texts, and expanding the elevator fault description text with the supplementary texts;
[0017] S13: The expanded elevator fault description text includes several categories of text sequences, each of which contains several words. The text sequence of each category is preprocessed to obtain an elevator fault description vocabulary list. The several categories of elevator fault description vocabulary lists are the text training sets.
[0018] Furthermore, in S13, the preprocessing performed on the text sequence of each category includes word segmentation and stop word filtering.
[0019] Furthermore, in S2, the text training set is input into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set, including the following steps:
[0020] S21: Divide the text training set into a test set, use the skipgram algorithm, and train the fastText word embedding model with the elevator fault description vocabulary list in the test set. The loss function used is: ;
[0021] In the formula, L is the loss value, T is the number of words in the text sequence, t is the word index of the text sequence, c is the preset sliding window size, j is the word index of the sliding window, is the likelihood function, The first t+j words, The first t words;
[0022] S22: Input the text training set into the trained fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set.
[0023] Furthermore, in S2, the word embedding matrix is input into the XGBoost model for training to obtain a target XGBoost model, including the following steps:
[0024] S23: extracting a number of words from the word embedding matrix by a random method with replacement, and obtaining classification prediction values of the words according to the words and a preset prediction algorithm, wherein the classification prediction values include classification prediction probability distribution values of a number of categories, and the classification model includes a number of decision trees;
[0025] The prediction algorithm is:
[0026] ;
[0027] In the formula, i is the index of the word, For the i The classification prediction value of words, is the total number of decision trees, is the index of the decision tree, Indicates k A decision tree, Indicates i words, Indicates i Words through k The result predicted by the decision tree;
[0028] S24: According to the classification prediction values of the plurality of words, the classification true probability distribution values and a preset gradient descent algorithm, the first classification model is trained to obtain a target first classification model, wherein the gradient descent algorithm is: ;
[0029] In the formula, is the loss value, are model parameters, n is the total number of words, is the log-likelihood loss function, For the i The true classification value of words, is the regularization function, For the i The classification prediction probability distribution value of each word.
[0030] Furthermore, the S2 also includes S25: using the RandomizedSearchCV method to optimize the parameters of the trained XGBoost model, and using the optimized model as the target XGBoost model.
[0031] Further, the text training set is input into the fastText model for training to obtain a target fastText model, including the following steps:
[0032] S31: Using a supervised learning mode, the text training set is input into the fastText model for training;
[0033] S32: using the RandomizedSearchCV method to optimize the parameters of the trained fastText model, and using the optimized model as the target fastText model.
[0034] Furthermore, in S6, the first classification result and the second classification result both include classification prediction probability distribution values of several categories;
[0035] Obtaining a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result includes the following steps:
[0036] S61: In the first classification result and the second classification result, the classification prediction probability distribution values of the same category are accumulated to obtain classification prediction probability distribution accumulated values of several categories;
[0037] S62: Compare the classification prediction probability distribution cumulative values of several categories, and take the category corresponding to the largest classification prediction probability distribution cumulative value as the final classification result of the elevator fault description text to be classified.
[0038] An elevator fault description text classification system, the elevator fault description text classification system is used to complete the above-mentioned elevator fault description text classification method;
[0039] The elevator fault description text classification system comprises:
[0040] A text preprocessing module is used to expand and preprocess the original text set to obtain a text training set; and to obtain elevator fault description text to be classified;
[0041] An encoding module is used to input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set;
[0042] A first classification module, used for inputting the word embedding matrix into an XGBoost model for training to obtain a target XGBoost model, and used for inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result;
[0043] A second classification module, used for training a fastText model using the text training set to obtain a target fastText model, and for inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result;
[0044] The classification result output module is used to obtain the final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result, and output the final classification result.
[0045] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned elevator fault description text classification method are implemented.
[0046] The technical solution provided by the present invention may include the following beneficial effects:
[0047] 1. The present invention can combine the advantages of parallel processing, strong interpretability and result stability of the XGBoost model, and the advantages of fast training speed and high accuracy of the fastText model, and adopt a dual classification method to make the final classification result stable and highly accurate, thereby improving the accuracy and efficiency of elevator fault description text classification. The fastText model and the XGBoost model use the same method to optimize parameters, further improving the accuracy of classification.
[0048] 2. The fastText word embedding model captures short text information by constructing word vectors and N-gram features. The word embedding matrix obtained by inputting the text training set into the fastText word embedding model is used as the training set of the XGBoost model, so that the method of the present invention can quickly process short texts and is suitable for the classification processing of a large number of elevator fault description texts.
[0049] 3. The present invention performs data enhancement on the text training set through the fine-tuned MOSS model, which not only facilitates the expansion of the rare categories of elevator fault description texts, but also effectively increases the data volume of the text training set. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 1 is a flow chart of a method for classifying elevator fault description text according to an embodiment of the present invention;
[0051] Figure 2 is a schematic diagram of the process of S1 in one embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the process of S2 in one embodiment of the present invention;
[0053] Figure 4 is a schematic diagram of the process of S3 in one embodiment of the present invention;
[0054] Figure 5 It is a schematic diagram of the process of S6 in one embodiment of the present invention. DETAILED DESCRIPTION
[0055] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0056] In the description of the present invention, unless otherwise specified, “several” means two or more.
[0057] Reference Figure 1 , a method for classifying elevator fault description texts of the present invention comprises the following steps:
[0058] S1: Obtain the original text set of elevator fault description text, expand and preprocess the original text set to obtain a text training set;
[0059] S2: Input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set, input the word embedding matrices into the XGBoost model (extreme gradient boosting model) for training, and obtain the target XGBoost model;
[0060] S3: inputting the text training set into the fastText model for training to obtain a target fastText model;
[0061] S4: inputting the elevator fault description text to be classified into the fastText word embedding model for encoding processing, obtaining the word embedding matrix of the elevator fault description text to be classified, and inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result;
[0062] S5: inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result;
[0063] S6: Obtain a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result.
[0064] The present invention uses a deep learning method to train a dual classification model by expanding and preprocessing the text training set, and can efficiently and accurately realize the automatic classification of batches of elevator fault description texts, thereby improving the accuracy and efficiency of elevator fault information management. It should be noted that in the above steps S1 to S6, the order of S2 and S3 can be swapped, and the order of S4 and S5 can be swapped.
[0065] In the present invention, the elevator fault description texts to be classified in S4 and S5 come from maintenance personnel of various maintenance companies, so that the number of elevator fault description texts to be classified is relatively large. The method of the present invention is used to process the collected elevator fault description texts to be classified, and the classification results can be obtained quickly and accurately, which is beneficial to the elevator supervision-related departments to efficiently manage and analyze elevator fault information.
[0066] Reference Figure 2 In one embodiment of the present invention, S1 includes the following sub-steps:
[0067] S11: obtaining an original text set of elevator fault description texts, training a MOSS model with the original text set, and obtaining a fine-tuned MOSS model and elevator fault description texts;
[0068] S12: obtaining the rare categories that need to be expanded according to the elevator fault description text in S11, inputting the rare category supplementary set into the fine-tuned MOSS model to obtain supplementary text, and expanding the elevator fault description text with the supplementary text;
[0069] S13: The expanded elevator fault description text includes several categories of text sequences, each of which contains several words. The text sequence of each category is preprocessed to obtain an elevator fault description vocabulary list. The several categories of elevator fault description vocabulary lists are the text training sets.
[0070] The execution subject of the elevator fault description text classification method of the present invention is a classification device of the elevator fault description text classification method, and the classification device can be a computer device, a server, or a server group formed by a combination of multiple computer devices. The original text set of the elevator fault description text is obtained through the classification device.
[0071] Specifically, the classification device also collects basic knowledge of elevators, knowledge of the entire elevator, and knowledge of its components. In S11, the elevator fault description text in the original text set is first organized into a question-and-answer format based on the basic knowledge of elevators, knowledge of the entire elevator, and knowledge of its components, and packaged in a JSON format string. It is then input into the MOSS model deployed by the classification device for training to obtain a fine-tuned MOSS model and elevator fault description text. It should be noted that when organized into a question-and-answer format, one answer should have three or more question description methods, which can generate a large amount of sample data after being input into the MOSS model. When training the MOSS model, only one sample is placed in each conversation, and only one conversation is conducted. This approach can reduce the amount of data in each training batch, effectively reduce computing resources and time consumption, and can make the MOSS model more focused on specific areas and topics.
[0072] The fine-tuned MOSS model can generate text that is similar in style to the elevator fault description text and is natural and fluent, and can expand the amount of data. Based on the categories contained in the elevator fault description text, in S12, the rare categories that need to be expanded are obtained according to the elevator fault description text in S11, and the rare category supplement set is input into the fine-tuned MOSS model, which can expand the elevator fault description text again and realize data enhancement of the elevator fault description text.
[0073] Specifically, the elevator fault description text is divided into fault text content and fault labels to obtain a text sequence of several categories, and rare categories that need to be expanded are obtained according to the text sequences of several categories, wherein the categories include safety protection system problems, guide system problems, traction system problems, machine room, shaft, pit space problems, user reasons, electric traction system problems, elevator hygiene reasons, elevator additional device problems, electrical control system problems, car system problems, weight balance system problems, door system problems and other problems, reflecting the specific problems of the more classic eight systems and four space classifications in the elevator industry. The present invention uses the elevator fault description text as the training data of the classification model, which can make the trained classification model better adapt to the current elevator industry and improve the efficiency and accuracy of the elevator fault description text classification.
[0074] Preferably, in said S13, the preprocessing of the text sequence of each category includes word segmentation and stop word filtering. Specifically, the classification device uses the jieba word segmentation tool to segment the text sequences of several categories in the elevator fault description text according to the preset elevator common noun term segmentation word library, which includes the common noun terms of the entire elevator and elevator parts, and filters out symbols, pure numbers, and words that are useless for text classification through regularization to achieve stop word filtering, obtain the preprocessed text sequences of several categories, and construct the elevator fault description vocabulary list.
[0075] Reference Figure 3 In one embodiment of the present invention, in S2, the text training set is input into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set, including the following steps:
[0076] S21: Divide the text training set into a test set, use the skipgram algorithm, and train the fastText word embedding model with the elevator fault description vocabulary list in the test set. The loss function used is: ;
[0077] In the formula, L is the loss value, T is the number of words in the text sequence, t is the word index of the text sequence, c is the preset sliding window size, j is the word index of the sliding window, is the likelihood function, The first t+j words, The first t words;
[0078] S22: Input the text training set into the trained fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set.
[0079] The skipgram algorithm is used to perform unsupervised learning training on the fastText word embedding model. After the training is completed, the trained fasttext word embedding model and its parameters are saved and packaged. Several word embedding matrices of the text training set obtained in S22 are used as training sets to train the XGBoost model.
[0080] Specifically, the fastText word embedding model is combined with the N-gram architecture. The fastText word embedding model captures short text information by constructing word vectors and N-gram features. At the same time, based on the simple structure of the fastText word embedding model and the XGBoost model, the present invention can quickly process short texts and is suitable for the classification processing of a large number of elevator fault description texts.
[0081] Preferably, in S2, inputting the word embedding matrix into an XGBoost model for training to obtain a target XGBoost model comprises the following steps:
[0082] S23: extracting a number of words from the word embedding matrix by a random method with replacement, and obtaining classification prediction values of the words according to the words and a preset prediction algorithm, wherein the classification prediction values include classification prediction probability distribution values of a number of categories, and the classification model includes a number of decision trees;
[0083] The prediction algorithm is:
[0084] In the formula, i is the index of the word, For the i The classification prediction value of words, is the total number of decision trees, is the index of the decision tree, Indicates k A decision tree, Indicates i words, Indicates i Words through k The result predicted by a decision tree.
[0085] In this embodiment, the classification device adopts a random replacement method to extract a number of words from the elevator fault description vocabulary list, and obtains classification prediction values of the several words based on the several words and a preset prediction algorithm, wherein the classification prediction values include classification prediction probability distribution values of several categories.
[0086] S24: According to the classification prediction values of the plurality of words, the classification true probability distribution values and a preset gradient descent algorithm, the first classification model is trained to obtain a target first classification model, wherein the gradient descent algorithm is: ;
[0087] In the formula, is the loss value, are model parameters, n is the total number of words, is the log-likelihood loss function, For the i The true classification value of words, is the regularization function, For the i The classification prediction probability distribution value of each word.
[0088] In this embodiment, the first classification model is trained according to the classification prediction values of several words, the classification true probability distribution values and the preset gradient descent algorithm to obtain the target first classification model, so as to solve the gradient vanishing and gradient exploding problems in the long time series training process, and have a stronger fitting ability for the time-dependent data set, so that the classification model can perform accurate classification, and effectively improve the classification accuracy and precision of the elevator fault description text.
[0089] In one embodiment of the present invention, S2 further includes S25: using the RandomizedSearchCV (hyperparameter optimization) method to optimize the parameters of the trained XGBoost model, and using the optimized model as the target XGBoost model. Thus, it is used to improve the classification accuracy and precision of the elevator fault description text, wherein the parameters of the XGBoost model include the number parameter of weak classifiers, the learning rate, the gamma parameter, the maximum depth parameter of the decision tree, and the minimum number of samples on the leaves of the decision tree.
[0090] Reference Figure 4 In one embodiment of the present invention, the text training set is input into the fastText model for training to obtain a target fastText model, comprising the following steps:
[0091] S31: Using a supervised learning mode, the text training set is input into the fastText model for training;
[0092] S32: using the RandomizedSearchCV method to optimize the parameters of the trained fastText model, and using the optimized model as the target fastText model.
[0093] The fastText model in this embodiment uses the same method as the XGBoost model to perform parameter optimization to improve the f1-score (score index), thereby improving the accuracy of classification.
[0094] Furthermore, in S6, the first classification result and the second classification result both include classification prediction probability distribution values of several categories;
[0095] Obtaining a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result includes the following steps:
[0096] S61: In the first classification result and the second classification result, the classification prediction probability distribution values of the same category are accumulated to obtain classification prediction probability distribution accumulated values of several categories;
[0097] S62: Compare the classification prediction probability distribution cumulative values of several categories, and take the category corresponding to the largest classification prediction probability distribution cumulative value as the final classification result of the elevator fault description text to be classified.
[0098] The present invention can combine the advantages of parallel processing, strong interpretability and result stability of the XGBoost model with the advantages of fast training speed and high accuracy of the fastText model, and adopt a dual classification method to make the final classification result stable and highly accurate, thereby improving the accuracy and efficiency of elevator fault description text classification.
[0099] Accordingly, an embodiment of the present invention further provides an elevator fault description text classification system, wherein the elevator fault description text classification system is used to complete the above-mentioned elevator fault description text classification method;
[0100] The elevator fault description text classification system comprises:
[0101] A text preprocessing module is used to expand and preprocess the original text set to obtain a text training set; and to obtain elevator fault description text to be classified;
[0102] An encoding module is used to input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set;
[0103] A first classification module, used for inputting the word embedding matrix into an XGBoost model for training to obtain a target XGBoost model, and used for inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result;
[0104] A second classification module, used for training a fastText model using the text training set to obtain a target fastText model, and for inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result;
[0105] The classification result output module is used to obtain the final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result, and output the final classification result.
[0106] In this embodiment, the MOSS model, the fastText word embedding model, the XGBoost model and the fastText model are trained through the text preprocessing module, the encoding module, the first classification module and the second classification module to obtain a model for efficiently and accurately realizing the automatic classification of the elevator fault description text; the text preprocessing module, the encoding module, the first classification module, the second classification module and the classification result output module are used to realize the rapid and accurate processing of a large number of elevator fault description texts to be classified.
[0107] Accordingly, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned elevator fault description text classification method are implemented. The computer device can store multiple instructions, and the instructions are suitable for the processor to load and execute the above-mentioned method steps S1 to S6. The specific execution process can refer to the specific description of the above-mentioned implementation method, which will not be repeated here.
[0108] Other structures and operations of an elevator fault description text classification method, system, and computer-readable storage medium according to an embodiment of the present invention are well known to those skilled in the art and will not be described in detail here.
[0109] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0110] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for classifying elevator fault description text, characterized in that: The following steps are involved: S1: Obtain the original text set of elevator fault description text, expand and preprocess the original text set to obtain a text training set; S2: Input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set, input the word embedding matrices into the XGBoost model for training, and obtain the target XGBoost model; S3: inputting the text training set into the fastText model for training to obtain a target fastText model; S4: inputting the elevator fault description text to be classified into the fastText word embedding model for encoding processing, obtaining the word embedding matrix of the elevator fault description text to be classified, and inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result; S5: inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result; S6: Obtaining a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result; The S1 comprises the following sub-steps: S11: obtaining an original text set of elevator fault description texts, training a MOSS model with the original text set, and obtaining a fine-tuned MOSS model and elevator fault description texts; organizing the elevator fault description texts in the original text set into a question-and-answer format based on basic elevator knowledge, overall elevator knowledge, and component knowledge. One answer should have three or more question description methods, and a large amount of sample data can be generated after being input into the MOSS model; S12: obtaining rare categories that need to be expanded according to the classification in the original text set, inputting the rare categories into the fine-tuned MOSS model to obtain supplementary texts, and expanding the elevator fault description text with the supplementary texts; S13: the expanded elevator fault description text includes several categories of text sequences, the text sequences include several words, each category of text sequence is preprocessed to obtain an elevator fault description vocabulary list, and the several categories of elevator fault description vocabulary lists are text training sets; The S2 comprises the following steps: S21: Divide the text training set into a test set, use the skipgram algorithm, and train the fastText word embedding model with the elevator fault description vocabulary list in the test set. The loss function used is: ; In the formula, L is the loss value, T is the number of words in the text sequence, t is the word index of the text sequence, c is the preset sliding window size, j is the word index of the sliding window, is the likelihood function, The first t+j words, The first t words; S22: Input the text training set into the trained fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set; S23: extracting a number of words from the word embedding matrix by a random method with replacement, and obtaining classification prediction values of the words according to the words and a preset prediction algorithm, wherein the classification prediction values include classification prediction probability distribution values of several categories, and the XGBoost model includes several decision trees; The prediction algorithm is: ; In the formula, i is the index of the word, For the i The classification prediction value of words, K is the total number of decision trees, k is the index of the decision tree, Indicates k A decision tree, Indicates i words, Indicates i Words through k The result predicted by the decision tree; S24: According to the classification prediction values of the plurality of words, the classification true probability distribution values and the preset gradient descent algorithm, the XGBoost model is trained to obtain a target XGBoost model, wherein the gradient descent algorithm is: , , In the formula, is the loss value, Model parameters, n is the total number of words, is the log-likelihood loss function, For the i The true classification value of words, is the regularization function, For the i The classification prediction probability distribution value of words; In S6, the first classification result and the second classification result both include classification prediction probability distribution values of several categories; Obtaining a final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result includes the following steps: S61: In the first classification result and the second classification result, the classification prediction probability distribution values of the same category are accumulated to obtain classification prediction probability distribution accumulated values of several categories; S62: comparing the classification prediction probability distribution cumulative values of several categories, and taking the category corresponding to the largest classification prediction probability distribution cumulative value as the final classification result of the elevator fault description text to be classified; In S13, the text sequence of each category is preprocessed, including word segmentation and stop word filtering; The S2 also includes S25: using the RandomizedSearchCV method to optimize the parameters of the trained XGBoost model, and using the optimized model as the target XGBoost model; Inputting the text training set into the fastText model for training to obtain a target fastText model comprises the following steps: S31: Using a supervised learning mode, the text training set is input into the fastText model for training; S32: using the RandomizedSearchCV method to optimize the parameters of the trained fastText model, and using the optimized model as the target fastText model.
2. An elevator fault description text classification system, characterized in that: The elevator fault description text classification system is used to complete the elevator fault description text classification method according to claim 1; The elevator fault description text classification system comprises: A text preprocessing module is used to expand and preprocess the original text set to obtain a text training set; and to obtain elevator fault description text to be classified; An encoding module is used to input the text training set into the fastText word embedding model for encoding processing to obtain several word embedding matrices of the text training set; A first classification module, used for inputting the word embedding matrix into an XGBoost model for training to obtain a target XGBoost model, and used for inputting the word embedding matrix of the elevator fault description text to be classified into the target XGBoost model to obtain a first classification result; A second classification module, used for training a fastText model using the text training set to obtain a target fastText model, and for inputting the elevator fault description text to be classified into the target fastText model to obtain a second classification result; The classification result output module is used to obtain the final classification result of the elevator fault description text to be classified according to the first classification result and the second classification result, and output the final classification result.
3. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the elevator fault description text classification method according to claim 1 are implemented.
Citation Information
Patent Citations
A method and system based on elevator customer service text classification
CN109947941A
Text classification method and device, storage medium and computer equipment
CN110209805A
XGBoost integrated credit evaluation system combined with deep neural network and method thereof
CN110472817A
Methods and apparatuses for troubleshooting a computer system
US20240045752A1