Data enhancement method of operation and maintenance data, text classification method and electronic equipment

By adopting multi-grained data augmentation method and text feature extraction technology in IT operation and maintenance data, the problems of data imbalance and sparse characteristics in the text classification of IT operation and maintenance data are solved, and the performance of the text classification model is improved.

CN120179986APending Publication Date: 2025-06-20E SURFING VISION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311698464.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-20

Smart Images

  • Figure CN120179986A_ABST
    Figure CN120179986A_ABST
Patent Text Reader

Abstract

The invention provides a data enhancement method for operation and maintenance data, a text classification method and electronic equipment. The enhancement method comprises the following steps: acquiring a plurality of pieces of operation and maintenance data as a text data set, and performing data enhancement on the text data set; performing data preprocessing on the text data set after data enhancement; converting the preprocessed text data set into text vector representation to obtain a text vector data set, and performing feature extraction on the text vector data set to obtain a text enhanced data set; the classification method comprises the steps that to-be-classified operation and maintenance data are obtained, text classification is conducted on the to-be-classified operation and maintenance data through a text classification model, and the text classification model is obtained through training of a text enhancement data set. According to the method, the problems of unbalanced sample categories, sparse short text features and the like in the data set can be relieved, redundant features which have small influence on text classification are effectively eliminated, the quality of the data set is improved, and the generalization ability and classification performance of a text classification model can be improved. The method is applied to the technical field of natural language processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a data enhancement method, a text classification method and an electronic device for operation and maintenance data. Background Art

[0002] IT (Information Technology) operation and maintenance data refers to the text data generated during the IT operation and maintenance process, mainly in the forms of work tickets, logs, etc. Its content consists of two text structures: structured text and unstructured text. Currently, deep learning models such as TextCNN, TextRNN, and Transformer are usually used to perform text classification processing on IT operation and maintenance data.

[0003] However, the unstructured text in IT operation and maintenance data is usually manually recorded, which contains a large number of IT operation and maintenance domain vocabulary, and there is a large gap in the text length between different unstructured texts. This will lead to frequent problems such as colloquialism and splicing grammar errors. Moreover, the labeled data set used for text classification usually has a relatively serious problem of data category imbalance, which undoubtedly causes a certain degree of difficulty in extracting effective text feature information, and further makes the classification effect of existing deep learning models in the text classification task of IT operation and maintenance data unsatisfactory.

[0004] To address the problem of data category imbalance in the labeled data set, some related technologies adopt data enhancement techniques to alleviate this problem. Currently, most commonly used text data enhancement techniques are text data enhancement techniques based on adding noise or back translation. However, these common text data enhancement techniques do not retain the important feature words that contribute greatly to text classification in the original text data when implementing data enhancement, and they have problems such as limited enhancement effect on short texts with sparse features and low diversification degree of the added text corpus. The data enhancement effect still needs to be improved. Summary of the Invention

[0005] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0006] To this end, an object of the present invention is to provide a data enhancement method, a text classification method and an electronic device for operation and maintenance data.

[0007] To achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:

[0008] On the one hand, an embodiment of the present invention provides a data enhancement method for operation and maintenance data, including the following steps:

[0009] Obtain a plurality of IT operation and maintenance data as a text data set;

[0010] Perform data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set;

[0011] Perform data preprocessing on the augmented text data set to obtain a preprocessed text data set;

[0012] Convert the preprocessed text data set into a text vector representation to obtain a text vector data set;

[0013] Extract features from the text vector data set to obtain an augmented text data set.

[0014] In addition, according to the data augmentation method for operation and maintenance data in the above embodiments of the present invention, the following additional technical features may also be included:

[0015] Further, in an embodiment of the present invention, the obtaining of several IT operation and maintenance data as a text data set includes:

[0016] Obtain several IT operation and maintenance data and their corresponding text categories as a text data set.

[0017] Further, in an embodiment of the present invention, the performing of data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set includes:

[0018] Obtain the IT operation and maintenance data belonging to the few-shot categories in the text data set as text sentences to be augmented, and each text sentence to be augmented is composed of several text words;

[0019] Randomly select multiple text words from multiple text sentences to be augmented, and perform data augmentation on the multiple text sentences to be augmented according to the multiple randomly selected text words to generate multiple first amplified text sentences;

[0020] Obtain the keyword table of the text data set, and perform data augmentation on the multiple text sentences to be augmented according to the keyword table to generate multiple second amplified text sentences;

[0021] Perform data augmentation on the multiple text sentences to be augmented by back-translation to generate multiple third amplified text sentences;

[0022] Add the multiple first amplified text sentences, second amplified text sentences, and third amplified text sentences to the text data set to generate an augmented text data set.

[0023] Further, in an embodiment of the present invention, randomly selecting multiple text words from multiple text sentences to be enhanced, and performing data enhancement on the multiple text sentences to be enhanced according to the multiple randomly selected text words to generate multiple first amplified text sentences, including:

[0024] Randomly select multiple text words from multiple text sentences to be enhanced as the first sample words;

[0025] Perform data enhancement on the multiple first sample words through the edit distance replacement method and the keyboard mis-touch replacement method to generate candidate words for the multiple first sample words;

[0026] In each text sentence to be enhanced, replace the first sample word with the candidate word of the first sample word to generate a first amplified text sentence corresponding to each text sentence to be enhanced, and then obtain multiple first amplified text sentences.

[0027] Further, in an embodiment of the present invention, performing data enhancement on the multiple first sample words through the edit distance replacement method and the keyboard mis-touch replacement method to generate candidate words for the multiple first sample words, including:

[0028] Obtain an IT operation and maintenance word database, where the IT operation and maintenance word database includes multiple IT operation and maintenance words, and then obtain the edit distance between each first sample word and each IT operation and maintenance word;

[0029] For each first sample word, obtain the IT operation and maintenance words with an edit distance less than a first value as the first candidate words of the first sample word;

[0030] Randomly select one letter from each first sample word as the letter to be replaced, and determine the candidate letters for the letter to be replaced in each first sample word;

[0031] Replace the letter to be replaced in each first sample word with the candidate letter for the letter to be replaced to generate a second candidate word for each first sample word;

[0032] Obtain the first candidate words and the second candidate words of the multiple first sample words as the candidate words of the multiple first sample words.

[0033] Further, in an embodiment of the present invention, obtaining the edit distance between each first sample word and each IT operation and maintenance word includes:

[0034] Obtain the number of character editing operations for converting the first sample word into an IT operation and maintenance word, where the number of character editing operations is the minimum value of the number of executions of the character editing operation;

[0035] Calculate the edit distance between each first sample word and each IT operation and maintenance word according to the character editing operand.

[0036] Further, in an embodiment of the present invention, the character editing operation includes at least one of inserting a character, deleting a character, and changing a character.

[0037] Further, in an embodiment of the present invention, the determining the candidate letters of the letter to be replaced in each first sample word includes:

[0038] Obtain the position information of the letter to be replaced on the keyboard;

[0039] According to the position information of the letter to be replaced on the keyboard, select multiple letters surrounding the letter to be replaced on the keyboard as the candidate letters of the letter to be replaced.

[0040] Further, in an embodiment of the present invention, the obtaining the keyword table of the text data set includes:

[0041] Process the text data set through the TextRank algorithm to generate the keyword table of the text data set.

[0042] Further, in an embodiment of the present invention, the data augmentation of multiple text statements to be enhanced according to the keyword table to generate multiple second amplified text statements includes at least one of the following:

[0043] For each text statement to be enhanced, randomly select a text word belonging to the keyword table from the text statement to be enhanced, replace the second sample word with a synonym of the second sample word, and generate a second amplified text statement corresponding to the text statement to be enhanced;

[0044] For each text statement to be enhanced, traverse the text words belonging to the keyword table in the text statement to be enhanced, and randomly insert synonyms of the text words belonging to the keyword table into the text statement to be enhanced to generate a second amplified text statement corresponding to the text statement to be enhanced;

[0045] For each text statement to be enhanced, traverse the text words in the text statement to be enhanced, and exchange the positions of two text words in the text statement to be enhanced to generate a second amplified text statement corresponding to the text statement to be enhanced;

[0046] Obtain the word deletion probability, and for each text statement to be enhanced, delete the text words not belonging to the keyword table from the text statement to be enhanced according to the word deletion probability to generate a second amplified text statement corresponding to the text statement to be enhanced.

[0047] Further, in an embodiment of the present invention, the data augmentation of multiple text statements to be enhanced by the back-translation method to generate multiple third augmented text statements includes:

[0048] Determine the initial language and the translated language of each text statement to be enhanced;

[0049] Translate each text statement to be enhanced represented by the initial language into the translated language to obtain each text statement to be enhanced represented by the translated language;

[0050] Translate each text statement to be enhanced represented by the translated language into the initial language to obtain the third augmented text statement corresponding to each text statement to be enhanced, and further generate multiple third augmented text statements.

[0051] Further, in an embodiment of the present invention, the initial language is English, and the translated language includes at least one of Chinese and French.

[0052] Further, in an embodiment of the present invention, the data preprocessing of the text data set after data augmentation to obtain the preprocessed text data set includes:

[0053] Perform text standardization and text cleaning on the text data set after data augmentation to obtain the text data set after text cleaning;

[0054] Perform string replacement processing on the text data set after text cleaning to obtain the text data set after replacement processing;

[0055] Perform word segmentation on the text data set after replacement processing to obtain the word-level tokens and character-level tokens of the text data set;

[0056] Perform word cleaning and normalization processing on the word-level tokens of the text data set to obtain the preprocessed text data set.

[0057] Further, in an embodiment of the present invention, the performing text standardization and text cleaning on the text data set after data augmentation to obtain the text data set after text cleaning includes:

[0058] Uniformly convert the characters in the text data set after data augmentation to half-width, and uniformly convert the English letters in the text data set after data augmentation to lowercase;

[0059] Eliminate the invalid characters in the text data set after data augmentation.

[0060] Further, in an embodiment of the present invention, the performing string replacement processing on the text data set after text cleaning to obtain the text data set after replacement processing includes:

[0061] Identify the rules for the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset;

[0062] Obtain the identifiers of the IT operation and maintenance feature words, and replace the IT operation and maintenance feature words in the text dataset with the identifiers of the IT operation and maintenance feature words, thereby obtaining the text dataset after replacement processing.

[0063] Further, in an embodiment of the present invention, the IT operation and maintenance feature words include at least one of date, time, percentage, IP address, memory address, memory capacity information, website address, telephone number, network status code, and error code.

[0064] Further, in an embodiment of the present invention, the identifying the rules for the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset includes:

[0065] Use regular expressions to identify the rules for the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset.

[0066] Further, in an embodiment of the present invention, the tokenizing the text dataset after replacement processing to obtain the word-level tokens and character-level tokens of the text dataset includes:

[0067] Use the NLTK natural language processing tool to tokenize the text dataset after replacement processing to obtain the word-level tokens and character-level tokens of the text dataset.

[0068] Further, in an embodiment of the present invention, the cleaning and normalizing the word-level tokens of the text dataset to obtain the preprocessed text dataset includes:

[0069] Clean the word-level tokens of the text dataset and restore the part of speech and tense of the word-level tokens of the text dataset to obtain the preprocessed text dataset.

[0070] Further, in an embodiment of the present invention, the methods of the word cleaning process include at least one of stop word removal, part-of-speech tagging, lemmatization, and splicing error correction.

[0071] Further, in an embodiment of the present invention, the converting the preprocessed text dataset into a text vector representation to obtain a text vector dataset includes:

[0072] Encode the character-level tokens of the preprocessed text dataset to generate the character vector features of the text dataset;

[0073] Encode the word-level tokens of the preprocessed text dataset to generate the word vector features of the text dataset;

[0074] Construct a text vector dataset based on the character vector features and word vector features of the text dataset.

[0075] Further, in an embodiment of the present invention, the encoding of the character-level tokens of the preprocessed text dataset to generate the character vector features of the text dataset includes:

[0076] Encode the character-level tokens of the preprocessed text dataset using one-hot encoding to generate the character vector features of the text dataset.

[0077] Further, in an embodiment of the present invention, the encoding of the word-level tokens of the preprocessed text dataset to generate the word vector features of the text dataset includes:

[0078] Encode the word-level tokens of the preprocessed text dataset using a pre-trained word vector model to generate the word vector features of the text dataset.

[0079] Further, in an embodiment of the present invention, the word vector model is a Word2Vec model.

[0080] Further, in an embodiment of the present invention, the feature extraction of the text vector dataset to obtain a text enhanced dataset includes:

[0081] Perform concatenation processing and text length normalization processing on the character vector features and word vector features in the text vector dataset to obtain a text enhanced dataset.

[0082] Further, in an embodiment of the present invention, the concatenation processing and text length normalization processing of the character vector features and word vector features in the text vector dataset include:

[0083] Concatenate the character vector features and word vector features in the text vector dataset in series to obtain the concatenated vector features of the text vector dataset;

[0084] Unify the text length of the concatenated vector features of the text vector dataset to a second value.

[0085] On the other hand, an embodiment of the present invention provides a text classification method for operation and maintenance data, including the following steps:

[0086] Obtain the IT operation and maintenance data to be classified;

[0087] Input the IT operation and maintenance data to be classified into the trained text classification model for text classification to obtain the text classification result;

[0088] Among them, the text classification model is obtained by pre-training the initial classification model using a text enhancement dataset, and the text enhancement dataset is obtained by the data enhancement method of the above-mentioned operation and maintenance data.

[0089] Further, in an embodiment of the present invention, the initial classification model is a TextRCNN model.

[0090] On the other hand, an embodiment of the present invention provides an electronic device, including:

[0091] At least one processor;

[0092] At least one memory for storing at least one program;

[0093] When the at least one program is executed by the at least one processor, the at least one processor implements the data enhancement method of the above-mentioned operation and maintenance data or the text classification method of the above-mentioned operation and maintenance data.

[0094] The beneficial effects of the present invention are as follows: A data augmentation method, a text classification method, and an electronic device for operation and maintenance data are provided. The data augmentation method includes: First, obtaining a number of IT operation and maintenance data as a text data set, then performing data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set, then performing data preprocessing on the augmented text data set to obtain a preprocessed text data set, and converting the preprocessed text data set into a text vector representation to obtain a text vector data set. Finally, feature extraction is performed on the text vector data set to obtain a text augmented data set, thereby realizing the text augmentation process of IT operation and maintenance data. The text classification method includes: Obtaining the IT operation and maintenance data to be classified, and using a text classification model to perform text classification on the IT operation and maintenance data to be classified. The text classification model is obtained by training an initial classification model with the text augmented data set, and the text augmented data set is obtained by the data augmentation method described above. The present invention performs data augmentation on IT operation and maintenance data at three different granularities of the character level, word level, and sentence level, increases the diversity of sentence patterns for the original data set, can effectively alleviate the problem of sample class imbalance in the original data set and the problem of feature sparsity of short text samples, eliminates redundant features in the original data set that have little impact on text classification, realizes multi-dimensional text feature extraction, and is conducive to generating a high-quality data set. In addition, using the data set generated by the data augmentation method of the present invention to train the text classification model can improve the generalization ability of the text classification model, and thus improve the classification accuracy of the text classification model in the operation and maintenance data text classification task.

[0095] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or can be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings

[0096] Figure 1 is a flowchart of a data augmentation method for operation and maintenance data provided by the present invention;

[0097] Figure 2 is a flowchart of data augmentation provided by the present invention;

[0098] Figure 3 is a flowchart of character-level data augmentation provided by the present invention;

[0099] Figure 4 is a schematic diagram of a 26-key keyboard provided by the present invention;

[0100] Figure 5 is a flowchart of sentence-level data augmentation provided by the present invention;

[0101] Figure 6 It is a schematic diagram of the application of data augmentation provided by the present invention;

[0102] Figure 7 It is a flowchart of the preprocessing of the dataset provided by the present invention;

[0103] Figure 8 It is a schematic diagram of the identifier provided by the present invention;

[0104] Figure 9 It is a schematic diagram of the principle of the data augmentation method for operation and maintenance data provided by the present invention;

[0105] Figure 10 It is a flowchart of the text classification method for operation and maintenance data provided by the present invention. Detailed implementation manners

[0106] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0107] The present application will be further described below with reference to the accompanying drawings of the specification and specific embodiments. The described embodiments should not be regarded as a limitation of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0108] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0109] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0110] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0111] 1) Text classification (TC), also known as automatic text classification, refers to the process by which a computer maps a piece of text containing information to a given category or categories of topics. The algorithm model that implements this process is called a classifier. The text classification problem is a classic problem in the field of natural language processing (NLP).

[0112] 2) Data Augmentation (DA) refers to a method of generating a large amount of labeled data from a small amount of labeled data while ensuring that the semantics of the labels remain unchanged as much as possible. Data augmentation can increase the number of training samples and enhance the generalization ability and robustness of the model.

[0113] IT (Information Technology) operation and maintenance data refers to text data generated during the IT operation and maintenance process, mainly in the form of work tickets, logs, etc. Its content consists of two text structures: structured text and unstructured text. Currently, deep learning models such as TextCNN and TextRNN are usually used to perform text classification on IT operation and maintenance data.

[0114] However, the unstructured text in IT operation and maintenance data is usually recorded manually, which contains a large number of IT operation and maintenance vocabulary, and the text length difference between different unstructured texts is large, which will lead to the frequent occurrence of problems such as colloquialism and splicing grammatical errors. In addition, the annotated data sets used for text classification usually have more serious data category imbalance problems, which undoubtedly creates a certain degree of difficulty in extracting effective text feature information, and thus makes the existing deep learning models have poor classification effect on the text classification task of IT operation and maintenance data.

[0115] Regarding the problem of data category imbalance in annotated datasets, some related technologies use data enhancement technology to alleviate this problem. Currently, most commonly used text data enhancement technologies are based on adding noise or back translation. On the one hand, related text data enhancement technologies lack the capture and utilization of feature words in the IT operation and maintenance field, which leads to the fact that they do not retain important feature words in the original text data that contribute greatly to text classification when performing data enhancement, resulting in feature redundancy. On the other hand, related text enhancement technologies have limited enhancement effects on short texts with sparse features, and there are problems such as low diversity of the increased text corpus, and the data enhancement effect still needs to be improved.

[0116] In view of the problems existing in the related technologies, such as the imbalance of dataset categories, the sparsity of short text features, the lack of capture and utilization of feature words in the IT operation and maintenance field during the data augmentation process, and the low accuracy of the classification model, the present invention provides a data augmentation method, a text classification method and an electronic device for operation and maintenance data. The IT operation and maintenance data is augmented at three different granularities: character level, word level and sentence level, and the text classification model is trained using the dataset generated by the data augmentation. The trained text classification model is used to implement the text classification of the IT operation and maintenance data. While increasing the data volume of the dataset, the present invention can effectively alleviate the problems of sample category imbalance and feature sparsity of short text samples in the original dataset, which helps to improve the generalization ability of the text classification model and the classification accuracy of the text classification model in the text classification task of operation and maintenance data.

[0117] The following further elaborates on the embodiments of the present application with reference to the accompanying drawings.

[0118] First, an implementation step of a data augmentation method for operation and maintenance data according to an embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0119] The method in the embodiments of the present invention can be applied to a terminal, or to a server, or can also be software running on a terminal or a server, etc. The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited thereto. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.

[0120] Refer to Figure 1 , Figure 1 which is a flowchart of a data augmentation method for operation and maintenance data provided by the present invention. The data augmentation method may include but is not limited to the following steps:

[0121] S101, obtain a plurality of IT operation and maintenance data as a text dataset.

[0122] It should be noted that the IT operation and maintenance data is mainly represented in the form of work tickets. Those skilled in the art can understand that other data forms such as logs are also applicable, and only corresponding adjustments need to be made.

[0123] S102. Perform data augmentation on the text dataset at the character level, word level, and sentence level to generate an augmented text dataset.

[0124] In this step, the data of the few-shot classes in the text dataset with class imbalance is respectively augmented at the character level, word level, and sentence level, and the newly added sample data is added to the text dataset, so as to obtain a text dataset with relatively sufficient and balanced data samples.

[0125] S103. Perform data preprocessing on the augmented text dataset to obtain a preprocessed text dataset.

[0126] Optionally, the methods of data preprocessing may include but are not limited to text normalization, data cleaning, word segmentation, text identifier replacement, etc.

[0127] In this step, after obtaining the augmented text dataset, perform data preprocessing on the text dataset to obtain a preprocessed text dataset, so as to eliminate the redundant features in the text dataset and improve the expression ability of the key features of IT operation and maintenance.

[0128] S104. Convert the preprocessed text dataset into a text vector representation to obtain a text vector dataset.

[0129] It should be noted that the text vector representation can be a discrete representation, and the discrete representation can be obtained through algorithms such as the bag-of-words model, one-hot encoding, TF-IDF, etc. Those skilled in the art can understand that the text vector representation can also be a distributed representation, and the distributed representation can be obtained through algorithms such as Word2Vec, BERT, etc. The present invention does not make specific limitations on this.

[0130] In this step, the text vector representation refers to representing the semantics of the text through a numerical vector. After obtaining the preprocessed text dataset, convert the text dataset into the form of a vector or matrix to generate a text vector dataset, so that the sample data of the text vector dataset conforms to the input data format of the classification model.

[0131] S105. Extract features from the text vector dataset to obtain an enhanced text dataset.

[0132] In this step, by combining the text vectors at the character level and word level, multi-dimensional text feature extraction is realized, which can alleviate the problem of feature sparsity of short text samples.

[0133] In some embodiments of the present invention, in step S101, the step of obtaining several IT operation and maintenance data as the text dataset mainly includes:

[0134] S1011, Obtain a number of IT operation and maintenance data and their corresponding text categories as a text data set.

[0135] Optionally, in step S101, a number of IT operation and maintenance data and their corresponding text categories are obtained by monitoring the operating conditions of devices or application systems such as servers, network devices, databases, middleware, etc.

[0136] Among them, the network device can be a bridge, or can also be a repeater, switch, hub, router, gateway, etc. The database can be a relational database such as SQLite, MySQL, etc., or can also be a non-relational database such as MongoDb, HBase, etc. The middleware can be a message-oriented middleware such as RocketMQ, Kafka, etc., or can also be an object request broker middleware, remote procedure call middleware, etc. Those skilled in the art can understand that the operation and maintenance data of one or more devices or application systems can be obtained according to the actual situation as IT operation and maintenance data, and the present invention does not make specific limitations on this.

[0137] It should be noted that the text category is the label data of the IT operation and maintenance data, and the text category can be appropriately adjusted according to the actual text classification problem. Optionally, the text classification problem in the embodiments of the present invention can be a binary classification problem, or can also be other classification problems such as multi-classification, multi-label classification, etc. The present invention does not make specific limitations on the type of the text classification problem.

[0138] Exemplarily, when the text classification problem is a binary classification problem, the text category includes any one of text type or non-text type, or the text type includes any one of operation and maintenance data type or non-operation and maintenance data type; when the text classification problem is a multi-classification problem, the text category can be the specific text type of the data, such as the IT operation and maintenance data of middleware, the IT operation and maintenance data of servers, etc.

[0139] In some embodiments of the present invention, referring to Figure 2 , Figure 2 is the flowchart of data augmentation provided by the present invention. In step S102, the process of performing data augmentation on the text data set at the character level, word level, and sentence level to generate the data-augmented text data set mainly includes the following steps:

[0140] S1021, Obtain the IT operation and maintenance data belonging to the few-shot categories in the text data set as the text statements to be augmented.

[0141] It should be noted that each text statement to be augmented is composed of a number of text words.

[0142] In this step, obtain the IT operation and maintenance data belonging to the few-shot categories in the text dataset, and use the IT operation and maintenance data belonging to the few-shot categories as the data objects for data augmentation.

[0143] Optionally, the few-shot category can be a category in the text dataset with the number of data less than a preset value, or a category in the text dataset with relatively difficult data acquisition. The few-shot category can be set according to the actual situation, and the present invention does not make specific limitations thereto.

[0144] S1022: Randomly select multiple text words from multiple text sentences to be augmented, and perform data augmentation on the multiple text sentences to be augmented according to the multiple randomly selected text words to generate multiple first amplified text sentences.

[0145] In this step, at the character level, perform data augmentation on the multiple text sentences to be augmented based on the data augmentation method of text edit distance and adjacent keyboard input to obtain multiple first amplified text sentences, so as to improve the generalization ability of the classification model to spelling mistakes in words.

[0146] S1023: Obtain the keyword table of the text dataset, and perform data augmentation on the multiple text sentences to be augmented according to the keyword table to generate multiple second amplified text sentences.

[0147] It should be noted that the keyword table includes multiple keywords, and the keywords can be understood as text words with key features in the field of IT operation and maintenance.

[0148] In this step, at the word level, improve the traditional EDA (Easy Data Augmentation) method based on the TextRank algorithm, propose an improved text data augmentation method, obtain the keyword table of the text dataset based on the improved text data augmentation method, and perform data augmentation on the multiple text sentences to be augmented according to the keyword table to obtain multiple second amplified text sentences. Compared with the traditional EDA method, the improved text data augmentation method proposed in the embodiments of the present invention focuses more on the extraction of important feature words in the dataset.

[0149] S1024: Perform data augmentation on the multiple text sentences to be augmented through back-translation to generate multiple third amplified text sentences.

[0150] It should be noted that the back-translation method refers to the process of translating the original sentence into another language and then translating it back to the original language.

[0151] In this step, at the sentence level, data augmentation is performed on multiple text sentences to be enhanced through back translation to obtain multiple third augmented text sentences, so that the text sentences obtained after data augmentation have stronger diversity.

[0152] S1025, Add multiple first augmented text sentences, second augmented text sentences, and third augmented text sentences to the text dataset to generate an augmented text dataset.

[0153] In this step, the newly added sample data obtained through the foregoing steps is added to the text dataset, thereby obtaining a text dataset with a relatively sufficient and balanced number of data samples.

[0154] In some embodiments of the present invention, refer to Figure 3 , Figure 3 is a flowchart of character-level data augmentation provided by the present invention. In step S1022, multiple text words are randomly selected from multiple text sentences to be enhanced, and data augmentation is performed on the multiple text sentences to be enhanced according to the multiple randomly selected text words to generate multiple first augmented text sentences. The process mainly includes the following steps:

[0155] A1, Randomly select multiple text words from multiple text sentences to be enhanced as the first sample words.

[0156] It should be noted that the number of the first sample words can be determined according to the actual situation, and the present invention embodiment does not specifically limit the number of the first sample words.

[0157] A2, Perform data augmentation on the multiple first sample words through the edit distance replacement method and the keyboard mis-touch replacement method to generate candidate words for the multiple first sample words.

[0158] It should be noted that the edit distance replacement method refers to a method of generating a new word candidate set by calculating the edit distance (EditDistance, ED) between two words, and the keyboard mis-touch replacement method refers to a method of randomly replacing a certain letter in a word with any adjacent letter of the letter on the keyboard to generate a new word.

[0159] In this step, aiming at the problem of spelling errors of words composed of similar characters during data entry, the edit distance replacement method is used to perform data augmentation on the multiple first sample words, thereby generating the first candidate set. At the same time, aiming at the situation of spelling errors caused by keyboard mis-touch operations during data entry, the keyboard mis-touch replacement method is used to perform data augmentation on the multiple first sample words, thereby generating the second candidate set. The candidate words for the multiple first sample words are constituted by the first candidate set and the second candidate set.

[0160] A3. In each text statement to be enhanced, replace the first sample word with a candidate word of the first sample word to generate a first amplified text statement corresponding to each text statement to be enhanced, thereby obtaining multiple first amplified text statements.

[0161] In this step, by replacing the first sample word in each text statement to be enhanced with a candidate word of the first sample word, multiple first amplified text statements are generated, realizing data augmentation of the text dataset at the character level.

[0162] Exemplarily, in step A2, any text statement to be enhanced includes M text words. First, select N text words from the M text words as the first sample words, and perform data augmentation on the N first sample words through the edit distance replacement method to generate a first candidate set; at the same time, perform data augmentation on the N first sample words through the keyboard mis-touch replacement method to generate a second candidate set. Then, select N candidate words from the total candidate set composed of the first candidate set and the second candidate set to replace the N first sample words in the text statement to be enhanced, thereby generating a new text statement, that is, the first amplified text statement.

[0163] In some embodiments of the present invention, step A2 is mainly divided into two parallel processing parts. One part is to implement data augmentation by the edit distance replacement method, and the other part is to implement data augmentation by the keyboard mis-touch replacement method. In step A2, the implementation process of performing data augmentation on multiple first sample words through the edit distance replacement method and the keyboard mis-touch replacement method to generate candidate words of multiple first sample words mainly includes the following steps:

[0164] A21. Obtain an IT operation and maintenance word database.

[0165] It should be noted that the IT operation and maintenance word database is a pre-constructed database, and the IT operation and maintenance word database includes multiple IT operation and maintenance words. It can be understood that IT operation and maintenance words refer to the words that make up IT operation and maintenance data.

[0166] A22. Obtain the edit distance between each first sample word and each IT operation and maintenance word.

[0167] It should be noted that the edit distance is a measure of the similarity between two character sequences. The character sequence can be understood as a word, and the edit distance is inversely proportional to the similarity between two words. That is, the greater the edit distance, the smaller the similarity between two words; the smaller the edit distance, the greater the similarity between two words.

[0168] Specifically, the edit distance between two words can be the minimum number of character edit operations required to convert between the two words. Based on this, first, obtain the number of character edit operations for converting the first sample word into an IT operation and maintenance word. Then, calculate the edit distance between each first sample word and each IT operation and maintenance word according to the number of character edit operations.

[0169] Optionally, the number of character edit operations is the minimum value of the number of executions of character edit operations, and the character edit operations can include but are not limited to at least one of inserting a character, deleting a character, and changing a character.

[0170] Exemplarily, obtain the first sample word abc and the IT operation and maintenance words abcd, ab, and abd. When the character edit operation is inserting a character, by inserting the character d into the first sample word abc, the first sample word abc is converted into the IT operation and maintenance word abcd; when the character edit operation is deleting a character, by deleting the character c in the first sample word abc, the first sample word abc is converted into the IT operation and maintenance word ab; when the character edit operation is changing a character, by changing the character c in the first sample word abc to the character d, the first sample word abc is converted into the IT operation and maintenance word abd.

[0171] More specifically, the edit distance between any first sample word and any IT operation and maintenance word satisfies the following formula:

[0172]

[0173] where lev a,b (i, j) represents the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b; a i represents the i-th character of the first sample word a, and b j represents the j-th character of the IT operation and maintenance word b.

[0174] In the formula, when a i ≠ b j , take the minimum value among lev a,b (i - 1, j) + 1, lev a,b (i, j - 1) + 1, and lev a,b (i - 1, j - 1) + 1 as the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b; when lev a,b (i - 1, j) + 1, lev a,b (i, j - 1) + 1, and lev a,bWhen the minimum value in (i - 1, j - 1)+1 is 0, max(i, j) represents the boundary condition of the edit distance, and the maximum value of i and j is taken as the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b.

[0175] A23. For each first sample word, obtain the IT operation and maintenance words with an edit distance less than the first value as the first candidate words of the first sample word.

[0176] It should be noted that the value of the first value can be set according to the actual situation, and the embodiments of the present invention do not specifically limit the first value.

[0177] In this step, first, a pre-set edit distance threshold, that is, the first value, is set. Then, select the IT operation and maintenance words with an edit distance less than the first value as the first candidate words of the first sample word.

[0178] A24. Randomly select a letter from each first sample word as the letter to be replaced, and determine the candidate letters of the letter to be replaced in each first sample word.

[0179] Specifically, after obtaining the letter to be replaced, first obtain the position information of the letter to be replaced on the keyboard. Here, the keyboard refers to a 26-key keyboard. Then, according to the position information of the letter to be replaced on the keyboard, select multiple letters surrounding the letter to be replaced on the keyboard as the candidate letters of the letter to be replaced. It can be understood that the multiple letters surrounding the letter to be replaced refer to the letters that are easily mis-touched when inputting the letter to be replaced through the 26-key keyboard.

[0180] Exemplarily, referring to Figure 4 , Figure 4 is a schematic diagram of the 26-key keyboard provided by the present invention. When the letter to be replaced is F, the letters R, T, D, G, C, V surrounding the letter F are the letters that are easily mis-touched when inputting the letter F through the 26-key keyboard. Therefore, according to the position information of the letter F on the 26-key keyboard, select the letters R, T, D, G, C, V as the candidate letters of the letter F.

[0181] Again exemplarily, referring to Figure 4 , when the letter to be replaced is H, the letters Y, U, I, J, N, B, G surrounding the letter H are the letters that are easily mis-touched when inputting the letter H through the 26-key keyboard. Therefore, according to the position information of the letter H on the 26-key keyboard, select the letters Y, U, I, J, N, B, G as the candidate letters of the letter H.

[0182] A25. Replace the letter to be replaced in each first sample word with the candidate letter of the letter to be replaced to generate the second candidate word of each first sample word.

[0183] In this step, new words are generated by replacing the letters to be replaced in each first sample word with any one of the candidate letters of the letter to be replaced, that is, the second candidate words of each first sample word.

[0184] A26. Obtain the first candidate words and the second candidate words of multiple first sample words as the candidate words of multiple first sample words.

[0185] In this step, the new words generated by the edit distance replacement method and the keyboard mis-touch replacement method are combined as the candidate words of multiple first sample words.

[0186] In some embodiments of the present invention, in step S1023, the implementation process of obtaining the keyword table of the text data set mainly includes the following steps:

[0187] Process the text data set through the TextRank algorithm to generate the keyword table of the text data set.

[0188] It should be noted that the TextRank algorithm is a graph-based ranking algorithm for keyword extraction and document summarization. For a given text, it uses the co-occurrence information or semantics between words within the text to extract the keywords or keyword groups of the text, and uses an extractive automatic summarization method to extract the key sentences of the text.

[0189] In this step, first, extract the keywords of the text data set through the TextRank algorithm to obtain the keyword table of the text data set, and the keyword table is composed of multiple keywords.

[0190] In some embodiments of the present invention, in step S1023, the steps of data augmentation of multiple text sentences to be enhanced according to the keyword table to generate multiple second amplified text sentences include at least one of the following:

[0191] B1. For each text sentence to be enhanced, randomly select the text words belonging to the keyword table from the text sentence to be enhanced as the second sample words, and replace the second sample words with the synonyms of the second sample words to generate the second amplified text sentence corresponding to the text sentence to be enhanced.

[0192] It should be noted that the number of the second sample words can be determined according to the actual situation, and the present invention does not make specific limitations on this.

[0193] In this step, data augmentation is performed on each text statement to be augmented through the Synonym Replacement (SR) method and a keyword list. Specifically, for each text statement to be augmented, a specified number of text words belonging to the keyword list in the text statement to be augmented are randomly selected as the second sample words, and synonyms of the second sample words are randomly selected to replace the second sample words, thereby generating a second augmented text statement corresponding to the text statement to be augmented.

[0194] B2. For each text statement to be augmented, traverse the text words belonging to the keyword list in the text statement to be augmented, and randomly insert synonyms of the text words belonging to the keyword list into the text statement to be augmented to generate a second augmented text statement corresponding to the text statement to be augmented.

[0195] In this step, data augmentation is performed on each text statement to be augmented through the Random Insertion (RI) method and a keyword list. Specifically, for each text statement to be augmented, traverse the text words belonging to the keyword list in the text statement to be augmented, randomly select a text word belonging to the keyword list from the text statement to be augmented, and insert a synonym of the randomly selected text word into any position in the text statement to be augmented, thereby generating a second augmented text statement corresponding to the text statement to be augmented.

[0196] Optionally, the embodiments of the present invention repeat step B2 to achieve multiple data augmentations. Those skilled in the art can understand that the specific method of random insertion can be changed according to the actual situation to achieve repeated execution. In addition, it should be noted that the number of second augmented text statements generated through this step is equal to the number of times this step is repeated.

[0197] Exemplarily, the repeated execution can be that in a single operation, any synonym of a text word belonging to the keyword list in the text statement to be augmented is inserted into a specified position in the text statement to be augmented, and this operation is repeated until all text words in the text statement to be augmented are traversed, thereby generating multiple second augmented text statements.

[0198] Another example is that the repeated execution can also be that in a single operation, a certain synonym of a text word belonging to the keyword list in the text statement to be augmented is inserted into a specified position in the text statement to be augmented, and this operation is repeated until all synonyms of the text word are traversed, thereby generating multiple second augmented text statements.

[0199] Exemplarily, the repeated execution can also be, in a single operation, inserting any synonym of a text word belonging to the keyword table in one of the specified positions of the text sentence to be enhanced, and repeating this operation until all the specified positions in the text sentence to be enhanced are traversed, thereby generating multiple second amplified text sentences.

[0200] B3. For each text sentence to be enhanced, traverse the text words in the text sentence to be enhanced, and swap the positions of two text words in the text sentence to be enhanced to generate a second amplified text sentence corresponding to the text sentence to be enhanced.

[0201] In this step, data enhancement is performed on each text sentence to be enhanced by the Random Swap (RS) method. Specifically, for each text sentence to be enhanced, randomly select two words in the sentence and swap their positions, and repeat this operation a specified number of times, thereby generating multiple second amplified text sentences corresponding to the text sentence to be enhanced.

[0202] It should be noted that the number of second amplified text sentences generated by this step is equal to the number of times this step is repeatedly operated.

[0203] B4. Obtain the word deletion probability. For each text sentence to be enhanced, delete the text words that do not belong to the keyword table from the text sentence to be enhanced according to the word deletion probability to generate a second amplified text sentence corresponding to the text sentence to be enhanced.

[0204] It should be noted that the word deletion probability can be determined according to the actual situation, and the present invention does not make specific limitations thereon.

[0205] In this step, data enhancement is performed on each text sentence to be enhanced by the Random Deletion (RD) method and the keyword table. Specifically, for each text sentence to be enhanced, delete the text words that do not belong to the keyword table from the text sentence to be enhanced according to the word deletion probability, thereby generating a second amplified text sentence corresponding to the text sentence to be enhanced.

[0206] In some embodiments of the present invention, referring to Figure 5 , Figure 5 is a flowchart of sentence-level data enhancement provided by the present invention. In step S1024, data enhancement is performed on multiple text sentences to be enhanced by the back translation method, and the implementation process of generating multiple third amplified text sentences mainly includes the following steps:

[0207] C1. Determine the initial language and the translated language of each text sentence to be enhanced.

[0208] It should be noted that the initial language and the translated language can be Chinese, or other languages such as English, French, German, etc. The present invention does not make specific limitations on this. However, it should be emphasized that the initial language includes one language, and the translated language includes at least one language.

[0209] Optionally, the initial language is English, and the translated language includes at least one of Chinese and French.

[0210] C2. Translate each text statement to be enhanced represented by the initial language into the translated language to obtain each text statement to be enhanced represented by the translated language.

[0211] This step is the forward translation process in the back-translation method. The forward translation process refers to translating a text sentence into another language.

[0212] C3. Translate each text statement to be enhanced represented by the translated language into the initial language to obtain the third amplified text statement corresponding to each text statement to be enhanced, and then generate multiple third amplified text statements.

[0213] This step is the reverse translation process in the back-translation method. The reverse translation process refers to translating a text sentence in another language back into the initial language.

[0214] The following takes an example to illustrate the implementation process of data augmentation at the character level, word level, and sentence level provided by the embodiments of the present invention.

[0215] Refer to Figure 6 , Figure 6 which is a schematic diagram of the application of data augmentation provided by the present invention. Assume that the IT operation and maintenance data “It’s getting very slow process. Some time it’s showing some errormassge.” in the text dataset is obtained as the text statement to be enhanced, and data augmentation is performed on this text statement to be enhanced at the character level (Char-level), word level (Word-level), and sentence level (Sentence-level).

[0216] Specifically, at the character level, based on the data augmentation method of text edit distance and adjacent keyboard input, data augmentation is performed on this text statement to be enhanced. The character “t” in “getting” is replaced with the character “g” to obtain the first amplified text statement of this text statement to be enhanced “It’s getging very slow process. Some time it’sshowing some error massge.”.

[0217] At the word level, the text statement to be enhanced is enhanced through four methods, and four second amplified text statements of the text statement to be enhanced can be obtained. Assuming that the words "error" and "very" belong to the keyword list, there are:

[0218] The text statement to be enhanced is enhanced through the synonym replacement method and the keyword list. The word "error" is replaced with its synonym "wrong", and then the second amplified text statement of the text statement to be enhanced can be obtained: "It’s getting very slow process. Some time it’s showing some wrong massge.";

[0219] The text statement to be enhanced is enhanced through the random insertion method and the keyword list. The synonym "much" of the word "very" is inserted into any position of the text statement to be enhanced, and then the second amplified text statement of the text statement to be enhanced can be obtained: "It’s getting very much slow process. Some time it’s showingsome error massge.";

[0220] The text statement to be enhanced is enhanced through the random swapping method and the keyword list. The positions of the words "slow" and "process" are swapped, and then the second amplified text statement of the text statement to be enhanced can be obtained: "It’sgetting very process slow. Some time it’s showing some error massge.";

[0221] The text statement to be enhanced is enhanced through the random deletion method and the keyword list. The word "slow" is deleted, and then the second amplified text statement of the text statement to be enhanced can be obtained: "It’s getting veryprocess. Some time it’sshowing some error massge."

[0222] At the sentence level, the text statement to be enhanced is enhanced through the back translation method, and the third amplified text statement of the text statement to be enhanced is generated through the forward translation process and the reverse translation process: "This is a very slowprocess. Sometimes some error messages are displayed."

[0223] By performing data augmentation on multiple text statements to be augmented in a text dataset at three granularities: character level, word level, and sentence level, multiple new augmented statements can be generated. Adding the new augmented statements to the original text dataset can constitute the text dataset after data augmentation, which can alleviate the problem of sample class imbalance in the original dataset and increase the data volume of the text dataset, facilitating the improvement of the classification accuracy of the text classification model.

[0224] In some embodiments of the present invention, referring to Figure 7 , Figure 7 is the flowchart of the dataset preprocessing provided by the present invention. In step S103, the implementation process of performing data preprocessing on the text dataset after data augmentation to obtain the preprocessed text dataset may include but is not limited to the following steps:

[0225] S1031, perform text normalization and text cleaning on the text dataset after data augmentation to obtain the text dataset after text cleaning.

[0226] Specifically, the process of text normalization may include but is not limited to uniformly converting the characters in the text dataset after data augmentation to half-width and uniformly converting the English letters in the text dataset after data augmentation to lowercase. The process of text cleaning may include but is not limited to removing invalid characters in the text dataset after data augmentation, such as non-text content, punctuation marks, etc.

[0227] S1032, perform string replacement processing on the text dataset after text cleaning to obtain the text dataset after replacement processing.

[0228] In this step, by performing string replacement processing on the text dataset after text cleaning, the expression of feature words in the IT operation and maintenance field is enhanced, and then the text dataset after replacement processing is obtained.

[0229] S1033, perform word segmentation on the text dataset after replacement processing to obtain the word-level tokens and character-level tokens of the text dataset.

[0230] Specifically, use the NLTK natural language processing tool to perform word segmentation on the text dataset after replacement processing to obtain the word-level tokens and character-level tokens of the text dataset. It can be understood that NLTK is short for Natural Language Toolkit. NLTK is one of the toolkits in the field of natural language processing, which includes many libraries and datasets and can be used to complete various natural language processing tasks.

[0231] S1034, perform word cleaning and normalization processing on the word-level tokens of the text dataset to obtain the preprocessed text dataset.

[0232] Specifically, the process of word cleaning may include, but is not limited to, cleaning the word-level tokens of the text dataset. The methods of word cleaning may include, but are not limited to, at least one of stop word removal, part-of-speech tagging, lemmatization, and spelling correction. The normalization process may include, but is not limited to, restoring the part-of-speech and tense of the word-level tokens of the text dataset.

[0233] In some embodiments of the present invention, the related data augmentation technology lacks the capture and utilization of the feature words in the field of IT operation and maintenance, resulting in a large number of redundant features in the dataset. To address this problem, the embodiments of the present invention propose a step of string replacement processing. Specifically, in step S1032, the implementation process of performing string replacement processing on the text dataset after text cleaning to obtain the text dataset after replacement processing may include, but is not limited to, the following steps:

[0234] First, perform rule recognition on the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset.

[0235] In this step, according to the characteristics of different types of IT operation and maintenance feature words, use regular expressions to perform rule recognition on the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset. Optionally, the IT operation and maintenance feature words may include, but are not limited to, at least one of date, time, percentage, IP address, memory address, memory capacity information (i.e., memory / file size), website URL, phone number, network status code, and error code.

[0236] It should be noted that the matching rules of the regular expressions can be set according to the actual situation, and the rule recognition of the text dataset after text cleaning is completed through the matching rules of the regular expressions. The present invention does not make specific limitations on this.

[0237] Exemplarily, when the IT operation and maintenance feature word is a percentage, the regular expression for identifying the percentage type is "\d{1,3}(?:\.\d{1,2})?%". The matching rules of this regular expression may include, but are not limited to, the integer part (i.e., "\d{1,3}"), the decimal point (i.e., "\."), the decimal part (i.e., "\d{1,2}"), and the percent sign (i.e., "%"). Among them, "\d{1,3}" means that the integer part contains 1-3 digits, "\d{1,2}" means that the decimal part contains 1-2 digits, ":" means the separator, and "?" means that both the decimal part and the percent sign part are optional parts.

[0238] Exemplarily, when the IT operation and maintenance feature word is a date, the regular expression for identifying the date type is "\d{4}-\d{1,2}-\d{1,2}". The matching rules of this regular expression can include but are not limited to year (i.e., "\d{4}"), month (i.e., "-\d{1,2}"), and day (i.e., "-\d{1,2}"). Among them, "\d{4}" means the year is represented by four digits, and "-\d{1,2}" means the month or day is represented by 1 - 2 digits.

[0239] Exemplarily, when the IT operation and maintenance feature word is a time, the regular expression for identifying the time type is "^\\d{1,2}:\\d{1,2}(:\\d{1,2}(.\\d{1,3})?)?$". The matching rules of this regular expression can include but are not limited to matching the entire string from the beginning to the end using "^" and "$", hour (i.e., "\\d{1,2}"), minute (i.e., "\\d{1,2}"), second (i.e., "\\d{1,2}"), and millisecond (i.e., "\\d{1,3}"). Among them, ":" represents the separator, "?" means both the second and millisecond are optional parts, "\\d{1,2}" means the hour, minute, or second is represented by 1 - 2 digits, and "\\d{1,3}" means the millisecond is represented by 1 - 3 digits.

[0240] Then, obtain the identifier of the IT operation and maintenance feature word, and replace the IT operation and maintenance feature word in the text dataset with the identifier of the IT operation and maintenance feature word, thereby obtaining the text dataset after replacement processing.

[0241] In this step, use the identifier corresponding to the IT operation and maintenance feature word identified by the rule to replace the IT operation and maintenance feature word, generate multiple new strings, and thereby obtain the text dataset after replacement processing.

[0242] Exemplarily, referring to Figure 8 , Figure 8 is a schematic diagram of the identifier provided by the present invention. Figure 8 In <date>”, using " <date>Replace "date" with "". For another example, for the data "0xc000000f" in the text dataset after text cleaning, by performing rule recognition on "0xc000000f", the IT operation and maintenance feature word "error code" can be obtained, and the identifier corresponding to "error code" can be retrieved <error>”, using " <error>Replace "error code" with "".

[0243] In some embodiments of the present invention, in step S104, the process of converting the preprocessed text data set into a text vector representation to obtain a text vector data set mainly includes the following steps:

[0244] S1041, Encode the character-level tokens of the preprocessed text data set to generate the character vector features of the text data set.

[0245] In this step, the one-hot encoding method is used to encode the character-level tokens of the preprocessed text data set to generate the character vector features of the text data set.

[0246] It should be noted that one-hot encoding refers to using 0 and 1 to represent some parameters, using a multi-bit status register to encode multiple states, each state has its own independent register bit, and at any time, only one bit is valid. That is, only one bit is 1 and the rest are 0.

[0247] Optionally, in other embodiments of the present invention, other methods such as the bag-of-words model, TF-IDF, etc. can also be used to encode the character-level tokens of the preprocessed text data set.

[0248] S1042, Encode the word-level tokens of the preprocessed text data set to generate the word vector features of the text data set.

[0249] In this step, the pre-trained word vector model is used to encode the word-level tokens of the preprocessed text data set to generate the word vector features of the text data set.

[0250] It should be noted that the word vector model can be a Word2Vec model or other word vector models such as CBOW, Skip-gram, etc. The present invention does not make specific limitations on this.

[0251] Exemplarily, taking the Word2Vec model as the word vector model, the pre-trained Word2Vec model is used to encode the word-level tokens of the preprocessed text data set to generate the Word2Vec vectors of the text data set as the word vector features.

[0252] S1043, Construct a text vector data set through the character vector features and word vector features of the text data set.

[0253] In some embodiments of the present invention, in step S105, the process of extracting features from the text vector data set to obtain a text enhancement data set may include but is not limited to the following steps:

[0254] S1051. Concatenate the character vector features and word vector features in the text vector dataset and perform text length normalization processing to obtain a text enhancement dataset.

[0255] It should be noted that text length normalization is to set the lengths of all sample data in the dataset to a reasonable unified length for easy input into the text classification model for classification. Here, the length refers to the total number of tokens in each sample data.

[0256] In this step, first, concatenate the character vector features and word vector features in the text vector dataset to obtain the concatenated vector features of the text vector dataset. Then, unify the text lengths of the concatenated vector features of the text vector dataset to a second value.

[0257] It can be understood that the second value can be determined according to the actual situation, and the present invention does not make specific limitations thereto.

[0258] Optionally, use the method of truncating and padding to unify the text lengths of the concatenated vector features of the text vector dataset to a second value. Specifically, when the lengths of the character vector features and word vector features of the text dataset are greater than the specified length, only intercept the part of the specified length from the character vector features and word vector features; when the lengths of the character vector features and word vector features of the text dataset are less than or equal to the specified length, use zero vectors to supplement the empty positions.

[0259] Next, an example will be used to illustrate the principle of the data enhancement method proposed in the embodiments of the present invention. Refer to Figure 9 , Figure 9 is the schematic diagram of the data enhancement method for operation and maintenance data provided by the present invention. The principle of data enhancement is as follows:

[0260] The first step is to obtain a number of IT operation and maintenance data as the text dataset.

[0261] The second step is to perform data enhancement on the text dataset at the character level, word level, and sentence level.

[0262] The third step is to preprocess the text dataset after data enhancement. The preprocessing process includes: first, perform text normalization and text cleaning, such as full-width and half-width character conversion, case conversion of letters, removal of non-text content, filtering of punctuation marks, etc.; then perform string replacement processing to replace the IT operation and maintenance feature words in the text dataset with corresponding identifiers; then perform word segmentation to obtain the word-level tokens and character-level tokens of the text dataset, and perform word cleaning and normalization on the word-level tokens, such as removing stop words, part-of-speech tagging, part-of-speech reduction, spelling correction, etc.

[0263] In the fourth step, perform text vector representation on the preprocessed text data set, that is, vectorize the tokens through one-hot encoding, Word2Vec, etc. to obtain a text vector data set.

[0264] In the fifth step, perform feature selection and combination on the text vector data set, and standardize the text length, thereby obtaining a text enhanced data set to achieve data enhancement.

[0265] Secondly, an implementation step of a text classification method for operation and maintenance data according to an embodiment of the present invention will be described in detail with reference to the accompanying drawings. Refer to Figure 10 , Figure 10 FIG. is a flowchart of a text classification method for operation and maintenance data provided by the present invention. The text classification method may include but is not limited to the following steps:

[0266] S201, obtain IT operation and maintenance data to be classified.

[0267] S202, input the IT operation and maintenance data to be classified into the trained text classification model for text classification to obtain a text classification result.

[0268] It should be noted that the text classification model is obtained by pre-training the initial classification model using the text enhanced data set, and the text enhanced data set is obtained through the data enhancement method for operation and maintenance data described above.

[0269] Optionally, the initial classification model may be a TextRCNN model, or other classification models for natural language processing such as TextCNN and TextRNN. The present invention does not make specific limitations thereto.

[0270] Exemplarily, taking the TextRCNN model as the initial classification model, obtain the text enhanced data set through the data enhancement method for operation and maintenance data described above. Use the text enhanced data set as the input of the TextRCNN model, train the TextRCNN model through the text enhanced data set, output the trained TextRCNN model as the text classification model, and use the text classification model to perform text classification on the IT operation and maintenance data to be classified, thereby obtaining a text classification result.

[0271] In summary, the present invention provides a data augmentation method, a text classification method, and an electronic device for operation and maintenance data. The data augmentation method includes: first, obtaining a plurality of IT operation and maintenance data as a text data set, then performing data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set, then performing data preprocessing on the augmented text data set to obtain a preprocessed text data set, converting the preprocessed text data set into a text vector representation to obtain a text vector data set, and finally extracting features from the text vector data set to obtain a text augmented data set, thereby realizing the text augmentation process of IT operation and maintenance data; the text classification method includes: obtaining the IT operation and maintenance data to be classified, and using a text classification model to perform text classification on the IT operation and maintenance data to be classified. Among them, the generation process of the text classification model includes: obtaining a text augmented data set through the data augmentation method described above, and then training an initial classification model using the text augmented data set to obtain a text classification model.

[0272] For the data augmentation method, on the one hand, the present invention performs data augmentation on IT operation and maintenance data at three different granularities of the character level, word level, and sentence level, increasing the syntactic diversity of the original data set and effectively alleviating the problem of sample class imbalance in the original data set; on the other hand, the present invention can eliminate redundant features in the original data set that have little impact on text classification and improve the contribution of specific type symbol representations to the classification of IT operation and maintenance data by adding string replacement processing during the preprocessing process; on the other hand, the present invention realizes multi-dimensional text feature extraction by combining text vector representations and standardizations at the character level and word level, overcoming the problem of feature sparsity of short text samples. The data augmentation method proposed by the present invention has higher inclusiveness for the text data set, which is conducive to generating high-quality data sets.

[0273] For the text classification method, the present invention uses the data set generated by the data augmentation method described above to train the text classification model, which can improve the generalization ability of the text classification model, and thus improve the classification accuracy of the text classification model in the operation and maintenance data text classification task.

[0274] In addition, an embodiment of the present invention also provides an electronic device, including:

[0275] At least one processor;

[0276] At least one memory for storing at least one program;

[0277] When the at least one program is executed by the at least one processor, the at least one processor implements the data augmentation method for operation and maintenance data described above, or the text classification method for operation and maintenance data described above.

[0278] The content in the above method embodiments is applicable to the embodiments of this electronic device. The functions specifically implemented in the embodiments of this electronic device are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0279] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0280] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.< / error> < / error> < / date> < / date>

Claims

1. A method for data augmentation of operation and maintenance data, characterized in that, Including the following steps: Obtain a number of IT operation and maintenance data as a text data set; Perform data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set; Perform data preprocessing on the augmented text data set to obtain a preprocessed text data set; Convert the preprocessed text data set into a text vector representation to obtain a text vector data set; Extract features from the text vector data set to obtain a text augmented data set.

2. The method for data augmentation of operation and maintenance data according to claim 1, characterized in that, The performing data augmentation on the text data set at the character level, word level, and sentence level to generate an augmented text data set includes: Obtain the IT operation and maintenance data belonging to the few-shot categories in the text data set as text statements to be augmented, and each text statement to be augmented consists of a number of text words; Randomly select a number of text words from multiple text statements to be augmented, and perform data augmentation on the multiple text statements to be augmented according to the randomly selected text words to generate multiple first amplified text statements; Obtain the keyword table of the text data set, and perform data augmentation on the multiple text statements to be augmented according to the keyword table to generate multiple second amplified text statements; Perform data augmentation on the multiple text statements to be augmented by back translation to generate multiple third amplified text statements; Add the multiple first amplified text statements, second amplified text statements, and third amplified text statements to the text data set to generate an augmented text data set.

3. The method for data augmentation of operation and maintenance data according to claim 2, characterized in that, The randomly selecting a number of text words from multiple text statements to be augmented and performing data augmentation on the multiple text statements to be augmented according to the randomly selected text words to generate multiple first amplified text statements includes: Randomly select a number of text words from multiple text statements to be augmented as first sample words; Perform data augmentation on the multiple first sample words by edit distance replacement method and keyboard mis-touch replacement method to generate candidate words for the multiple first sample words; In each text statement to be augmented, replace the first sample word with the candidate word of the first sample word to generate the first amplified text statement corresponding to each text statement to be augmented, and thus obtain multiple first amplified text statements.

4. The method for data augmentation of operation and maintenance data according to claim 3, characterized in that, The performing data augmentation on the multiple text statements to be augmented according to the keyword table to generate multiple second amplified text statements includes at least one of the following: For each text statement to be augmented, randomly select the text words belonging to the keyword table from the text statement to be augmented as second sample words, and replace the second sample words with synonyms of the second sample words to generate the second amplified text statement corresponding to the text statement to be augmented; For each text statement to be augmented, traverse the text words belonging to the keyword table in the text statement to be augmented, and randomly insert the synonyms of the text words belonging to the keyword table into the text statement to be augmented to generate the second amplified text statement corresponding to the text statement to be augmented; For each text statement to be enhanced, traverse the text words in the text statement to be enhanced, swap the positions of two text words in the text statement to be enhanced, and generate a second augmented text statement corresponding to the text statement to be enhanced; Obtain the word deletion probability. For each text statement to be enhanced, delete the text words that do not belong to the keyword table from the text statement to be enhanced according to the word deletion probability, and generate a second augmented text statement corresponding to the text statement to be enhanced.

5. The method for data augmentation of operation and maintenance data according to claim 1, characterized in that, The data preprocessing of the text data set after data augmentation to obtain the preprocessed text data set includes: Perform text normalization and text cleaning on the text data set after data augmentation to obtain the text data set after text cleaning; Perform string replacement processing on the text data set after text cleaning to obtain the text data set after replacement processing; Perform word segmentation on the text data set after replacement processing to obtain the word-level tokens and character-level tokens of the text data set; Perform word cleaning and normalization processing on the word-level tokens of the text data set to obtain the preprocessed text data set.

6. The method for data augmentation of operation and maintenance data according to claim 5, characterized in that, The string replacement processing of the text data set after text cleaning to obtain the text data set after replacement processing includes: Perform rule recognition on the text data set after text cleaning to obtain the IT operation and maintenance feature words of the text data set; Obtain the identifiers of the IT operation and maintenance feature words, and replace the IT operation and maintenance feature words in the text data set with the identifiers of the IT operation and maintenance feature words, so as to obtain the text data set after replacement processing.

7. The method for data augmentation of operation and maintenance data according to claim 1, characterized in that, The conversion of the preprocessed text data set into a text vector representation to obtain a text vector data set includes: Encode the character-level tokens of the preprocessed text data set to generate the character vector features of the text data set; Encode the word-level tokens of the preprocessed text data set to generate the word vector features of the text data set; Construct a text vector data set through the character vector features and word vector features of the text data set.

8. A method for data enhancement of operation and maintenance data according to claim 7, wherein, The feature extraction of the text vector data set to obtain a text augmentation data set includes: Perform splicing processing and text length normalization processing on the character vector features and word vector features in the text vector data set to obtain a text augmentation data set.

9. A method for text classification of operation and maintenance data, wherein, Includes the following steps: Obtain the IT operation and maintenance data to be classified; Input the IT operation and maintenance data to be classified into the trained text classification model for text classification to obtain a text classification result; Wherein, the text classification model is pre-trained on the initial classification model by using the text augmentation data set, and the text augmentation data set is obtained by using a data augmentation method for operation and maintenance data as described in any one of claims 1-8.

10. An electronic device, wherein, Includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a data augmentation method for operation and maintenance data as described in any one of claims 1-8, or implements a text classification method for operation and maintenance data as described in any one of claims 9.