Data augmentation method for operation and maintenance data, text classification method and electronic device

By adopting multi-grained data augmentation method and text vector representation technology in IT operation and maintenance data, the problems of text imbalance and sparse characteristics in IT operation and maintenance data are solved, and the performance of text classification model is improved.

WO2025124150A1PCT designated stage expired Publication Date: 2025-06-19E SURFING VISION TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135191
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-11-28
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The unstructured text in IT operation and maintenance data contains a large number of vocabulary in the IT operation and maintenance field due to manual recording. The text length gap is large, resulting in frequent colloquial and splicing grammar errors, and the data category imbalance in the labeled data set, which affects the classification effect of text feature extraction and deep learning models.

Method used

The data enhancement method at character level, word level and sentence level is adopted to enhance IT operation and maintenance data by multi-angle enhancement, and the text data set after data enhancement is generated, and text features are extracted through data preprocessing and text vector representation, and finally the text classification model is trained.

Benefits of technology

It effectively alleviates the problem of sample category imbalance in the original dataset and the sparse feature of short text samples, eliminates redundant features, and improves the generalization ability and classification accuracy of the text classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135191_19062025_PF_FP_ABST
    Figure CN2024135191_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a data augmentation method for operation and maintenance data, a text classification method and an electronic device. The augmentation method comprises: acquiring a plurality of pieces of operation and maintenance data as a text data set, and performing data augmentation on same; performing data preprocessing on the text data set after data augmentation; and converting the preprocessed text data set into a text vector representation to obtain a text vector data set, and performing feature extraction on same to obtain a text augmented data set. The classification method comprises: acquiring operation and maintenance data to be classified, and using a text classification model to perform text classification on said operation and maintenance data, the text classification model being obtained by training a text augmented data set. The present invention can ameliorate the problems such as unbalanced sample categories in data sets and sparse short-text features, and effectively eliminates redundant features having small impact on text classification, thus improving the quality of data sets, and improving the generalization ability and classification performance of text classification models. The present invention is applied to the technical field of natural language processing.
Need to check novelty before this filing date? Find Prior Art

Description

Data enhancement method for operation and maintenance data, text classification method, and electronic device Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a data enhancement method, a text classification method, and an electronic device for operation and maintenance data. Background Art

[0002] IT (Information Technology) operations data refers to text data generated during IT operations, primarily in the form of work tickets and logs. Its content consists of both structured and unstructured text. Currently, deep learning models such as TextCNN, TextRNN, and Transformer are commonly used to perform text classification on IT operations data.

[0003] However, the unstructured text in IT operation and maintenance data is usually recorded manually, which contains a large number of IT operation and maintenance vocabulary, and the text length differences between different unstructured texts are large, which will lead to the frequent occurrence of problems such as colloquialism and splicing grammatical errors. Moreover, the annotated datasets used for text classification usually have a more serious data category imbalance problem, which undoubtedly creates a certain degree of difficulty in extracting effective text feature information, and thus makes the existing deep learning models have poor classification effects on the text classification task of IT operation and maintenance data.

[0004] In response to the problem of data category imbalance in labeled datasets, some related technologies use data augmentation technology to alleviate this problem. Currently, most commonly used text data augmentation technologies are based on adding noise or back-translation methods. However, these common text data augmentation technologies do not retain important feature words in the original text data that contribute greatly to text classification when implementing data augmentation. Moreover, they have limited enhancement effects on short texts with sparse features and low diversity of the added text corpus. The effect of data augmentation still needs to be improved. Summary of the Invention

[0005] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.

[0006] To this end, the purpose of the present invention is to provide a data enhancement method for operation and maintenance data, a text classification method and an electronic device.

[0007] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:

[0008] In one aspect, an embodiment of the present invention provides a method for enhancing operation and maintenance data, comprising the following steps:

[0009] Obtain several IT operation and maintenance data as text datasets;

[0010] Performing data enhancement on the text dataset at the character level, word level, and sentence level to generate a data-enhanced text dataset;

[0011] Perform data preprocessing on the data-enhanced text dataset to obtain a preprocessed text dataset;

[0012] Convert the preprocessed text dataset into a text vector representation to obtain a text vector dataset;

[0013] Feature extraction is performed on the text vector dataset to obtain a text enhancement dataset.

[0014] In addition, the data enhancement method for operation and maintenance data according to the above embodiment of the present invention may also have the following additional technical features:

[0015] Furthermore, in one embodiment of the present invention, the acquiring of a plurality of IT operation and maintenance data as a text data set includes:

[0016] Obtain several IT operation and maintenance data and corresponding text categories as text datasets.

[0017] Furthermore, in one embodiment of the present invention, performing data enhancement on the text dataset at the character level, word level, and sentence level to generate the data-enhanced text dataset includes:

[0018] Obtaining IT operation and maintenance data belonging to a small sample category in the text dataset as text sentences to be enhanced, each of which is composed of a plurality of text words;

[0019] Randomly selecting a plurality of text words from a plurality of text sentences to be enhanced, performing data enhancement on the plurality of text sentences to be enhanced according to the plurality of randomly selected text words, and generating a plurality of first augmented text sentences;

[0020] Obtaining a keyword table of the text data set, performing data enhancement on a plurality of text sentences to be enhanced according to the keyword table, and generating a plurality of second augmented text sentences;

[0021] Performing data augmentation on multiple text sentences to be augmented by back-translation to generate multiple third augmented text sentences;

[0022] A plurality of first augmented text sentences, second augmented text sentences, and third augmented text sentences are added to the text dataset to generate a data-enhanced text dataset.

[0023] Furthermore, in one embodiment of the present invention, randomly selecting a plurality of text words from a plurality of text sentences to be enhanced, performing data enhancement on the plurality of text sentences to be enhanced according to the plurality of randomly selected text words, and generating a plurality of first augmented text sentences, includes:

[0024] Randomly selecting multiple text words from multiple text sentences to be enhanced as first sample words;

[0025] Performing data enhancement on the plurality of first sample words by using an edit distance replacement method and a keyboard accidental touch replacement method to generate candidate words for the plurality of first sample words;

[0026] In each text sentence to be enhanced, the first sample word is replaced with a candidate word of the first sample word to generate a first augmented text sentence corresponding to each text sentence to be enhanced, thereby obtaining multiple first augmented text sentences.

[0027] Furthermore, in one embodiment of the present invention, performing data enhancement on the plurality of first sample words by using the edit distance replacement method and the keyboard accidental touch replacement method to generate candidate words for the plurality of first sample words includes:

[0028] Obtaining an IT operation and maintenance word database, the IT operation and maintenance word database including a plurality of IT operation and maintenance words, and then obtaining an edit distance between each first sample word and each IT operation and maintenance word;

[0029] For each first sample word, obtaining an IT operation and maintenance word having an edit distance less than a first value as a first candidate word for the first sample word;

[0030] Randomly selecting a letter from each first sample word as a letter to be replaced, and determining a candidate letter for the letter to be replaced in each first sample word;

[0031] Replacing the letters to be replaced in each first sample word with candidate letters of the letters to be replaced, and generating a second candidate word for each first sample word;

[0032] A first candidate word and a second candidate word of the plurality of first sample words are obtained as candidate words of the plurality of first sample words.

[0033] Furthermore, in one embodiment of the present invention, obtaining the edit distance between each first sample word and each IT operation and maintenance word includes:

[0034] Obtaining a character editing operation number for converting the first sample word into an IT operation and maintenance word, where the character editing operation number is a minimum value of the number of execution times of the character editing operation;

[0035] The edit distance between each first sample word and each IT operation and maintenance word is calculated according to the number of character edit operations.

[0036] Furthermore, in one embodiment of the present invention, the character editing operation includes at least one of inserting a character, deleting a character, and changing a character.

[0037] Furthermore, in one embodiment of the present invention, determining candidate letters for letters to be replaced in each first sample word includes:

[0038] Obtaining position information of the letter to be replaced on the keyboard;

[0039] According to the position information of the letter to be replaced on the keyboard, a plurality of letters surrounding the letter to be replaced on the keyboard are selected as candidate letters for the letter to be replaced.

[0040] Furthermore, in one embodiment of the present invention, obtaining the keyword table of the text dataset includes:

[0041] The text dataset is processed using a TextRank algorithm to generate a keyword table for the text dataset.

[0042] Furthermore, in one embodiment of the present invention, performing data enhancement on a plurality of text sentences to be enhanced according to the keyword table to generate a plurality of second augmented text sentences includes at least one of the following:

[0043] For each text sentence to be enhanced, randomly selecting a text word from the keyword table as a second sample word from the text sentence to be enhanced, replacing the second sample word with a synonym of the second sample word, and generating a second augmented text sentence corresponding to the text sentence to be enhanced;

[0044] For each text sentence to be enhanced, traverse the text words in the keyword table in the text sentence to be enhanced, randomly insert synonyms of the text words in the keyword table into the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced;

[0045] For each text sentence to be enhanced, traverse the text words in the text sentence to be enhanced, swap the positions of two text words in the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced;

[0046] Obtain a word deletion probability, and for each text sentence to be enhanced, delete text words that do not belong to the keyword table from the text sentence to be enhanced according to the word deletion probability, and generate a second augmented text sentence corresponding to the text sentence to be enhanced.

[0047] Furthermore, in one embodiment of the present invention, performing data enhancement on a plurality of text sentences to be enhanced by back-translation to generate a plurality of third augmented text sentences includes:

[0048] Determine the original language and the translated language of each text sentence to be enhanced;

[0049] Translate each text sentence to be enhanced represented by the initial language into a translated language to obtain each text sentence to be enhanced represented by the translated language;

[0050] Each text sentence to be enhanced represented by the translation language is translated into the original language to obtain a third augmented text sentence corresponding to each text sentence to be enhanced, thereby generating a plurality of third augmented text sentences.

[0051] Furthermore, in one embodiment of the present invention, the original language is English, and the translated language includes at least one of Chinese and French.

[0052] Furthermore, in one embodiment of the present invention, performing data preprocessing on the data-enhanced text dataset to obtain a preprocessed text dataset includes:

[0053] Perform text standardization and text cleaning on the text dataset after data enhancement to obtain a text dataset after text cleaning;

[0054] Perform string replacement processing on the text dataset after text cleaning to obtain a text dataset after replacement processing;

[0055] Performing word segmentation on the replaced text dataset to obtain word-level lemmas and character-level lemmas of the text dataset;

[0056] The word-level tokens of the text dataset are cleaned and normalized to obtain a preprocessed text dataset.

[0057] Furthermore, in one embodiment of the present invention, the text data set after data enhancement is subjected to text standardization and text cleaning to obtain a text data set after text cleaning, including:

[0058] Convert the characters of the data-enhanced text dataset to half-width characters and convert the English letters of the data-enhanced text dataset to lowercase.

[0059] Remove invalid characters from the text dataset after data augmentation.

[0060] Furthermore, in one embodiment of the present invention, performing string replacement processing on the text dataset after text cleaning to obtain the replaced text dataset includes:

[0061] Perform rule recognition on the cleaned text dataset to obtain IT operation and maintenance feature words of the text dataset;

[0062] An identifier of the IT operation and maintenance feature word is obtained, and the IT operation and maintenance feature word in the text data set is replaced with the identifier of the IT operation and maintenance feature word, thereby obtaining a text data set after replacement.

[0063] Furthermore, in one embodiment of the present invention, the IT operation and maintenance feature words include at least one of date, time, percentage, IP address, memory address, memory capacity information, website, telephone number, network status code and error code.

[0064] Furthermore, in one embodiment of the present invention, the step of performing rule recognition on the cleaned text dataset to obtain IT operation and maintenance feature words of the text dataset includes:

[0065] Regular expressions are used to perform rule recognition on the text dataset after text cleaning to obtain IT operation and maintenance feature words of the text dataset.

[0066] Furthermore, in one embodiment of the present invention, performing word segmentation on the replaced text dataset to obtain word-level tokens and character-level tokens of the text dataset includes:

[0067] The NLTK natural language processing tool is used to perform word segmentation on the replaced text dataset to obtain word-level lemmas and character-level lemmas of the text dataset.

[0068] Furthermore, in one embodiment of the present invention, the word-level word-grams of the text dataset are cleaned and normalized to obtain a preprocessed text dataset, including:

[0069] The word-level word-grams of the text data set are subjected to word cleaning processing, and the part of speech and tense of the word-level word-grams of the text data set are restored to obtain a preprocessed text data set.

[0070] Furthermore, in one embodiment of the present invention, the word cleaning process includes at least one of stop word removal, part-of-speech tagging, morphological restoration, and splicing error correction.

[0071] Furthermore, in one embodiment of the present invention, converting the preprocessed text dataset into a text vector representation to obtain the text vector dataset includes:

[0072] Encoding the character-level tokens of the preprocessed text dataset to generate word vector features of the text dataset;

[0073] Encoding the word-level tokens of the preprocessed text dataset to generate word vector features of the text dataset;

[0074] A text vector dataset is constructed using the character vector features and word vector features of the text dataset.

[0075] Furthermore, in one embodiment of the present invention, encoding the character-level tokens of the preprocessed text dataset to generate word vector features of the text dataset includes:

[0076] The character-level tokens of the preprocessed text dataset are encoded using a one-hot encoding method to generate word vector features of the text dataset.

[0077] Furthermore, in one embodiment of the present invention, encoding the word-level tokens of the preprocessed text dataset to generate word vector features of the text dataset includes:

[0078] The pre-trained word vector model is used to encode the word-level tokens of the pre-processed text dataset to generate word vector features of the text dataset.

[0079] Furthermore, in one embodiment of the present invention, the word vector model is a Word2Vec model.

[0080] Furthermore, in one embodiment of the present invention, the feature extraction of the text vector dataset to obtain the text enhancement dataset includes:

[0081] The character vector features and word vector features in the text vector dataset are concatenated and the text length is normalized to obtain a text enhancement dataset.

[0082] Furthermore, in one embodiment of the present invention, the concatenation processing and text length normalization processing of the character vector features and word vector features in the text vector dataset include:

[0083] Concatenate the character vector features and word vector features in the text vector dataset to obtain the concatenated vector features of the text vector dataset;

[0084] The text lengths of the concatenated vector features of the text vector dataset are unified into a second value.

[0085] On the other hand, an embodiment of the present invention provides a text classification method for operation and maintenance data, comprising the following steps:

[0086] Obtain IT operation and maintenance data to be classified;

[0087] Inputting the IT operation and maintenance data to be classified into the trained text classification model for text classification to obtain text classification results;

[0088] The text classification model is obtained by pre-training an initial classification model using a text enhancement dataset, and the text enhancement dataset is obtained by the data enhancement method of the operation and maintenance data described above.

[0089] Furthermore, in one embodiment of the present invention, the initial classification model is a TextRCNN model.

[0090] In another aspect, an embodiment of the present invention provides an electronic device, including:

[0091] at least one processor;

[0092] at least one memory for storing at least one program;

[0093] When the at least one program is executed by the at least one processor, the at least one processor implements the aforementioned data enhancement method for operation and maintenance data, or the aforementioned text classification method for operation and maintenance data.

[0094] The beneficial effects of the present invention are: providing a data enhancement method for operation and maintenance data, a text classification method and an electronic device, the data enhancement method comprising: first obtaining a number of IT operation and maintenance data as a text data set, then performing data enhancement on the text data set at the character level, word level and sentence level to generate a data-enhanced text data set, then performing data preprocessing on the data-enhanced text data set to obtain a preprocessed text data set, and converting the preprocessed text data set into a text vector representation to obtain a text vector data set, and finally performing feature extraction on the text vector data set to obtain a text enhancement data set, thereby realizing text enhancement processing of the IT operation and maintenance data; the text classification method comprising: obtaining IT operation and maintenance data to be classified, performing text classification on the IT operation and maintenance data to be classified using a text classification model, the text classification model being obtained by training an initial classification model using the text enhancement data set, and the text enhancement data set being obtained by the data enhancement method described above. The present invention performs data enhancement on IT operation and maintenance data at three different granularities: character level, word level, and sentence level, thereby increasing sentence diversity for the original data set. This can effectively alleviate the sample category imbalance problem and the feature sparsity problem of short text samples in the original data set, eliminate redundant features in the original data set that have little impact on text classification, realize multi-dimensional text feature extraction, and facilitate the generation of high-quality data sets. In addition, using the data set generated by the data enhancement method of the present invention to train a text classification model can improve the generalization ability of the text classification model, thereby improving the classification accuracy of the text classification model in the operation and maintenance data text classification task.

[0095] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] FIG1 is a flow chart of a method for enhancing operation and maintenance data provided by the present invention;

[0097] FIG2 is a flow chart of data enhancement provided by the present invention;

[0098] FIG3 is a flow chart of character-level data enhancement provided by the present invention;

[0099] FIG4 is a schematic diagram of a 26-key keyboard provided by the present invention;

[0100] FIG5 is a flowchart of sentence-level data enhancement provided by the present invention;

[0101] FIG6 is a schematic diagram of an application of data enhancement provided by the present invention;

[0102] FIG7 is a flow chart of data set preprocessing provided by the present invention;

[0103] FIG8 is a schematic diagram of an identifier provided by the present invention;

[0104] FIG9 is a schematic diagram of a data enhancement method for operation and maintenance data provided by the present invention;

[0105] FIG10 is a flowchart of a text classification method for operation and maintenance data provided by the present invention. DETAILED DESCRIPTION

[0106] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0107] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be considered as limiting the present application. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0108] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0109] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0110] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0111] 1) Text classification (TC), also known as automatic text classification, is the process by which a computer maps a piece of information-containing text to a predetermined category or categories. The algorithmic model that implements this process is called a classifier. Text classification is a classic problem in the field of natural language processing (NLP).

[0112] 2) Data augmentation (DA) refers to a method of generating a large amount of labeled data from a small amount of labeled data while ensuring that the label semantics remain unchanged as much as possible. Data augmentation can increase the number of training samples and enhance the generalization ability and robustness of the model.

[0113] IT (Information Technology) operations data refers to text data generated during IT operations, primarily in the form of work tickets, logs, and other data. Its content consists of both structured and unstructured text. Currently, deep learning models such as TextCNN and TextRNN are commonly used to perform text classification on IT operations data.

[0114] However, the unstructured text in IT operation and maintenance data is usually recorded manually, which contains a large number of IT operation and maintenance vocabulary, and the text length differences between different unstructured texts are large, which will lead to the frequent occurrence of problems such as colloquialism and splicing grammatical errors. Moreover, the annotated datasets used for text classification usually have a more serious data category imbalance problem, which undoubtedly creates a certain degree of difficulty in extracting effective text feature information, and thus makes the existing deep learning models have poor classification effects on the text classification task of IT operation and maintenance data.

[0115] To address the problem of data category imbalance in labeled datasets, some related technologies use data enhancement technology to alleviate this problem. Currently, most commonly used text data enhancement technologies are based on adding noise or back-translation methods. On the one hand, related text data enhancement technologies lack the capture and utilization of feature words in the IT operation and maintenance field, which leads to the fact that they do not retain important feature words in the original text data that contribute greatly to text classification when performing data enhancement, resulting in feature redundancy. On the other hand, related text enhancement technologies have limited enhancement effects on short texts with sparse features, and there are problems such as low diversity of the increased text corpus. The effect of data enhancement still needs to be improved.

[0116] In response to the problems existing in related technologies, such as imbalanced dataset categories and sparse short text features, a lack of capture and utilization of characteristic words in the IT operation and maintenance field during the data enhancement process, and low classification model accuracy, the present invention provides a data enhancement method, a text classification method, and an electronic device for operation and maintenance data. The method enhances IT operation and maintenance data at three different granularities: character level, word level, and sentence level. The data set generated by the data enhancement is used to train a text classification model, and the trained text classification model is used to implement text classification of IT operation and maintenance data. While increasing the data volume of the dataset, the present invention can effectively alleviate the sample category imbalance problem in the original dataset and the feature sparsity problem of short text samples, thereby helping to improve the generalization ability of the text classification model and the classification accuracy of the text classification model in the text classification task of operation and maintenance data.

[0117] The embodiments of the present application are further described below with reference to the accompanying drawings.

[0118] First, the following describes in detail an implementation step of a data enhancement method for operation and maintenance data proposed in an embodiment of the present invention with reference to the accompanying drawings.

[0119] The methods in the embodiments of the present invention can be applied to terminals or servers, or can be software running on terminals or servers. A terminal can be, but is not limited to, a tablet computer, laptop computer, or desktop computer. A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Furthermore, a server can be, but is not limited to, a node server in a blockchain network. Blockchain represents a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0120] 1 , which is a flow chart of a method for enhancing operation and maintenance data provided by the present invention, the method may include but is not limited to the following steps:

[0121] S101, obtaining a number of IT operation and maintenance data as a text dataset.

[0122] It should be noted that IT operation and maintenance data is mainly represented in the form of work ticket data. Those skilled in the art will understand that other data forms such as logs are also applicable and only need to be adjusted accordingly.

[0123] S102, performing data enhancement on the text dataset at the character level, word level, and sentence level to generate a data-enhanced text dataset.

[0124] In this step, the minority sample category data in the category-imbalanced text dataset are enhanced at three granularities: character level, word level, and sentence level, and the newly added sample data are added to the text dataset to obtain a text dataset with a relatively sufficient and balanced number of data samples.

[0125] S103: performing data preprocessing on the data-enhanced text dataset to obtain a preprocessed text dataset.

[0126] Optionally, data preprocessing methods may include but are not limited to text standardization, data cleaning, word segmentation, text tag replacement, etc.

[0127] In this step, after obtaining the data-enhanced text dataset, the text dataset is preprocessed to obtain a preprocessed text dataset, so as to remove redundant features in the text dataset and improve the expression ability of key features of IT operation and maintenance.

[0128] S104, converting the preprocessed text dataset into a text vector representation to obtain a text vector dataset.

[0129] It should be noted that the text vector representation can be a discrete representation, which can be obtained through algorithms such as the bag-of-words model, one-hot encoding, and TF-IDF. Those skilled in the art will appreciate that the text vector representation can also be a distributed representation, which can be obtained through algorithms such as Word2Vec and BERT, and this invention does not specifically limit this.

[0130] In this step, text vector representation refers to representing the semantics of text through numerical vectors. After obtaining the preprocessed text dataset, the text dataset is converted into a vector or matrix form to generate a text vector dataset so that the sample data of the text vector dataset conforms to the input data format of the classification model.

[0131] S105: Perform feature extraction on the text vector dataset to obtain a text enhancement dataset.

[0132] In this step, multi-dimensional text feature extraction is achieved by combining character-level and word-level text vectors, which can alleviate the feature sparsity problem of short text samples.

[0133] In some embodiments of the present invention, in step S101, the step of obtaining a plurality of IT operation and maintenance data as a text dataset mainly includes:

[0134] S1011, obtaining a number of IT operation and maintenance data and corresponding text categories as a text dataset.

[0135] Optionally, in step S101 , a number of IT operation and maintenance data and corresponding text categories are obtained by monitoring the operating status of devices such as servers, network devices, databases, middleware, or application systems.

[0136] Among them, the network device can be a bridge, or a repeater, switch, hub, router, gateway, etc. The database can be a relational database such as SQLite, MySQL, or a non-relational database such as MongoDB, HBase, etc. The middleware can be a message-oriented middleware such as RocketMQ, Kafka, or an object request broker middleware, remote procedure call middleware, etc. It will be understood by those skilled in the art that the operation and maintenance data of one or more devices or application systems can be obtained as IT operation and maintenance data according to actual circumstances, and the present invention does not specifically limit this.

[0137] It should be noted that the text category refers to the label data of IT operation and maintenance data and can be adjusted appropriately based on the actual text classification problem. Optionally, the text classification problem in the embodiments of the present invention can be a binary classification problem or other classification problems such as multi-classification or multi-label classification. The present invention does not specifically limit the type of text classification problem.

[0138] Exemplarily, when the text classification problem is a binary classification problem, the text category includes any one of the text type and the non-text type, or the text type includes any one of the operation and maintenance data type and the non-operation and maintenance data type; when the text classification problem is a multi-classification problem, the text category can be the specific text type of the data, such as IT operation and maintenance data of the middleware, IT operation and maintenance data of the server, etc.

[0139] In some embodiments of the present invention, referring to FIG. 2 , which is a flowchart of data enhancement provided by the present invention, in step S102 , data enhancement is performed on a text dataset at the character level, word level, and sentence level to generate a data-enhanced text dataset. The process mainly includes the following steps:

[0140] S1021: Obtain IT operation and maintenance data belonging to a minority sample category in the text dataset as text sentences to be enhanced.

[0141] It should be noted that each text sentence to be enhanced is composed of several text words.

[0142] In this step, the IT operation and maintenance data belonging to the minority sample category in the text dataset is obtained, and the IT operation and maintenance data belonging to the minority sample category is used as the data object for data enhancement.

[0143] Optionally, the few-sample category can be a category in the text dataset where the amount of data is less than a preset value, or it can refer to a category in the text dataset where data is more difficult to obtain. The few-sample category can be set according to actual conditions, and the present invention does not make any specific limitations on this.

[0144] S1022: Randomly select multiple text words from multiple text sentences to be enhanced, perform data enhancement on the multiple text sentences to be enhanced based on the multiple randomly selected text words, and generate multiple first augmented text sentences.

[0145] In this step, at the character level, data enhancement is performed on multiple text sentences to be enhanced based on the data enhancement method of text editing distance and keyboard adjacent input to obtain multiple first augmented text sentences, so as to improve the generalization ability of the classification model for spelling errors in words.

[0146] S1023: Obtain a keyword table of the text data set, perform data enhancement on a plurality of text sentences to be enhanced according to the keyword table, and generate a plurality of second augmented text sentences.

[0147] It should be noted that the keyword table includes multiple keywords, and keywords can be understood as text words with key characteristics in the field of IT operation and maintenance.

[0148] In this step, at the word level, the traditional Easy Data Augmentation (EDA) method is improved based on the TextRank algorithm. This method proposes an improved text data augmentation method. This method obtains a keyword table for the text dataset and then performs data augmentation on multiple text sentences to be augmented based on the keyword table, resulting in multiple second augmented text sentences. Compared to the traditional EDA method, the improved text data augmentation method proposed in this embodiment of the present invention places greater emphasis on extracting important feature words from the dataset.

[0149] S1024: Perform data enhancement on the multiple text sentences to be enhanced by back-translation to generate multiple third augmented text sentences.

[0150] It should be noted that back translation refers to the process of translating the original sentence into other languages ​​and then translating it back to the original language.

[0151] In this step, at the sentence level, data enhancement is performed on multiple text sentences to be enhanced through back translation to obtain multiple third augmented text sentences, so that the text sentences obtained after data enhancement have greater diversity.

[0152] S1025 , adding the plurality of first augmented text sentences, the second augmented text sentences, and the third augmented text sentences to the text dataset to generate a data-enhanced text dataset.

[0153] In this step, the newly added sample data obtained through the previous steps are added to the text dataset, thereby obtaining a text dataset with a relatively sufficient and balanced number of data samples.

[0154] In some embodiments of the present invention, referring to FIG. 3 , which is a flowchart of character-level data enhancement provided by the present invention, in step S1022 , a plurality of text words are randomly selected from a plurality of text sentences to be enhanced, and data enhancement is performed on the plurality of text sentences to be enhanced based on the plurality of randomly selected text words. The process of generating a plurality of first augmented text sentences mainly includes the following steps:

[0155] A1: Randomly select multiple text words from multiple text sentences to be enhanced as the first sample words.

[0156] It should be noted that the number of the first sample words may be determined according to actual conditions, and the embodiment of the present invention does not impose any specific limitation on the number of the first sample words.

[0157] A2: Perform data enhancement on the multiple first sample words by using an edit distance replacement method and a keyboard accidental touch replacement method to generate candidate words for the multiple first sample words.

[0158] It should be noted that the edit distance replacement method refers to a method of generating a new word candidate set by calculating the edit distance (ED) between two words, and the keyboard accidental touch replacement method refers to a method of randomly replacing a letter in a word with any adjacent letter of the letter on the keyboard to generate a new word.

[0159] In this step, to address misspellings caused by similar characters during data entry, the edit distance replacement method is used to enhance the data of multiple first sample words, thereby generating a first candidate set. Furthermore, to address misspellings caused by accidental keyboard touches during data entry, the accidental keyboard touch replacement method is used to enhance the data of multiple first sample words, thereby generating a second candidate set. The first and second candidate sets are combined to form candidate words for the multiple first sample words.

[0160] A3: In each text sentence to be enhanced, the first sample word is replaced with a candidate word of the first sample word to generate a first augmented text sentence corresponding to each text sentence to be enhanced, thereby obtaining multiple first augmented text sentences.

[0161] In this step, multiple first augmented text sentences are generated by replacing the first sample word in each text sentence to be enhanced with a candidate word of the first sample word, thereby achieving data enhancement of the text dataset at the character level.

[0162] For example, in step A2, any text sentence to be enhanced includes M text words. First, N text words are selected from the M text words as first sample words. Data augmentation is performed on the N first sample words using the edit distance replacement method to generate a first candidate set. Simultaneously, data augmentation is performed on the N first sample words using the keyboard accidental touch replacement method to generate a second candidate set. Then, N candidate words are selected from the total candidate set consisting of the first and second candidate sets to replace the N first sample words in the text sentence to be enhanced, thereby generating a new text sentence, i.e., the first augmented text sentence.

[0163] In some embodiments of the present invention, step A2 is mainly divided into two parallel processing parts, one of which is to implement data enhancement using the edit distance replacement method, and the other is to implement data enhancement using the keyboard accidental touch replacement method. In step A2, the data enhancement of multiple first sample words by the edit distance replacement method and the keyboard accidental touch replacement method is performed to generate candidate words for multiple first sample words. The implementation process mainly includes the following steps:

[0164] A21, obtain the IT operation and maintenance word database.

[0165] It should be noted that the IT operation and maintenance word database is a pre-built database, and the IT operation and maintenance word database includes a plurality of IT operation and maintenance words. It is understandable that the IT operation and maintenance words refer to words that constitute IT operation and maintenance data.

[0166] A22, obtaining the edit distance between each first sample word and each IT operation and maintenance word.

[0167] It should be noted that the edit distance is a measure of the similarity between two character sequences. Character sequences can be understood as words, and the edit distance is inversely proportional to the similarity between the two words. That is, the larger the edit distance, the smaller the similarity between the two words; the smaller the edit distance, the greater the similarity between the two words.

[0168] Specifically, the edit distance between two words can be the minimum number of character edit operations required to convert the two words. Based on this, first, the number of character edit operations required to convert the first sample word into the IT operation and maintenance word is obtained. Then, based on the number of character edit operations, the edit distance between each first sample word and each IT operation and maintenance word is calculated.

[0169] Optionally, the character editing operation count is the minimum number of execution times of the character editing operation, and the character editing operation may include but is not limited to at least one of inserting a character, deleting a character, and changing a character.

[0170] Exemplarily, a first sample word abc and IT operation and maintenance words abcd, ab, and abd are obtained. When the character editing operation is inserting a character, the first sample word abc is converted into the IT operation and maintenance word abcd by inserting the character d into the first sample word abc; when the character editing operation is deleting a character, the first sample word abc is converted into the IT operation and maintenance word ab by deleting the character c in the first sample word abc; when the character editing operation is changing a character, the first sample word abc is converted into the IT operation and maintenance word abd by changing the character c in the first sample word abc to the character d.

[0171] More specifically, the edit distance between any first sample word and any IT operation and maintenance word satisfies the following formula:

[0172] Among them, lev a,b (i, j) represents the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b; a i represents the i-th character of the first sample word a, b j Indicates the j-th character of the IT operation and maintenance word b.

[0173] In the formula, when a i ≠b j When taking lev a,b (i-1, j)+1, lev a,b (i, j-1)+1, lev a,b The minimum value among (i-1, j-1)+1 is taken as the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b; when lev a,b (i-1, j)+1, lev a,b (i, j-1)+1, lev a,b When the minimum value in (i-1, j-1)+1 is 0, max(i, j) represents the boundary condition of the edit distance, and the maximum value between i and j is taken as the edit distance between the i-th character of the first sample word a and the j-th character of the IT operation and maintenance word b.

[0174] A23 : For each first sample word, obtain an IT operation and maintenance word whose edit distance is less than a first value as a first candidate word of the first sample word.

[0175] It should be noted that the value of the first numerical value can be set according to actual conditions, and the embodiment of the present invention does not specifically limit the first numerical value.

[0176] In this step, first, an edit distance threshold, i.e., a first value, is pre-set, and then an IT operation and maintenance word with an edit distance less than the first value is selected as a first candidate word for the first sample word.

[0177] A24, randomly selecting a letter from each first sample word as a letter to be replaced, and determining a candidate letter for the letter to be replaced in each first sample word.

[0178] Specifically, after obtaining the letter to be replaced, the position information of the letter to be replaced on the keyboard is first obtained. The keyboard is a 26-key keyboard. Then, based on the position information of the letter to be replaced on the keyboard, multiple letters surrounding the letter to be replaced on the keyboard are selected as candidate letters for the letter to be replaced. It is understood that the multiple letters surrounding the letter to be replaced are letters that are easily accidentally touched when inputting the letter to be replaced on a 26-key keyboard.

[0179] For example, referring to FIG4 , which is a schematic diagram of a 26-key keyboard provided by the present invention, when the letter to be replaced is F, the letters R, T, D, G, C, and V surrounding F are easily mistakenly touched when inputting F on the 26-key keyboard. Therefore, based on the position information of F on the 26-key keyboard, the letters R, T, D, G, C, and V are selected as candidates for F.

[0180] For another example, referring to FIG4 , when the letter to be replaced is H, the letters Y, U, I, J, N, B, and G surrounding H are letters that are easily accidentally touched when inputting H using a 26-key keyboard. Therefore, based on the position information of H on the 26-key keyboard, the letters Y, U, I, J, N, B, and G are selected as candidate letters for H.

[0181] A25 , replacing the letters to be replaced in each first sample word with candidate letters to be replaced, to generate a second candidate word for each first sample word.

[0182] In this step, a new word, ie, a second candidate word of each first sample word, is generated by replacing the to-be-replaced letter in each first sample word with any candidate letter of the to-be-replaced letter.

[0183] A26 , obtaining first candidate words and second candidate words of the plurality of first sample words as candidate words of the plurality of first sample words.

[0184] In this step, new words generated by the edit distance replacement method and the keyboard accidental touch replacement method are merged as candidate words of the plurality of first sample words.

[0185] In some embodiments of the present invention, in step S1023, the process of obtaining the keyword table of the text dataset mainly includes the following steps:

[0186] The text dataset is processed using the TextRank algorithm to generate a keyword table for the text dataset.

[0187] It should be noted that the TextRank algorithm is a graph-based ranking algorithm for keyword extraction and document summarization. For a given text, it uses the co-occurrence information or semantics between words within the text to extract the keywords or keyword groups of the text, and uses an extractive automatic summarization method to extract the key sentences of the text.

[0188] In this step, the keywords of the text dataset are first extracted using the TextRank algorithm to obtain a keyword table of the text dataset, which consists of multiple keywords.

[0189] In some embodiments of the present invention, in step S1023, performing data enhancement on the plurality of text sentences to be enhanced according to the keyword table to generate a plurality of second augmented text sentences includes at least one of the following:

[0190] B1, for each text sentence to be enhanced, randomly select a text word belonging to the keyword table from the text sentence to be enhanced as a second sample word, replace the second sample word with a synonym of the second sample word, and generate a second augmented text sentence corresponding to the text sentence to be enhanced.

[0191] It should be noted that the number of the second sample words may be determined according to actual conditions, and the present invention does not impose any specific limitation thereto.

[0192] In this step, data augmentation is performed on each text sentence to be augmented using Synonym Replacement (SR) and a keyword table. Specifically, for each text sentence to be augmented, a specified number of text words from the keyword table are randomly selected as second sample words. Synonyms of the second sample words are randomly selected to replace the second sample words, thereby generating a second augmented text sentence corresponding to the text sentence to be augmented.

[0193] B2, for each text sentence to be enhanced, traverse the text words belonging to the keyword table in the text sentence to be enhanced, randomly insert synonyms of the text words belonging to the keyword table into the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced.

[0194] In this step, data augmentation is performed on each text sentence to be augmented using Random Replacement (RI) and a keyword table. Specifically, for each text sentence to be augmented, the method traverses the text words in the keyword table in the text sentence to be augmented, randomly selects a text word from the keyword table from the text sentence to be augmented, and inserts a synonym of the randomly selected text word into any position in the text sentence to be augmented, thereby generating a second augmented text sentence corresponding to the text sentence to be augmented.

[0195] Optionally, in this embodiment of the present invention, step B2 is repeated to implement multiple data augmentations. Those skilled in the art will appreciate that the specific method of random insertion can be varied according to actual circumstances to achieve repeated execution. Furthermore, it should be noted that the number of second augmented text sentences generated by this step is equal to the number of times this step is repeated.

[0196] Exemplarily, repeated execution can be in a single operation, inserting any synonym of a text word belonging to the keyword table in the text sentence to be enhanced into a specified position of the text sentence to be enhanced, repeating this operation until all text words of the text sentence to be enhanced are traversed, and then generating multiple second augmented text sentences.

[0197] As another example, repeated execution can also be in a single operation, in which a synonym of a text word belonging to the keyword table in the text sentence to be enhanced is inserted into a specified position of the text sentence to be enhanced, and this operation is repeated until all synonyms of the text word are traversed, thereby generating multiple second augmented text sentences.

[0198] As another example, repeated execution can also be in a single operation, where any synonym of a text word belonging to the keyword table in the text sentence to be enhanced is inserted into one of the specified positions of the text sentence to be enhanced, and this operation is repeated until all the specified positions in the text sentence to be enhanced are traversed, thereby generating multiple second augmented text sentences.

[0199] B3: For each text sentence to be enhanced, traverse the text words in the text sentence to be enhanced, swap the positions of two text words in the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced.

[0200] In this step, data augmentation is performed on each sentence to be augmented using the Random Swap (RS) method. Specifically, for each sentence to be augmented, two words are randomly selected and swapped. This operation is repeated a specified number of times to generate multiple second augmented sentences corresponding to the sentence to be augmented.

[0201] It should be noted that the number of second augmented text sentences generated through this step is equal to the number of times this step is repeated.

[0202] B4, obtaining word deletion probability, for each text sentence to be enhanced, deleting text words that do not belong to the keyword table from the text sentence to be enhanced according to the word deletion probability, and generating a second augmented text sentence corresponding to the text sentence to be enhanced.

[0203] It should be noted that the word deletion probability may be determined according to actual conditions, and the present invention does not impose any specific limitation thereto.

[0204] In this step, data enhancement is performed on each text sentence to be enhanced using the Random Deletion (RD) method and the keyword table. Specifically, for each text sentence to be enhanced, words that do not belong to the keyword table are deleted from the text sentence to be enhanced based on the word deletion probability, thereby generating a second augmented text sentence corresponding to the text sentence to be enhanced.

[0205] In some embodiments of the present invention, referring to FIG. 5 , which is a flowchart of sentence-level data enhancement provided by the present invention, in step S1024, the process of performing data enhancement on multiple text sentences to be enhanced by back-translation to generate multiple third augmented text sentences mainly includes the following steps:

[0206] C1, determine the original language and translated language of each text sentence to be enhanced.

[0207] It should be noted that the original language and the translated language can be Chinese, or other languages ​​such as English, French, German, etc., and the present invention does not make specific limitations on this. However, it should be emphasized that the original language includes one language and the translated language includes at least one language.

[0208] Optionally, the original language is English, and the translated language includes at least one of Chinese and French.

[0209] C2: Translate each text sentence to be enhanced represented by the initial language into the translated language to obtain each text sentence to be enhanced represented by the translated language.

[0210] This step is the forward translation process in the back translation method. The forward translation process refers to translating text sentences into other languages.

[0211] C3, translating each text sentence to be enhanced represented by the translated language into the original language to obtain a third augmented text sentence corresponding to each text sentence to be enhanced, and then generating multiple third augmented text sentences.

[0212] This step is the reverse translation process in the back translation method. The reverse translation process refers to translating text sentences in other languages ​​back into the original language.

[0213] The following example illustrates the implementation process of data enhancement at the character level, word level, and sentence level provided by the embodiment of the present invention.

[0214] 6 , which is a schematic diagram of an application of data enhancement provided by the present invention, assumes that the IT operation and maintenance data “It's getting very slow process. Some time it's showing some error massge.” in a text data set is obtained as a text sentence to be enhanced, and data enhancement is performed on the text sentence to be enhanced at the character level (Char-level), word level (Word-level), and sentence level (Sentence-level).

[0215] Specifically, at the character level, the data enhancement method based on text editing distance and keyboard adjacent input is used to enhance the text sentence to be enhanced, and the character "t" in "getting" is replaced with the character "g", thereby obtaining the first augmented text sentence of the text sentence to be enhanced: "It's getting a very slow process. Some time it's showing some error massge."

[0216] At the word level, four methods are used to perform data augmentation on the text sentence to be augmented, and four second augmented text sentences of the text sentence to be augmented can be obtained. Assuming that the words "error" and "very" are both in the keyword table, we have:

[0217] Data augmentation is performed on the text sentence to be augmented by using a synonym replacement method and a keyword table, replacing the word "error" with its synonym "wrong", thereby obtaining a second augmented text sentence of the text sentence to be augmented: "It's getting very slow process. Some time it's showing some wrong massge."

[0218] Data augmentation is performed on the text sentence to be augmented by using a random insertion method and a keyword table. The synonym of the word "very" "much" is inserted into any position of the text sentence to be augmented, thereby obtaining a second augmented text sentence of the text sentence to be augmented: "It's getting very much slow process. Some time it's showing some error massge.";

[0219] Data enhancement is performed on the text sentence to be enhanced by using a random swap method and a keyword table, swapping the positions of the word "slow" and the word "process", thereby obtaining a second augmented text sentence of the text sentence to be enhanced: "It's getting very process slow. Some time it's showing some error massge.";

[0220] Data enhancement is performed on the text sentence to be enhanced by using a random deletion method and a keyword table, and the word "slow" is deleted, thereby obtaining a second augmented text sentence of the text sentence to be enhanced: "It's getting very process. Some time it's showing some error massge."

[0221] At the sentence level, data enhancement is performed on the text sentence to be enhanced through back translation, and a third augmented text sentence "This is a very slow process. Sometimes some error messages are displayed." of the text sentence to be enhanced is generated through forward translation and backward translation.

[0222] By performing data augmentation on multiple text sentences to be augmented in a text dataset at three granularities: character level, word level, and sentence level, multiple new augmented sentences can be generated. By adding the new augmented sentences to the original text dataset, a data-augmented text dataset can be formed. This can alleviate the sample category imbalance problem in the original dataset and increase the data volume of the text dataset, which is conducive to improving the classification accuracy of the text classification model.

[0223] In some embodiments of the present invention, referring to FIG. 7 , which is a flowchart of data set preprocessing provided by the present invention, in step S103 , data preprocessing is performed on the data-enhanced text data set to obtain the preprocessed text data set, which may include but is not limited to the following steps:

[0224] S1031, performing text standardization and text cleaning on the data-enhanced text dataset to obtain a text-cleaned text dataset.

[0225] Specifically, the text normalization process may include, but is not limited to, converting characters in the data-enhanced text dataset to half-width characters and converting English letters in the data-enhanced text dataset to lowercase. The text cleaning process may include, but is not limited to, removing invalid characters in the data-enhanced text dataset, such as non-text content and punctuation marks.

[0226] S1032: Perform string replacement processing on the text dataset after text cleaning to obtain a text dataset after replacement processing.

[0227] In this step, string replacement processing is performed on the text dataset after text cleaning to enhance the expression of feature words in the IT operation and maintenance field, thereby obtaining a text dataset after replacement processing.

[0228] S1033 , performing word segmentation processing on the text dataset after the replacement processing to obtain word-level lemmas and character-level lemmas of the text dataset.

[0229] Specifically, we use the NLTK natural language processing tool to perform word segmentation on the replaced text dataset to obtain word-level and character-level tokens. As you can understand, NLTK stands for Natural Language Toolkit. NLTK is a toolkit in the field of natural language processing. It includes numerous libraries and datasets that can be used to complete various natural language processing tasks.

[0230] S1034 , performing word cleaning and normalization processing on the word-level tokens of the text dataset to obtain a preprocessed text dataset.

[0231] Specifically, the word cleaning process may include, but is not limited to, performing word cleaning on word-level lemmas in a text dataset. The word cleaning process may include, but is not limited to, at least one of: stop word removal, part-of-speech tagging, lemmatization, and concatenation correction. The normalization process may include, but is not limited to, restoring the part of speech and tense of word-level lemmas in a text dataset.

[0232] In some embodiments of the present invention, related data enhancement technologies lack the ability to capture and utilize key words in the IT operations field, resulting in a large number of redundant features in the dataset. To address this issue, embodiments of the present invention propose a string replacement step. Specifically, in step S1032, string replacement is performed on the cleaned text dataset to obtain the replaced text dataset, which may include but is not limited to the following steps:

[0233] First, rule recognition is performed on the text dataset after text cleaning to obtain the IT operation and maintenance feature words of the text dataset.

[0234] In this step, regular expressions are used to identify rules based on the characteristics of different types of IT operation and maintenance feature words in the cleaned text dataset to obtain IT operation and maintenance feature words in the text dataset. Optionally, IT operation and maintenance feature words may include, but are not limited to, at least one of date, time, percentage, IP address, memory address, memory capacity information (i.e., memory / file size), website address, phone number, network status code, and error code.

[0235] It should be noted that the matching rules of the regular expression can be set according to the actual situation, and the rule recognition of the text data set after text cleaning can be completed through the matching rules of the regular expression. The present invention does not make specific limitations on this.

[0236] Exemplarily, when the IT operation and maintenance feature word is a percentage, the regular expression used to identify the percentage type is "\d{1,3}(?:\.\d{1,2})?%", and the matching rules of the regular expression may include but are not limited to the integer part (i.e. "\d{1,3}"), the decimal point (i.e. "\."), the decimal part (i.e. "\d{1,2}"), and the percent sign (i.e. "%"), where "\d{1,3}" means that the integer part contains 1-3 digits, "\d{1,2}" means that the decimal part contains 1-2 digits, ":" means a separator, and "?" means that both the decimal part and the percent sign are optional parts.

[0237] As another example, when the IT operation and maintenance feature word is date, the regular expression used to identify the date type is "\d{4}-\d{1,2}-\d{1,2}". The matching rules of this regular expression may include but are not limited to year (i.e., "\d{4}"), month (i.e., "-\d{1,2}") and day (i.e., "-\d{1,2}"), where "\d{4}" means that the year is represented by four digits, and "-\d{1,2}" means that the month or day is represented by 1-2 digits.

[0238] As another example, when the IT operation and maintenance feature word is time, the regular expression used to identify the time type is "^\d{1,2}:\d{1,2}(:\d{1,2}(.\d{1,3})?)?$", and the matching rules of the regular expression may include but are not limited to using "^" and "$" to match the entire string from the beginning to the end, hours (i.e., "\d{1,2}"), minutes (i.e., "\d{1,2}"), seconds (i.e., "\d{1,2}"), and milliseconds (i.e., "\d{1,3}"), where ":" represents a separator, "?" represents that seconds and milliseconds are optional parts, "\d{1,2}" represents that hours, minutes or seconds are represented by 1-2 digits, and "\d{1,3}" represents that milliseconds are represented by 1-3 digits.

[0239] Then, the identifier of the IT operation and maintenance feature word is obtained, and the IT operation and maintenance feature word in the text data set is replaced with the identifier of the IT operation and maintenance feature word, thereby obtaining the text data set after replacement processing.

[0240] In this step, the identifiers corresponding to the IT operation and maintenance feature words obtained through rule identification are used to replace the IT operation and maintenance feature words, generate multiple new character strings, and then obtain a text dataset after replacement processing.

[0241] For example, referring to FIG8 , FIG8 is a schematic diagram of the identifier provided by the present invention, and the “example” in FIG8 represents certain data in the text data set after text cleaning, the “meaning” represents the IT operation and maintenance feature words corresponding to these data, and the “identifier” represents the identifier corresponding to the IT operation and maintenance feature words. For example, for the data “22 / 11 / 2022” in the text data set after text cleaning, rule recognition is performed on “22 / 11 / 2022”, and the IT operation and maintenance feature word is obtained as “date”, and the identifier corresponding to “date” is obtained. <date>",use" <date>" to replace "date". For another example, for the data "0xc000000f" in the text data set after text cleaning, rule recognition of "0xc000000f" can be performed, and the IT operation and maintenance feature word "error code" can be obtained. The identifier corresponding to "error code" is obtained. <error>",use" <error>" to replace "error code".

[0242] In some embodiments of the present invention, in step S104, the process of converting the preprocessed text dataset into a text vector representation to obtain the text vector dataset mainly includes the following steps:

[0243] S1041, encode the character-level word units of the preprocessed text dataset to generate word vector features of the text dataset.

[0244] This step uses one-hot encoding to encode the character-level tokens of the preprocessed text dataset to generate word vector features of the text dataset.

[0245] It should be noted that one-hot encoding refers to the use of 0s and 1s to represent some parameters, using a multi-bit state register to encode multiple states. Each state has its own independent register bit, and at any time, only one of the bits is valid. That is, only one bit is 1, and the rest are 0.

[0246] Optionally, in other embodiments of the present invention, other methods such as bag-of-words model, TF-IDF, etc. may be used to encode the character-level tokens of the preprocessed text dataset.

[0247] S1042, encode the word-level tokens of the preprocessed text dataset to generate word vector features of the text dataset.

[0248] In this step, the pre-trained word vector model is used to encode the word-level tokens of the preprocessed text dataset to generate the word vector features of the text dataset.

[0249] It should be noted that the word vector model can be a Word2Vec model, or other word vector models such as CBOW, Skip-gram, etc., and the present invention does not make specific limitations on this.

[0250] Exemplarily, a Word2Vec model is used as a word vector model, and a pre-trained Word2Vec model is used to encode word-level tokens of a preprocessed text dataset to generate a Word2Vec vector of the text dataset as a word vector feature.

[0251] S1043, constructing a text vector dataset using the character vector features and word vector features of the text dataset.

[0252] In some embodiments of the present invention, in step S105, the process of performing feature extraction on the text vector dataset to obtain the text enhanced dataset may include but is not limited to the following steps:

[0253] S1051, concatenate the character vector features and word vector features in the text vector dataset and perform text length normalization to obtain a text enhancement dataset.

[0254] It should be noted that text length normalization is based on the length of all sample data in the dataset, setting it to a reasonable uniform length to facilitate input into the text classification model for classification. The length refers to the total number of tokens in each sample data.

[0255] In this step, first, the character vector features and word vector features in the text vector dataset are concatenated to obtain a concatenated vector feature of the text vector dataset. Then, the text length of the concatenated vector feature of the text vector dataset is unified to the second value.

[0256] It is understandable that the second value may be determined according to actual conditions, and the present invention does not impose any specific limitation on this.

[0257] Optionally, the text length of the concatenated vector features of the text vector dataset is unified to a second value by adopting a truncation and short-filling method. Specifically, when the length of the character vector features and word vector features of the text dataset is greater than a specified length, only the portion of the specified length is truncated from the character vector features and word vector features; when the length of the character vector features and word vector features of the text dataset is less than or equal to the specified length, the remaining position is filled with a zero vector.

[0258] The following example illustrates the principle of the data enhancement method proposed in an embodiment of the present invention. Referring to FIG9 , FIG9 is a schematic diagram of a data enhancement method for operation and maintenance data provided by the present invention. The principle of data enhancement is as follows:

[0259] The first step is to obtain some IT operation and maintenance data as a text dataset.

[0260] The second step is to perform data augmentation on the text dataset at the character level, word level, and sentence level.

[0261] The third step is to preprocess the text dataset after data enhancement. The preprocessing process includes: first, text standardization and text cleaning, such as full-width and half-width character conversion, uppercase and lowercase letter conversion, removal of non-text content, filtering of punctuation marks, etc.; then, string replacement processing is performed to replace the IT operation and maintenance feature words in the text dataset with corresponding identifiers; then, word segmentation is performed to obtain word-level lemmas and character-level lemmas of the text dataset, and word-level lemmas are cleaned and normalized, such as removing stop words, part-of-speech tagging, part-of-speech restoration, spelling correction, etc.

[0262] The fourth step is to perform text vector representation on the preprocessed text dataset, that is, to vectorize the word units through one-hot encoding, Word2Vec, etc. to obtain a text vector dataset.

[0263] The fifth step is to perform feature selection and combination on the text vector dataset and standardize the text length to obtain the text enhancement dataset and achieve data enhancement.

[0264] Next, referring to the accompanying drawings, a detailed description of the implementation steps of a text classification method for operation and maintenance data proposed in accordance with an embodiment of the present invention is provided. Referring to FIG10 , FIG10 is a flowchart of a text classification method for operation and maintenance data provided by the present invention. The text classification method may include but is not limited to the following steps:

[0265] S201, obtaining IT operation and maintenance data to be classified.

[0266] S202: Input the IT operation and maintenance data to be classified into a trained text classification model to perform text classification and obtain a text classification result.

[0267] It should be noted that the text classification model is obtained by pre-training the initial classification model using the text enhancement dataset, and the text enhancement dataset is obtained through the data enhancement method of the operation and maintenance data described above.

[0268] Optionally, the initial classification model may be a TextRCNN model, or other classification models for natural language processing such as TextCNN, TextRNN, etc., which is not specifically limited in the present invention.

[0269] Exemplarily, the TextRCNN model is used as the initial classification model, and a text enhancement dataset is obtained through the data enhancement method of operation and maintenance data described above. The text enhancement dataset is used as the input of the TextRCNN model, and the TextRCNN model is trained using the text enhancement dataset. The trained TextRCNN model is used as the output of the text classification model, and the text classification model is used to perform text classification on the IT operation and maintenance data to be classified, thereby obtaining the text classification results.

[0270] In summary, the present invention provides a data enhancement method, a text classification method, and an electronic device for operation and maintenance data. The data enhancement method includes: first obtaining a number of IT operation and maintenance data as a text dataset, then performing data enhancement on the text dataset at the character level, word level, and sentence level to generate a data-enhanced text dataset, then performing data preprocessing on the data-enhanced text dataset to obtain a preprocessed text dataset, and converting the preprocessed text dataset into a text vector representation to obtain a text vector dataset, and finally performing feature extraction on the text vector dataset to obtain a text enhancement dataset, thereby realizing text enhancement processing of IT operation and maintenance data; the text classification method includes: obtaining IT operation and maintenance data to be classified, and performing text classification on the IT operation and maintenance data to be classified using a text classification model. Among them, the generation process of the text classification model includes: obtaining a text enhancement dataset through the data enhancement method described above, and then training the initial classification model using the text enhancement dataset to obtain a text classification model.

[0271] As for the data enhancement method, on the one hand, the present invention enhances IT operation and maintenance data at three different granularities: character level, word level, and sentence level, thereby increasing sentence diversity for the original data set and effectively alleviating the sample category imbalance problem in the original data set. On the other hand, the present invention can remove redundant features that have little impact on text classification in the original data set by adding string replacement processing during the preprocessing process, thereby improving the contribution of specific types of symbol representation to IT operation and maintenance data classification. On the other hand, the present invention realizes multi-dimensional text feature extraction by combining text vector representation and standardization at the character level and word level, thereby overcoming the feature sparsity problem of short text samples. The data enhancement method proposed in the present invention has higher inclusiveness for text data sets and is conducive to generating high-quality data sets.

[0272] Regarding the text classification method, the present invention uses the data set generated by the data enhancement method described above to train the text classification model, which can improve the generalization ability of the text classification model and thereby improve the classification accuracy of the text classification model in the operation and maintenance data text classification task.

[0273] In addition, an embodiment of the present invention further provides an electronic device, including:

[0274] at least one processor;

[0275] at least one memory for storing at least one program;

[0276] When the at least one program is executed by the at least one processor, the at least one processor implements the aforementioned data enhancement method for operation and maintenance data, or the aforementioned text classification method for operation and maintenance data.

[0277] The contents of the above method embodiments are all applicable to the present electronic device embodiment. The functions specifically implemented by the present electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0278] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

[0279] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.< / error> < / error> < / date> < / date>

Claims

1. A data enhancement method for operation and maintenance data, characterized in that: The following steps are involved: Obtain several IT operation and maintenance data as text data sets; Performing data enhancement on the text dataset at the character level, word level, and sentence level to generate a data-enhanced text dataset; Perform data preprocessing on the data-enhanced text dataset to obtain a preprocessed text dataset; Convert the preprocessed text dataset into a text vector representation to obtain a text vector dataset; Feature extraction is performed on the text vector dataset to obtain a text enhancement dataset.

2. The data enhancement method for operation and maintenance data according to claim 1, characterized in that: The step of performing data enhancement on the text dataset at the character level, word level and sentence level to generate a data enhanced text dataset includes: Obtain IT operation and maintenance data belonging to a small sample category in the text data set as text sentences to be enhanced, each of which is composed of a number of text words; Randomly selecting a plurality of text words from a plurality of text sentences to be enhanced, performing data enhancement on the plurality of text sentences to be enhanced according to the plurality of randomly selected text words, and generating a plurality of first augmented text sentences; Acquire a keyword table of the text data set, perform data enhancement on a plurality of text sentences to be enhanced according to the keyword table, and generate a plurality of second augmented text sentences; Performing data augmentation on a plurality of text sentences to be augmented by a back-translation method to generate a plurality of third augmented text sentences; A plurality of first augmented text sentences, second augmented text sentences and third augmented text sentences are added to the text dataset to generate a data-enhanced text dataset.

3. The data enhancement method for operation and maintenance data according to claim 2, characterized in that: The method of randomly selecting a plurality of text words from a plurality of text sentences to be enhanced, performing data enhancement on the plurality of text sentences to be enhanced according to the plurality of randomly selected text words, and generating a plurality of first augmented text sentences includes: Randomly select multiple text words from multiple text sentences to be enhanced as first sample words; Performing data enhancement on the plurality of first sample words by using an edit distance replacement method and a keyboard mis-touch replacement method to generate candidate words of the plurality of first sample words; In each text sentence to be enhanced, the first sample word is replaced with a candidate word of the first sample word to generate a first augmented text sentence corresponding to each text sentence to be enhanced, thereby obtaining a plurality of first augmented text sentences.

4. The data enhancement method for operation and maintenance data according to claim 3, characterized in that: The step of performing data enhancement on the plurality of text sentences to be enhanced according to the keyword table to generate a plurality of second augmented text sentences includes at least one of the following: For each text sentence to be enhanced, randomly select a text word in the keyword table from the text sentence to be enhanced as a second sample word, replace the second sample word with a synonym of the second sample word, and generate a second augmented text sentence corresponding to the text sentence to be enhanced; For each text sentence to be enhanced, traverse the text words in the keyword table in the text sentence to be enhanced, randomly insert synonyms of the text words in the keyword table into the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced; For each text sentence to be enhanced, traverse the text words in the text sentence to be enhanced, swap the positions of two text words in the text sentence to be enhanced, and generate a second augmented text sentence corresponding to the text sentence to be enhanced; A word deletion probability is obtained, and for each text sentence to be enhanced, text words that do not belong to the keyword table are deleted from the text sentence to be enhanced according to the word deletion probability, so as to generate a second augmented text sentence corresponding to the text sentence to be enhanced.

5. The data enhancement method for operation and maintenance data according to claim 1, characterized in that: The data preprocessing is performed on the data-enhanced text dataset to obtain a preprocessed text dataset, including: Perform text standardization and text cleaning on the data-enhanced text dataset to obtain a text-cleaned text dataset; Perform string replacement processing on the text data set after text cleaning to obtain a text data set after replacement processing; Performing word segmentation processing on the replaced text dataset to obtain word-level lemmas and character-level lemmas of the text dataset; The word-level word-units of the text data set are cleaned and normalized to obtain a preprocessed text data set.

6. The data enhancement method for operation and maintenance data according to claim 5, characterized in that: The step of performing string replacement processing on the text data set after text cleaning to obtain the text data set after replacement processing includes: Perform rule recognition on the text data set after text cleaning to obtain IT operation and maintenance feature words of the text data set; An identifier of the IT operation and maintenance feature word is obtained, and the IT operation and maintenance feature word in the text data set is replaced with the identifier of the IT operation and maintenance feature word, thereby obtaining a text data set after replacement processing.

7. The data enhancement method for operation and maintenance data according to claim 1, characterized in that: The preprocessed text data set is converted into a text vector representation to obtain a text vector data set, including: Encoding the character-level word units of the preprocessed text dataset to generate word vector features of the text dataset; Encoding the word-level tokens of the preprocessed text dataset to generate word vector features of the text dataset; A text vector dataset is constructed using the character vector features and word vector features of the text dataset.

8. The data enhancement method for operation and maintenance data according to claim 7, characterized in that: The step of extracting features from the text vector dataset to obtain a text enhancement dataset includes: The character vector features and word vector features in the text vector data set are concatenated and the text length is normalized to obtain a text enhancement data set.

9. A text classification method for operation and maintenance data, characterized in that: The following steps are involved: Obtain IT operation and maintenance data to be classified; Inputting the IT operation and maintenance data to be classified into a trained text classification model for text classification to obtain a text classification result; The text classification model is obtained by pre-training an initial classification model using a text enhancement data set, and the text enhancement data set is obtained by a data enhancement method for operation and maintenance data as described in any one of claims 1-8.

10. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a data enhancement method for operation and maintenance data as described in any one of claims 1 to 8, or implements a text classification method for operation and maintenance data as described in any one of claim 9.

Citation Information

Patent Citations

  • Text classification method oriented to healthy public opinion

    CN108829810A

  • Text classification method and device based on multi-channel deep learning model

    CN110851594A

  • Data processing method and device based on classification model, electronic equipment and medium

    CN111881983A

  • Question matching task-oriented data enhancement method

    CN115510863A

  • Substation alarm event identification method based on pre-training model and data enhancement

    CN116304041A