Operation and maintenance operation data processing method and device and electronic equipment
Through the classified operation model and named entity recognition model, the operation and maintenance operation data is processed, and the operation types, actions and objects are identified, which solves the problem that private domain operation and maintenance data cannot be parsed, and the streamlined storage and effective analysis of data is realized.
Patent Information
- Application Number
- CN202510095423.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology cannot effectively parse private domain operation and maintenance data, resulting in the impact of the time space and performance of data stored on the blockchain.
The operation and maintenance operation data are processed using the classified operation model and the named entity recognition model, identify operation types, actions and objects, and stored in the database.
It realizes accurate analysis and streamlined storage of private domain operation and maintenance data, solving data storage space and performance problems.
Smart Images

Figure CN119989054A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a method, device and electronic equipment for processing operation and maintenance data. Background Art
[0002] As the requirements for safe production security continue to increase, the authenticity of data in the field of operation and maintenance has become particularly important. In order to prevent data tampering, storing it on the blockchain is an effective means, but the data types are complex and the scale is too large. Direct storage will seriously affect the space and performance of the blockchain. Therefore, it is necessary to parse and simplify the data and store it in the blockchain. In the field of operation and maintenance, there are several commonly used data parsing methods and technologies: regular expression-based parsing, rule-based parsing, clustering and data mining-based parsing, natural language processing-based parsing, machine learning-based parsing, etc. However, for private domain operation and maintenance data, there is still a problem that private domain operation and maintenance data cannot be parsed.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present invention provide a method, device and electronic device for processing operation and maintenance data, so as to at least solve the technical problem in the related art that private domain operation and maintenance data cannot be parsed.
[0005] According to one aspect of an embodiment of the present invention, a method for processing operation and maintenance operation data is provided, comprising: obtaining target operation and maintenance operation data from a software transaction service interface, and reading out an operation subject and operation time from the target operation and maintenance operation data; classifying the target operation and maintenance operation data using a classification operation model, and determining an operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; when the operation type is a write operation type, extracting an operation action and an operation object from the target operation and maintenance operation data using a named entity recognition model; and storing the target operation and maintenance operation data, the operation subject, the operation time, the operation object, and the operation action correspondingly in a first database.
[0006] Furthermore, target operation and maintenance data is obtained from the software transaction service interface, including: obtaining initial operation and maintenance data from the software transaction service interface; determining a target conversion rule corresponding to the initial operation and maintenance data based on a data type corresponding to the initial operation and maintenance data, wherein the data type includes one of the following: a structured type and an unstructured type; and converting the initial operation and maintenance data using the target conversion rule to obtain target operation and maintenance data.
[0007] Furthermore, the initial operation and maintenance operation data is converted using target conversion rules to obtain target operation and maintenance operation data, including: converting the initial operation and maintenance operation data using target conversion rules to obtain converted operation and maintenance operation data; storing the converted operation and maintenance operation data in a second database; and reading the target operation and maintenance operation data from the second database based on a scheduled task.
[0008] Furthermore, the target operation and maintenance operation data is classified using the classification operation model to determine the operation type corresponding to the target operation and maintenance operation data, including: encoding the target operation and maintenance operation data to obtain a digital coding vector; inputting the digital coding vector into the classification operation model to obtain the operation type output by the classification operation model.
[0009] Furthermore, the named entity recognition model is used to extract operation actions and operation objects in the target operation and maintenance operation data, including: using the named entity model to perform entity recognition on the target operation and maintenance operation data to obtain the operation object; using the natural language processing library to perform part-of-speech recognition on the next word after the operation object to obtain the target part-of-speech of the next word; based on the target part-of-speech, determining the operation action from the target operation and maintenance operation data.
[0010] Furthermore, the target operation and maintenance data is subjected to entity recognition using a named entity model to obtain an operation object, including: using a named entity model to perform entity recognition on the target operation and maintenance data to determine a target entity in the target operation and maintenance data, and a target position of the target entity in the target operation and maintenance data; when the target position is the starting position of the target operation and maintenance data, determining that the target entity is an operation object.
[0011] Furthermore, based on the target part of speech, an operation action is determined from the target operation and maintenance data, including: when the target part of speech is a preset part of speech, determining a verb adjacent to the operation object as an operation action; when the target part of speech is not a preset part of speech, replacing the operation object and the next word to obtain first text data, adding a subject to the first text data to obtain second text data, and identifying the operation object and the operation action from the second text data.
[0012] Furthermore, the method also includes: obtaining an original data set; marking the read and write operation types of the original data set to obtain a first data set; dividing the first data set to obtain a first training set and a first test set; using the first training set to train the pre-trained language model multiple times, and storing the first training model whose first evaluation indicator during the multiple training processes is greater than the first preset indicator, wherein the first evaluation indicator includes one of the following: recall rate and accuracy; obtaining the first training model corresponding to the maximum first evaluation indicator from the stored first training models to obtain a classification operation model.
[0013] Furthermore, the method also includes: performing entity labeling on the original data set to obtain a second data set; dividing the second data set to obtain a second training set and a second test set; using the second training set to train the pre-trained language model multiple times, and storing the second training model whose second evaluation index during the multiple training processes is greater than the second preset index, wherein the second evaluation index includes one of the following: recall rate and accuracy; obtaining the second training model corresponding to the maximum second evaluation index from the stored second training models to obtain a named entity recognition model.
[0014] Furthermore, the method also includes: in response to receiving a query instruction sent by the client, determining a data source and a time period corresponding to the query instruction; querying in the first database based on the data source and the time period to obtain a query result; and feeding back the query result to the client.
[0015] According to another aspect of an embodiment of the present invention, there is also provided a device for processing operation and maintenance operation data, including: an acquisition module, used to acquire target operation and maintenance operation data from a software transaction service interface, and read out the operation subject and operation time from the target operation and maintenance operation data; a classification module, used to classify the target operation and maintenance operation data using a classification operation model, and determine the operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; an extraction module, used to extract the operation action and operation object in the target operation and maintenance operation data using a named entity recognition model when the operation type is a write operation type; and a storage module, used to store the target operation and maintenance operation data, operation subject, operation time, operation object and operation action correspondingly in a first database.
[0016] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the method in each embodiment of the present invention is executed when the program is running.
[0017] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present invention.
[0018] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method in each embodiment of the present invention is implemented.
[0019] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present invention is implemented.
[0020] According to another aspect of the embodiments of the present invention, a computer program is further provided. When the computer program is executed by a processor, the methods in the embodiments of the present invention are implemented.
[0021] In an embodiment of the present invention, the target operation and maintenance operation data from the software transaction service interface is obtained, and the operation subject and operation time are read from the target operation and maintenance operation data; the target operation and maintenance operation data are classified by using a classification operation model to determine the operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; when the operation type is a write operation type, the operation action and operation object in the target operation and maintenance operation data are extracted by using a named entity recognition model; the target operation and maintenance operation data, the operation subject, the operation time, the operation object and the operation action are stored in the first database in correspondence. It is easy to notice that by using the classification operation model, the write operation type in the target operation and maintenance operation data can be accurately identified, so as to further analyze and process the write operation data in a targeted manner. Further, the named entity recognition model can be used to accurately extract the operation action and operation object from the write operation data, so as to achieve the purpose of being able to parse the private domain operation and maintenance data, thereby achieving the technical effect of being able to parse the private domain operation and maintenance data, and thus solving the technical problem that the private domain operation and maintenance data cannot be parsed in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0023] Figure 1 is a flow chart of a method for processing operation and maintenance data according to an embodiment of the present invention;
[0024] Figure 2 is a schematic diagram of a training method for a named entity recognition model according to an embodiment of the present invention;
[0025] Figure 3 is a flowchart of an optional key operation and maintenance operation extraction program execution method according to an embodiment of the present invention;
[0026] Figure 4 is a flow chart of an optional hypertext transfer protocol calling method according to an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of an optional key operation and maintenance operation extraction program architecture according to an embodiment of the present invention;
[0028] Figure 6 is a schematic diagram of a device for processing operation and maintenance data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] According to an embodiment of the present invention, an embodiment of a method for processing operation and maintenance data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0032] Figure 1 is a flow chart of a method for processing operation and maintenance data according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:
[0033] Step S102, obtaining target operation and maintenance data from the software transaction service interface, and reading the operation subject and operation time from the target operation and maintenance operation data.
[0034] The above-mentioned software transaction service interface usually refers to the application programming interface (Application Programming Interface, API) designed in the software system to realize the transaction or service function. In the embodiment of the present invention, it can be the transaction service interface (Service API of TravelSky, referred to as SAT) of the company's internal key basic software. It is an interface used within the company to process transaction requests and provide services. The above-mentioned operation and maintenance operation data can be any type of private domain operation and maintenance data. The specific operation and maintenance data content is not limited in this embodiment and can be set according to actual usage requirements. Among them, private domain operation and maintenance data refers to operation and maintenance data generated and managed within a specific organization or enterprise. This type of data usually involves the operating status, operation subject, operation time, operation records, monitoring information, log files, configuration changes, performance indicators, error reports, etc. of the internal system of the organization.
[0035] In an optional embodiment, in order to solve the technical problem that private domain operation and maintenance data cannot be parsed in the prior art, an embodiment of the present invention discloses a method for processing operation and maintenance operation data. First, the target operation and maintenance operation data from the software transaction service interface can be obtained through the internal system, and the operation subject and operation time can be read from the target operation and maintenance operation data.
[0036] It should be noted that, when the target operation and maintenance data is obtained, the target operation and maintenance data can also be preprocessed to obtain the processed target operation and maintenance data. Specifically, the preprocessing method includes: using specific processing methods for different types of operation and maintenance data of the software SAT, and converting semi-structured and unstructured data into text data that approximates natural language.
[0037] Step S104: classify the target operation and maintenance operation data using the classification operation model to determine the operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type.
[0038] The above classification operation model is a machine learning model designed to identify and distinguish read and write operations in operation and maintenance operation data, and is trained in advance by technical personnel. The model here is mainly based on deep learning technology, especially the Bidirectional Encoder Representations from Transformers (BERT) model is fine-tuned to adapt to specific read and write operation classification requirements.
[0039] In an optional embodiment, when the target operation and maintenance data is obtained, the acquired target operation and maintenance operation data can be input into a classification operation model trained in advance, and the target operation and maintenance operation data can be classified by the classification operation model to determine the operation type corresponding to the target operation and maintenance operation data. For example, data preprocessing can be performed first, which may include but is not limited to: cleaning and standardization, word segmentation and vectorization, and the cleaned text is converted into a digital vector form that the model can understand. Secondly, the vectorized target operation and maintenance operation data can be input into the classification model, and the model will output a probability distribution, in which the probabilities of write operations and read operations correspond respectively. According to the output probability, the operation type of each data can be determined. For example, a threshold value (such as 0.5, but not limited to this) is set. If the model predicts that the probability of a write operation is greater than the threshold value, the data is determined to be a write operation type, otherwise it is a read operation type. And the threshold value is not fixed. The threshold value can be dynamically adjusted according to the performance of the model, such as the accuracy or recall rate, to improve the classification results.
[0040] It should be noted that when the operation type corresponding to the target operation and maintenance operation data is obtained, the classification result can be saved, that is, the classification result (operation type) can be saved together with the original operation and maintenance operation data for subsequent processing, such as key information extraction, data storage, etc.
[0041] It should be noted that the training steps of the operation classification model include: For the typical text binary classification problem of dividing data into read and write, we use the BERT model for training. The specific training content includes: From the preprocessed text, manually select some data for labeling, mark non-write operations as 0, and write operations as 1. In order to ensure the effect of model training, try to ensure that the ratio of the two types of data is 1:1. The labeled data is divided into a training set and a test set in a ratio of 6:4. The training set is larger to ensure that the model has enough samples for training, and a certain proportion of the test set should also be reserved to detect the effect of model training. For the read-write classification model, we want to identify write operations as comprehensively as possible, and under this premise, allow a small number of read operations to be identified as write operations. Therefore, we use the recall rate (precision) as the evaluation indicator (recall rate: based on actual samples, the proportion of predicted correct positive examples in the total actual positive samples among the samples that are actually positive examples). For example, there are 200 actual write operations, and 206 predicted write operations, of which 190 are real write operations. The recall rate at this time is 190 / 200 = 0.95. We use the basic case-free BERT model (Bert-base-uncased) as the pre-trained language model, and load the tokenizer based on Bert-base-uncased. Through the tokenizer, the training text data is converted into a digital encoding vector, the training round is set to 300, the initial recall rate is 0.5, the Bert-base-uncased pre-trained model is used, and the final number of categories of the classification model is set to 2, and fine-tuned training is performed in combination with our labeled labels. After each round of training, the model is evaluated using the test set. If the recall rate of the test result is greater than the current recall rate value, the model is stored and the recall rate value is updated to the current value. After the training is completed, a better read-write classification model can be obtained.
[0042] Step S106: When the operation type is a write operation type, the named entity recognition model is used to extract the operation action and operation object in the target operation and maintenance operation data.
[0043] The above-mentioned named entity model can be trained in advance by technicians to identify and classify proper nouns in text, namely named entities. In the processing of operation and maintenance data, named entities usually include key entities such as server names, database table names, and configuration file names. These entities have special meanings in operation and maintenance records and can directly point out the objects of operation.
[0044] In an optional embodiment, when the operation type is obtained as a write operation type through the classification operation model, the named entity recognition model can be used to extract the operation actions and operation objects in the target operation and maintenance operation data. For example, data preprocessing can be performed first, which may include but is not limited to: removing irrelevant information, standardizing formats, etc., to ensure that the data is suitable for model input. Secondly, the preprocessed data can be input into the named entity recognition model, at which time the named entity recognition model will analyze each operation data, identify and mark the named entities therein. The model output is usually a sequence, and each word or phrase in the sequence will be marked with an entity type or a "non-entity" type. Then, the legality of the sentence can be judged and normalized through the named entity model, including: for some non-standard sentences (such as missing subjects, inverted sentences, etc.), it is necessary to further judge their legality and try to normalize them into standard sentence structures. This step may use the spaCy natural language processing library to judge whether the entity position is reasonable through part-of-speech tagging and syntactic analysis, and to correct the missing subject, chaotic word order, etc. Then, the named entity recognition model can be used to perform entity position analysis: determine the position of the named entity (operation object) in the sentence, which helps to identify the verbs (operation actions) related to it. Finally, the named entity recognition model can be used for verb extraction: combining the entity position and spaCy's part-of-speech tagging results to identify the verbs most related to the entity, that is, the operation actions.
[0045] It should be noted that spaCy is an open source natural language processing (NLP) library, which is mainly used to process and understand human language and provide high-performance text processing capabilities. spaCy is well-known for its speed and efficiency and is widely used in industry and academia.
[0046] It should be noted that for SAT operation data, the entities that need to be identified are various servers, database tables, and files. The overall process of training the named entity recognition model is as follows: Figure 2 shown. Figure 2 is a schematic diagram of a training method for a named entity recognition model according to an embodiment of the present invention, such as Figure 2As shown in the figure, from the preprocessed text, some data is manually selected for labeling, and different types of entities are marked with custom tags. We represent various types of servers as SVR, databases as TBL, and configurations as CFG. However, some entities are composed of multiple words. We use B to represent the entity start mark and I to represent the subsequent part of the entity. If there is only a single-word entity in the sentence, such as server jboss, we mark it as B-SVR. This recognition model is actually a multi-classification model, which classifies each word in the text into three categories: SVR, TBL, and CFG. The labeled data is divided into a training set and a test set in a ratio of 6:4. The Bert-base-uncased pre-trained language model and the tokenizer based on it are loaded. The accuracy is used as the evaluation indicator (accuracy: the proportion of correct samples to the total samples based on the overall samples). The reason for using this indicator is that each category is very important to the model. Similarly, set the training rounds to 300 (not limited to this), the initial accuracy to 0.5 (not limited to this), use the Bert-base-uncased pre-training model and set the final number of categories of the classification model to 4, and fine-tune the BERT model with the labeled label and training set. After each training round, use the test set to test the trained BERT model, that is, to determine whether the test result is greater than the set recall rate / accuracy. If the accuracy of the test result is greater than the current accuracy value, store the model and update the recall rate / accuracy value to the current value. After training, a better named entity recognition model can be obtained.
[0047] Step S108, storing the target operation and maintenance operation data, operation subject, operation time, operation object and operation action in the first database in correspondence.
[0048] The first database is used to preliminarily store and manage key information extracted from the operation and maintenance data. In this embodiment, it can be an open source relational database management system (MySQL database for short).
[0049] In an optional embodiment, when the operation action and operation object of the target operation and maintenance operation data are obtained, the target operation and maintenance operation data, operation subject, operation time, operation object and operation action can be stored in the first database through the internal system. For example, after extracting the key information, the identified operation object, operation action, text content and the operation subject, time and index name associated with the text content are stored in the MySQL database. The functional interface is implemented to query the data in MySQL, and the database table is retrieved according to the data source and time field by receiving the parameters of the data source and time period, and the query result value is returned to the browser in the form of a JSON array. Among them, JSON is a lightweight data exchange format that is easy for people to read and write, and is also easy for machines to parse and generate. It is based on a subset of JavaScript, but as a data format, it is widely used in various programming languages and environments.
[0050] It should be noted that in order to concise the stored data as much as possible, only the write operation content that actually changes the production environment can be retained, and the key time, operation subject, operation object, and operation action can be extracted from it. The operation subject and the corresponding field from the owner of the log file or the database, and the time from the log header information or the specific field of the data, these two parts can be identified in the data collection stage, so the key goal of the present invention is to classify the read and write operations and identify the operation object and operation action in the write operation.
[0051] In an embodiment of the present invention, the target operation and maintenance operation data from the software transaction service interface is obtained, and the operation subject and operation time are read from the target operation and maintenance operation data; the target operation and maintenance operation data are classified by using a classification operation model to determine the operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; when the operation type is a write operation type, the operation action and operation object in the target operation and maintenance operation data are extracted by using a named entity recognition model; the target operation and maintenance operation data, the operation subject, the operation time, the operation object and the operation action are stored in the first database in correspondence. It is easy to notice that by using the classification operation model, the write operation type in the target operation and maintenance operation data can be accurately identified, so as to further analyze and process the write operation data in a targeted manner. Further, the named entity recognition model can be used to accurately extract the operation action and operation object from the write operation data, so as to achieve the purpose of being able to parse the private domain operation and maintenance data, thereby achieving the technical effect of being able to parse the private domain operation and maintenance data, and thus solving the technical problem that the private domain operation and maintenance data cannot be parsed in the related technology.
[0052] Optionally, obtaining target operation and maintenance data from a software transaction service interface includes: obtaining initial operation and maintenance data from the software transaction service interface; determining a target conversion rule corresponding to the initial operation and maintenance data based on a data type corresponding to the initial operation and maintenance data, wherein the data type includes one of the following: a structured type and an unstructured type; and converting the initial operation and maintenance data using the target conversion rule to obtain target operation and maintenance data.
[0053] The above-mentioned initial operation and maintenance data may be unprocessed operation and maintenance data. The above-mentioned processing rules may be rules set in advance by a technician for processing different types of initial operation and maintenance data to obtain target operation and maintenance data.
[0054] In an optional embodiment, the initial operation and maintenance data can be first obtained from the software transaction service interface; based on the data type corresponding to the initial operation and maintenance data, the target conversion rule corresponding to the initial operation and maintenance data is determined, wherein the data type includes one of the following: structured type and unstructured type. For example, for the semi-structured data (including relevant tags to separate semantic elements and stratify records and fields) recorded in the server, it has certain standardized data header information, which can be segmented according to the data content body start marker. Here, the data header is removed by using a regular expression method, and the data body information part is retained. For another example, for the unstructured data (irregular or incomplete data structure, no predefined data model) recorded in the SAT operation and maintenance tool, all its formats can be manually identified, and each format is processed separately. The program uses a rule-based multi-condition judgment method for parsing. For example, if Chinese appears, it is converted into the corresponding English. If the text contains both underscores and uppercase letters, it is split according to the underscore, etc., and it is uniformly converted into a form that is biased towards natural language for subsequent NLP processing.
[0055] Furthermore, when the data type and target rule of the initial operation and maintenance data are obtained, the initial operation and maintenance data may be converted according to the target conversion rule to obtain the target operation and maintenance data.
[0056] It should be noted that the preprocessed data results are relatively closer to natural language, but there are still problems such as missing subjects and illegal inversion sentences. However, this does not affect the training of the named entity recognition model and the read-write operation classification model, and further processing will be done when the operation action is extracted.
[0057] Optionally, the initial operation and maintenance data is converted using target conversion rules to obtain target operation and maintenance data, including: converting the initial operation and maintenance data using target conversion rules to obtain converted operation and maintenance data; storing the converted operation and maintenance data in a second database; and reading the target operation and maintenance data from the second database based on a scheduled task.
[0058] The second database is used to store the converted operation and maintenance data. The specific time period of the scheduled task can be set according to actual use requirements, which is not limited in this embodiment and can be half an hour, but not limited to this.
[0059] In an optional embodiment, when the data type and target rule of the initial operation and maintenance operation data are obtained, the initial operation and maintenance operation data can be first converted using the target conversion rule to obtain the converted operation and maintenance operation data, and then the converted operation and maintenance operation data can be stored in the second database, and then the target operation and maintenance operation data can be read from the second database based on the scheduled task. For example, after training and obtaining the read-write operation classification model and the operation object recognition model, the model can be used to simplify the data. The data that needs to be simplified has been collected and stored in Elasticsearch (ES) as an index according to the data source. The program uses a scheduled task method to simplify and store data, reads nearly half an hour of data (operation subject, time, text content, index name) from ES every half an hour, associates and binds each operation subject, time, and text content for subsequent storage, and performs subsequent processing on the text content. Input the text content, and call different preprocessing methods for processing according to different indexes of the data source.
[0060] Optionally, the target operation and maintenance operation data is classified using a classification operation model to determine the operation type corresponding to the target operation and maintenance operation data, including: encoding the target operation and maintenance operation data to obtain a digital coding vector; inputting the digital coding vector into the classification operation model to obtain the operation type output by the classification operation model.
[0061] In an optional embodiment, when the target operation and maintenance data is obtained, the target operation and maintenance data can be first encoded to obtain a digital encoding vector, and then the digital encoding vector can be input into the classification operation model to obtain the operation type output by the classification operation model. For example, the processed data is uniformly stored in the list L1, the Bert-base-uncased Tokenizer is loaded, the text data in the list L1 is converted into a digital encoding vector, the read-write operation classification model is used for prediction, and the source list is found through the subscript of the vector data predicted to be 1.
[0062] The text data in L1 is stored in a new list L2, and the above Tokenizer is used to segment the list.
[0063] The text data in L2 is converted into digital encoding vectors and predicted using the operational object recognition model.
[0064] Optionally, a named entity recognition model is used to extract operation actions and operation objects in the target operation and maintenance operation data, including: using the named entity model to perform entity recognition on the target operation and maintenance operation data to obtain the operation object; using a natural language processing library to perform part-of-speech recognition on the next word after the operation object to obtain a target part-of-speech of the next word; based on the target part-of-speech, determining the operation action from the target operation and maintenance operation data.
[0065] The natural language processing library mentioned above is the spaCy natural language processing library.
[0066] In an optional embodiment, when the target operation and maintenance data is obtained, the named entity model can be used to perform entity recognition on the target operation and maintenance data to obtain the operation object; the natural language processing library can be used to perform part of speech recognition on the next word after the operation object to obtain the target part of speech of the next word; based on the target part of speech, the operation action is determined from the target operation and maintenance data. For example, the location information of the entity and the part of speech of the sentence can be determined, that is, the entity position of SVR or TBL or CFG in the sentence is determined (there is only one entity tag in each text data). If it is at the beginning of the sentence, we will load the preview model through spaCy to perform part of speech recognition on the sentence. If the part of speech of the next word of the entity is AUX auxiliary verb, it is determined to be a completely normal natural language sentence, and the phrase of the entity tag is directly extracted as the operation object, and the VERB verb closest to it is used as the operation action. If it is at the beginning of the sentence and the word after spaCy recognizes the entity is not an auxiliary verb, it is determined to be an illegal inverted sentence and its position is swapped. In order to make the spaCy library recognize predicate verbs more accurately, add the first-person subject to the text data.
[0067] Optionally, the target operation and maintenance data is subjected to entity recognition using a named entity model to obtain an operation object, including: performing entity recognition on the target operation and maintenance data using a named entity model to determine a target entity in the target operation and maintenance data, and a target position of the target entity in the target operation and maintenance data; when the target position is the starting position of the target operation and maintenance data, determining that the target entity is the operation object.
[0068] In an optional embodiment, first, a named entity model can be used to perform entity recognition on the target operation and maintenance data to determine the target entity in the target operation and maintenance data, as well as the target position of the target entity in the target operation and maintenance data; secondly, when the target position is the starting position of the target operation and maintenance data, the target entity is determined to be an operation object.
[0069] Optionally, based on the target part of speech, an operation action is determined from the target operation and maintenance data, including: when the target part of speech is a preset part of speech, determining a verb adjacent to the operation object as an operation action; when the target part of speech is not a preset part of speech, replacing the operation object and the next word to obtain first text data, adding a subject to the first text data to obtain second text data, and identifying the operation object and the operation action from the second text data.
[0070] In an optional embodiment, when the target part of speech is a preset part of speech, the verb adjacent to the operation object can be determined as an operation action. When the target part of speech is not a preset part of speech, the operation object and the next word can be replaced to obtain first text data, a subject can be added to the first text data to obtain second text data, and the operation object and the operation action can be identified from the second text data.
[0071] Optionally, the method also includes: obtaining an original data set; marking the read and write operation types of the original data set to obtain a first data set; dividing the first data set to obtain a first training set and a first test set; using the first training set to train the pre-trained language model multiple times, and storing the first training model whose first evaluation indicator during the multiple training processes is greater than the first preset indicator, wherein the first evaluation indicator includes one of the following: recall rate and accuracy; obtaining the first training model corresponding to the maximum first evaluation indicator from the stored first training models to obtain a classification operation model.
[0072] In an optional embodiment, the initial classification operation model can also be trained to obtain a classification operation model. Specifically, the training content may include: first, the original data set can be obtained, and the read and write operation types of the original data set can be marked to obtain a first data set. Secondly, the first data set can be divided to obtain a first training set and a first test set. Then, the pre-trained language model can be trained multiple times using the first training set, and the first training model whose recall rate or accuracy rate during the multiple training processes is greater than that of the first preset indicator is stored. Then, the first training model corresponding to the maximum first evaluation indicator is obtained from the stored first training model to obtain the classification operation model.
[0073] That is, the model is trained so that the model will continue to train after meeting the first preset indicator. For example, the first indicator is set to 50%. When the accuracy reaches 63%, the model will be saved when it exceeds 50%, and the first indicator will be updated to 63%. This cycle repeats until the preset training rounds are reached. The first training model corresponding to the maximum first evaluation indicator is obtained to obtain the classification operation model. For example, the model has a total of 200 training rounds. When the model is trained for 73 times, the accuracy reaches 98%. After the 127 training rounds, the accuracy does not exceed 98%. The model result of the 73rd round will be retained.
[0074] Optionally, the method also includes: labeling entities on the original data set to obtain a second data set; dividing the second data set to obtain a second training set and a second test set; using the second training set to train the pre-trained language model multiple times, and storing the second training model whose recall rate or accuracy rate during the multiple training processes is greater than that of the second preset indicator, and then obtaining the second training model corresponding to the maximum second evaluation indicator from the stored second training models to obtain a named entity recognition model.
[0075] In an optional embodiment, the initial named entity recognition model can also be trained to obtain a named entity recognition model. Specifically, the training content may include: first, the original data set can be entity labeled to obtain a second data set, and then the second data set can be divided to obtain a second training set and a second test set. Then, the second training set can be used to train the pre-trained language model multiple times to obtain a second evaluation index for the training. The second training model whose second evaluation index during the multiple training processes is greater than that meeting the second preset index is stored. The second training model can be stored in a cloud storage platform or in a database for subsequent call or tracing of the training process, wherein the second evaluation index includes one of the following: recall rate and accuracy rate; the second training model corresponding to the maximum second evaluation index is obtained from the stored second training model to obtain a named entity recognition model.
[0076] That is, the model is trained so that the model will continue to train after meeting the second preset indicator. For example, the second indicator is set to 50%. When the accuracy reaches 63%, the model will be saved when it exceeds 50%, and the second indicator will be updated to 63%. This cycle repeats until the preset training round is reached. The second training model corresponding to the maximum second evaluation indicator is obtained to obtain the classification operation model. For example, the model has a total of 200 training rounds. When the model is trained for 73 times, the accuracy reaches 98%. After the 127 training rounds, the accuracy does not exceed 98%. The model result of the 73rd round will be retained.
[0077] Optionally, the method further includes: in response to receiving a query instruction sent by the client, determining a data source and a time period corresponding to the query instruction; querying in the first database based on the data source and the time period to obtain a query result; and feeding back the query result to the client.
[0078] In an optional embodiment, when a query instruction sent by a client is received, the data source and time period corresponding to the query instruction can be determined first, and then a query can be performed in the first database based on the data source and time period to obtain the query result, and finally the query result can be fed back to the client. For example, by selecting the data source and time period on the front-end page to query the processed data, the program will return the corresponding content stored in the database table according to the passed parameters and display it on the page in the form of a list.
[0079] This embodiment proposes a method for extracting key operation and maintenance operations from non-standard complex structure data based on natural language processing (NLP) technology. This method adopts a combination of BERT deep learning model and spaCy natural language processing library, and uses the operation and maintenance operation data of the internal key software SAT (Service API of TravelSky, transaction service interface) as the data source. The pre-processed data is marked using the BERT model for training of read and write operation classification model and named entity recognition model respectively, and all write operation data are screened out using the trained read and write operation classification model, and then the named entity recognition model is combined with the spaCy natural language processing library to extract the operation action and operation object, so as to achieve the effect of streamlining data. The main technical solutions of this embodiment include: selection of data source, selection of key operation and maintenance operation extraction method, and method implementation process.
[0080] Among them, the selection of data sources includes: the operation and maintenance data of the software SAT is relatively diversified, including multiple types of read and write operations; secondly, the structure of the data and the stored data is complex, including semi-structured and unstructured data, which is very challenging to parse and extract. It is more difficult and more popular to simplify such complex and diverse data. Therefore, the present invention selects the operation and maintenance data of the software SAT as the data source.
[0081] Among them, the selection of key operation and maintenance operation extraction methods includes: in order to condense the stored data as much as possible and retain the key information that can be understood by humans, the write operation data that actually changes the production environment should be screened out, and the four parts of time, operation subject, operation object and operation action should be extracted from it. The specific content of the operation can be understood from these four parts of information. The operation subject and the corresponding field from the log file or the database, the time comes from the log header information or the specific field of the data. These two parts can be identified in the data collection stage, so we will focus on the extraction of operation actions and operation objects. SAT operation and maintenance operation data has the characteristics of semi-structured and unstructured, and a wide variety of read and write operations. In addition, SAT operation and maintenance operation data is recorded in human-written code, so it will be more inclined to natural language, but at the same time, there will be problems such as grammatical inversion, abbreviation, spelling errors, and word creation, and the data in the operation and maintenance field itself contains many proprietary words. Therefore, we adopt a combination of multiple NLP technologies: the BERT deep learning model is used to classify read and write operations. Because it covers multiple layers of neural networks, it can also greatly reduce the negative impact of human-written problems such as abbreviations, spelling errors, and word creation. Due to writing problems such as missing subjects and illegal inversion, directly using the spaCy natural language processing library to directly identify SAT operation and maintenance operation data will produce a large number of errors. Therefore, the named entity recognition model is combined with spaCy to convert illegal sentences into natural language text, and then spaCy is used to load the en_core_web_sm model to identify verbs-operation actions, and then the named entity recognition model is used to extract proper nouns-operation objects.
[0082] Among them, the implementation process of the method includes: in the preprocessing stage, specific processing methods are used for different types of operation and maintenance data of the software SAT, and semi-structured and unstructured data are converted into text data that approximates natural language. After obtaining the preprocessed data, enter the model training stage, screen out a small amount of data for read and write operation labeling processing, and try to ensure that the ratio of the number of read and write operation data is 1:1. The labeled data is divided into a training set and a test set, and the BERT model is used for read and write operation classification training, and the model with the highest recall rate is retained; then a small amount of data is screened out for proper noun labeling processing, and different types of entities are marked with custom labels. The labeled data is divided into a training set and a test set, and the BERT model is used for named entity recognition training, and the model with the highest accuracy is retained. After the model training is completed, the evaluation modes of these two models can be used in combination with the spaCy natural language processing library to classify read and write operations and extract operation actions and operation objects. The overall execution process is as follows. Figure 3As shown in the scheduled tasks in . The present invention only extracts key operations from the data (the default data has been collected into ES (Elasticsearch)), reads data from ES at regular intervals, preprocesses it according to its type, and classifies the preprocessed text data using a read-write operation classification model, filters out the content for write operations, and then uses a named entity recognition model to extract the operation object. Next, the entity location information obtained by the named entity recognition model is combined with the part-of-speech recognition of the spaCy library to normalize the data into correct natural language text, and then spaCy is used to extract the required operation actions from it according to the part-of-speech. The operation objects and operation actions streamlined by the program processing and the original data, operation subject, time, and data source index directly read from ES are stored in the MySQL database. Through the above method, while retaining key information, the operation data of the operation and maintenance can be extremely streamlined, providing an effective solution to the problem of the amount of data on the chain of operation and maintenance data.
[0083] Figure 3 is a flowchart of an optional key operation and maintenance operation extraction program execution method according to an embodiment of the present invention, such as Figure 3 As shown, the method comprises the following steps:
[0084] Step S31, setting a scheduled task;
[0085] Step S32, reading data from elastic search and extracting the original data, operation subject, time, and data source;
[0086] Step S33, data preprocessing;
[0087] Step S34, using the read-write classification model (i.e. the above-mentioned operation classification model) to classify the write operation data;
[0088] Step S35, extracting the operation object using a named entity recognition model;
[0089] Step S36, combining the named entity recognition model with the spaCy library (i.e. the above-mentioned natural language processing library) to transform the data and adjust it into natural language text;
[0090] Step S37, using the spaCy library to identify the verb closest to the entity in the text as the operation action;
[0091] Step S38, storing the original data, operation subject, time, data source, operation object and operation action in the MySQL database (ie the first database mentioned above).
[0092] Figure 4FIG. 1 is a flow chart of an optional method for calling a Hypertext Transfer Protocol (HTTP) according to an embodiment of the present invention. Figure 4 As shown, the method comprises the following steps:
[0093] Step S41, first make an HTTP call;
[0094] Step S42, selecting a data source and a time period;
[0095] Step S43, reading the MySQL database;
[0096] Step S44, returning the result to the browser.
[0097] Figure 5 FIG. 1 is a schematic diagram of an optional key operation and maintenance operation extraction program architecture according to an embodiment of the present invention. Figure 5 As shown, it includes a terminal layer, a gateway layer, an application service layer and a database layer. Among them, the terminal layer is connected to the gateway layer, the gateway layer is connected to the application service layer, and the application service layer is connected to the database layer. The personal computer (PC) browser in the terminal layer sends an HTTP request as a terminal. The request must contain the data source and time period parameters in json format. It is transmitted to the back-end service layer based on the Django framework through the gateway layer of the universal web server gateway interface (uSWGI). The service layer parses the data and converts it into a structured query language (SQL) statement to query the MySQL of the storage layer. The query results are returned to the browser in a list form through the uSWGI gateway layer, thereby completing a request interaction. In addition, the service layer itself contains a scheduled task to simplify the data and store key information in the MySQL database. Among them, the Django framework is an advanced server-side programming language web (Python Web) framework that follows the model-view-controller (MVC) design pattern and allows the rapid development of high-performance and scalable web applications.
[0098] The advantage of the technical solution of this embodiment is that it can not only parse complex semi-structured and unstructured data, but also extract the required key information from the data. Parsing based on regular expressions and rules cannot process unstructured data, and for semi-structured data such as application data, it can only extract the data body and cannot be further streamlined. However, this embodiment can be applied to data of all types and structures. Parsing based on clustering and data mining uses unsupervised learning algorithms, which are suitable for classification, but for the case where the read and write operations in the operation and maintenance data are only one word different, it is easy to classify them into one category. The read-write classification model in this embodiment is trained using the BERT model, which is based on supervised multiple neural network training and can effectively divide the read and write operations into two categories. Traditional machine learning is not suitable for extracting and representing features from high-dimensional data such as text, and the single natural language processing technology has a poor recognition effect on operation and maintenance data. The method in this embodiment uses the deep learning-based BERT model for training, which is suitable for processing high-dimensional data such as data text, and uses the technology of named entity recognition to help restore the word order of operation and maintenance data. Combined with the natural language processing library of spaCy, it can effectively identify the parts of speech of text data and extract keywords.
[0099] According to an embodiment of the present invention, an embodiment of a device for processing operation and maintenance data is provided. It should be noted that the device can be used to execute the above-mentioned method for processing operation and maintenance data. Figure 6 Schematic diagram of a device for processing operation and maintenance data according to an embodiment of the present invention. Figure 6 As shown, the device includes: an acquisition module 62, which is used to acquire target operation and maintenance operation data from a software transaction service interface, and read out the operation subject and operation time from the target operation and maintenance operation data; a classification module 64, which is used to classify the target operation and maintenance operation data using a classification operation model, and determine the operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; an extraction module 66, which is used to extract the operation action and operation object in the target operation and maintenance operation data using a named entity recognition model when the operation type is a write operation type; a storage module 68, which is used to store the target operation and maintenance operation data, the operation subject, the operation time, the operation object and the operation action correspondingly in the first database.
[0100] Optionally, the acquisition module includes: an acquisition unit, used to acquire initial operation and maintenance data from a software transaction service interface; a first determination unit, used to determine a target conversion rule corresponding to the initial operation and maintenance operation data based on a data type corresponding to the initial operation and maintenance operation data, wherein the data type includes one of the following: a structured type and an unstructured type; and a conversion unit, used to convert the initial operation and maintenance operation data using the target conversion rule to obtain target operation and maintenance operation data.
[0101] Optionally, the conversion unit includes: a conversion subunit, used to convert the initial operation and maintenance operation data using target conversion rules to obtain converted operation and maintenance operation data; a storage subunit, used to store the converted operation and maintenance operation data to a second database; and a reading subunit, used to read target operation and maintenance operation data from the second database based on a scheduled task.
[0102] Optionally, the classification module includes: an encoding unit, used to encode the target operation and maintenance data to obtain a digital encoding vector; and an input unit, used to input the digital encoding vector into a classification operation model to obtain an operation type output by the classification operation model.
[0103] Optionally, the extraction module includes: a first recognition unit, used to perform entity recognition on the target operation and maintenance data using a named entity model to obtain an operation object; a second recognition unit, used to perform part-of-speech recognition on the next word after the operation object using a natural language processing library to obtain a target part-of-speech of the next word; and a second determination unit, used to determine the operation action from the target operation and maintenance data based on the target part-of-speech.
[0104] Optionally, the first identification unit includes: an identification unit, used to perform entity recognition on the target operation and maintenance data using a named entity model, determine the target entity in the target operation and maintenance data, and the target position of the target entity in the target operation and maintenance data; a first determination subunit, used to determine that the target entity is an operation object when the target position is the starting position of the target operation and maintenance data.
[0105] Optionally, the second determination unit includes: a second determination subunit, used to determine that the verb adjacent to the operation object is an operation action when the target part of speech is a preset part of speech; a replacement subunit, used to replace the operation object and the next word when the target part of speech is not a preset part of speech, to obtain first text data, add a subject to the first text data, to obtain second text data, and identify the operation object and the operation action from the second text data.
[0106] Optionally, the device also includes: a data set acquisition module, used to acquire the original data set; a first labeling module, used to label the read and write operation types of the original data set to obtain a first data set; a first partitioning module, used to partition the first data set to obtain a first training set and a first test set; a first training module, used to use the first training set to train the pre-trained language model multiple times, and store the first training model whose first evaluation indicator during the multiple training processes is greater than the first preset indicator, wherein the first evaluation indicator includes one of the following: recall rate and accuracy; a first acquisition module, used to obtain the first training model corresponding to the maximum first evaluation indicator from the stored first training models to obtain a classification operation model.
[0107] Optionally, the device also includes: a second labeling module, used to perform entity labeling on the original data set to obtain a second data set; a second partitioning module, used to partition the second data set to obtain a second training set and a second test set; a second training module, used to use the second training set to train the pre-trained language model multiple times, and store the second training model whose second evaluation index during the multiple training processes is greater than the second preset index, wherein the second evaluation index includes one of the following: recall rate and accuracy; a second acquisition module, used to obtain the second training model corresponding to the maximum second evaluation index from the stored second training models to obtain a named entity recognition model.
[0108] Optionally, the device also includes: a data determination module, used to determine the data source and time period corresponding to the query instruction in response to receiving the query instruction sent by the client; a query module, used to query in the first database based on the data source and time period to obtain the query result; and a feedback module, used to feed back the query result to the client.
[0109] An embodiment of the present application further provides an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of the present invention when running.
[0110] An embodiment of the present application further provides a computer-readable storage medium, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present invention.
[0111] An embodiment of the present application further provides a computer program product, including a computer program, which implements the methods in various embodiments of the present invention when executed by a processor.
[0112] An embodiment of the present application further provides a computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present invention is implemented.
[0113] The embodiments of the present application further provide a computer program, which implements the methods in the above-mentioned embodiments of the present invention when executed by a processor.
[0114] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0116] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0117] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0118] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0119] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for processing operation and maintenance data, characterized in that: include: Obtaining target operation and maintenance data from the software transaction service interface, and reading the operation subject and operation time from the target operation and maintenance data; Classifying the target operation and maintenance operation data using a classification operation model to determine an operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; In the case where the operation type is a write operation type, extracting the operation action and the operation object in the target operation and maintenance operation data by using a named entity recognition model; The target operation and maintenance data, the operation subject, the operation time, the operation object and the operation action are stored in a first database in correspondence.
2. The method according to claim 1, characterized in that Obtain target operation and maintenance data from the software transaction service interface, including: Acquiring initial operation and maintenance data from the software transaction service interface; Determining a target conversion rule corresponding to the initial operation and maintenance operation data based on a data type corresponding to the initial operation and maintenance operation data, wherein the data type includes one of the following: a structured type and an unstructured type; The initial operation and maintenance data is converted using the target conversion rule to obtain the target operation and maintenance data.
3. The method according to claim 2, characterized in that The initial operation and maintenance data is converted by using the target conversion rule to obtain the target operation and maintenance data, including: The initial operation and maintenance data is converted by using the target conversion rule to obtain converted operation and maintenance data; storing the converted operation and maintenance data in a second database; The target operation and maintenance data is read from the second database based on a scheduled task.
4. The method according to claim 1, characterized in that: Classifying the target operation and maintenance operation data by using a classification operation model to determine the operation type corresponding to the target operation and maintenance operation data includes: Encoding the target operation and maintenance data to obtain a digital encoding vector; The digital coding vector is input into the classification operation model to obtain the operation type output by the classification operation model.
5. The method according to claim 1, characterized in that The named entity recognition model is used to extract the operation actions and operation objects in the target operation and maintenance operation data, including: Using the named entity model to perform entity recognition on the target operation and maintenance operation data to obtain the operation object; Using a natural language processing library to perform part-of-speech recognition on the next word after the operation object to obtain a target part-of-speech of the next word; Based on the target part of speech, the operation action is determined from the target operation and maintenance operation data.
6. The method according to claim 5, characterized in that Using the named entity model to perform entity recognition on the target operation and maintenance operation data to obtain the operation object includes: Performing entity recognition on the target operation and maintenance data using the named entity model to determine a target entity in the target operation and maintenance data and a target position of the target entity in the target operation and maintenance data; In a case where the target position is the starting position of the target operation and maintenance data, the target entity is determined to be the operation object.
7. The method according to claim 5, characterized in that Determining the operation action from the target operation and maintenance data based on the target part of speech includes: In the case where the target part of speech is a preset part of speech, determining a verb adjacent to the operation object as the operation action; When the target part of speech is not a preset part of speech, the operation object and the next word are replaced to obtain first text data, a subject is added to the first text data to obtain second text data, and the operation object and the operation action are identified from the second text data.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Get the original data set; Marking the original data set with a read and write operation type to obtain a first data set; Dividing the first data set into a first training set and a first test set; The pre-trained language model is trained multiple times using the first training set, and a first training model having a first evaluation index greater than a first preset index during the multiple training processes is stored, wherein the first evaluation index includes one of the following: recall rate and precision rate; The first training model corresponding to the maximum first evaluation index is obtained from the stored first training models to obtain the classification operation model.
9. The method according to claim 8, characterized in that The method further comprises: Performing entity labeling on the original data set to obtain a second data set; Dividing the second data set to obtain a second training set and a second test set; The pre-trained language model is trained multiple times using the second training set, and a second training model whose second evaluation index is greater than the second preset index during the multiple training processes is stored, wherein the second evaluation index includes one of the following: recall rate and precision rate; The second training model corresponding to the maximum second evaluation index is obtained from the stored second training models to obtain the named entity recognition model.
10. The method according to any one of claims 1 to 7, characterized in that The method further comprises: In response to receiving a query instruction sent by a client, determining a data source and a time period corresponding to the query instruction; Performing a query in the first database based on the data source and the time period to obtain a query result; The query result is fed back to the client.
11. A device for processing operation and maintenance data, characterized in that: include: An acquisition module, used to acquire target operation and maintenance data from a software transaction service interface, and read out an operation subject and an operation time from the target operation and maintenance data; A classification module, used to classify the target operation and maintenance operation data using a classification operation model, and determine an operation type corresponding to the target operation and maintenance operation data, wherein the operation type includes one of the following: a write operation type or a read operation type; An extraction module, used for extracting operation actions and operation objects in the target operation and maintenance operation data by using a named entity recognition model when the operation type is a write operation type; The storage module is used to store the target operation and maintenance operation data, the operation subject, the operation time, the operation object and the operation action in a first database in correspondence.
12. An electronic device, characterized in that: include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 10 when running.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that It comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.