Operation and maintenance data classification method and device

By extracting the key category features in the operation and maintenance data text and calculating the importance of keywords, combining feature fusion processing and classifier output, the problem of inaccurate classification of operation and maintenance data in traditional methods is solved, and high-accuracy refined classification is achieved.

CN120045714APending Publication Date: 2025-05-27BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063641.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional operation and maintenance data classification methods are difficult to understand the deep semantics and contextual relationships of professional knowledge content, resulting in inaccurate classification results.

Method used

A method of classification of operation and maintenance data is proposed, which can obtain operation and maintenance data text, preprocess, extract key category characteristics, calculate the importance value of keywords, feature fusion processing, and input a pre-constructed classifier to output the corresponding label category.

Benefits of technology

It realizes the accurate and refined classification of operation and maintenance data, and improves the refined management level of operation and maintenance data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045714A_ABST
    Figure CN120045714A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide an operation and maintenance data classification method and apparatus. The method comprises the steps of obtaining an operation and maintenance data text; preprocessing the operation and maintenance data text to obtain a preprocessed operation and maintenance data text; key category features are extracted from the preprocessed operation and maintenance data text; wherein the key category feature comprises a plurality of keywords; calculating an importance degree value of each keyword, and selecting at least one keyword meeting an importance condition according to the importance degree value of each keyword; performing feature fusion processing on the key category feature and the importance degree value of the at least one keyword to obtain a classification feature; and inputting the classification features into a pre-constructed classifier, and outputting a corresponding label category by the classifier. According to the method and the device, accurate and refined classification can be realized by combining the statistical characteristics and the contextual relationship of the keywords in the operation and maintenance data text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method and device for classifying operation and maintenance data. Background Art

[0002] The operation and maintenance data of engineering projects in the energy industry has the characteristics of strong professionalism, rich technical details, precise equipment specifications, and strict specification standards. The data volume is huge and the data types are complex and diverse. Using traditional numerical features or keyword matching methods for data classification, due to the difficulty in understanding the deep semantics and context relationships of professional knowledge content, the classification results are not accurate enough. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present application is to propose a method and device for classifying operation and maintenance data to solve the problem of inaccurate classification of operation and maintenance data.

[0004] Based on the above purpose, the embodiments of the present application provide a method for classifying operation and maintenance data, including:

[0005] Obtain operation and maintenance data text;

[0006] Preprocess the operation and maintenance data text to obtain the preprocessed operation and maintenance data text;

[0007] Extract key category features from the preprocessed operation and maintenance data text; wherein, the key category features include multiple keywords;

[0008] Calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance value of each keyword;

[0009] Perform feature fusion processing on the key category features and the importance values of at least one keyword to obtain classification features;

[0010] Input the classification features into a pre-constructed classifier, and the classifier outputs the corresponding label category.

[0011] Optionally, extracting key category features from the preprocessed operation and maintenance data text includes:

[0012] Input the preprocessed operation and maintenance data text into a pre-constructed language representation model, use the language representation model to perform text vector conversion on the preprocessed operation and maintenance data, and extract key category features from the converted text vector representation.

[0013] Optionally, calculating the importance value of each keyword, and selecting at least one keyword that meets the importance condition according to the importance value of each keyword, includes:

[0014] Calculate the TF-IDF value of each keyword;

[0015] Sort the keywords in descending order according to their TF-IDF values;

[0016] Based on the sorted keywords, select a predetermined number of the top-ranked keywords, or select the keywords whose TF-IDF values are greater than a preset importance threshold.

[0017] Optionally, the key category features include the scores of each keyword; calculating the TF-IDF value of each keyword includes:

[0018] Determine the weight of a keyword according to the score of the keyword;

[0019] Determine the importance-weighted term frequency according to the weight of the keyword and the calculated term frequency;

[0020] Determine the importance inverse document frequency according to the importance level of the operation and maintenance data text and the calculated inverse document frequency;

[0021] Calculate the TF-IDF value according to the importance-weighted term frequency and the importance inverse document frequency.

[0022] An embodiment of the present application further provides an operation and maintenance data classification device, including:

[0023] An acquisition module, configured to acquire an operation and maintenance data text;

[0024] A preprocessing module, configured to preprocess the operation and maintenance data text to obtain a preprocessed operation and maintenance data text;

[0025] An extraction module, configured to extract key category features from the preprocessed operation and maintenance data text; wherein, the key category features include a plurality of keywords;

[0026] A calculation module, configured to calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance values of the keywords;

[0027] A fusion module, configured to perform feature fusion processing on the key category features and the importance values of at least one keyword to obtain classification features;

[0028] A classification module, configured to input the classification features into a pre-constructed classifier, and output a corresponding label category by the classifier.

[0029] Optionally, the extraction module is configured to input the preprocessed operation and maintenance data text into a pre-constructed language representation model, use the language representation model to perform text vector conversion on the preprocessed operation and maintenance data, and extract key category features from the converted text vector representation.

[0030] Optionally, the calculation module is configured to calculate the TF-IDF value of each keyword; sort the keywords in descending order according to the TF-IDF value; based on the sorted keywords, select a predetermined number of keywords ranked at the front, or select keywords whose TF-IDF value is greater than a preset importance threshold.

[0031] Optionally, the key category features include the scores of each keyword;

[0032] The calculation module is configured to determine the weight of a keyword according to the score of each keyword; determine the importance weighted term frequency according to the weight of the keyword and the calculated term frequency; determine the importance inverse document frequency according to the importance level of the operation and maintenance data text and the calculated inverse document frequency; calculate the TF-IDF value according to the importance weighted term frequency and the importance inverse document frequency.

[0033] As can be seen from the above, the operation and maintenance data classification method and device provided by the embodiments of the present application extract key category features from the preprocessed operation and maintenance data text, calculate the importance degree value of each keyword, select at least one keyword that meets the importance condition therefrom, perform feature fusion processing on the key category features and the importance degree value of the at least one keyword, and input the fused classification features into a pre-constructed classifier, and the classifier outputs the corresponding label category. The present application combines the statistical characteristics and context relationships of keywords in the operation and maintenance data text, can achieve accurate and refined classification, and improve the refined management level of operation and maintenance data. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic flowchart of the method of the embodiment of the present application;

[0036] Figure 2 It is a block diagram of the device structure of the embodiment of the present application;

[0037] Figure 3 It is a block diagram of the electronic device structure of the embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0039] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the embodiments of the present application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0040] As Figure 1 shown, the embodiments of the present application provide an operation and maintenance data classification method, including:

[0041] S101: Obtain operation and maintenance data texts;

[0042] S102: Preprocess the operation and maintenance data texts to obtain preprocessed operation and maintenance data texts;

[0043] In this embodiment, the original operation and maintenance data texts are obtained from a data source. For example, relevant operation and maintenance data are obtained from an engineering project system. The operation and maintenance data are generally in an unstructured text form, including but not limited to technical documents, operation manuals, maintenance records, fault logs, etc. Optionally, in the power industry scenario, to ensure the comprehensiveness and diversity of classification, the operation and maintenance data in various power engineering systems are collected for the operation and maintenance data classification in the power industry.

[0044] For the convenience of subsequent processing, the obtained operation and maintenance data texts are preprocessed to obtain normalized and standardized operation and maintenance data texts. Among them, preprocessing the operation and maintenance data texts includes data cleaning. As shown in Table 1, data cleaning includes but is not limited to deleting null values, invalid redundant data, removing stop words, correcting spelling mistakes, etc., and according to industry knowledge, the technical terms are uniformly normalized to ensure data quality; then, the cleaned data is converted into a unified text format, such as a plain text or a formatted text form, and the specific text format is not limited.

[0045] Table 1 Partial data cleaning content

[0046]

[0047] S103: Extract key category features from the preprocessed operation and maintenance data text; among them, the key category features include multiple keywords.

[0048] In this embodiment, after preprocessing the operation and maintenance data text, a pre-constructed language representation model is used to extract key category features. That is, the preprocessed operation and maintenance data text is input into the pre-constructed language representation model, and the preprocessed operation and maintenance data text is converted into a text vector representation by the language representation model. Key category features are extracted from the converted text vector representation. Among them, the key category features output by the language representation model include multiple keywords and the scores of each keyword. The keyword is an important professional word in the operation and maintenance data that plays an important role in classification.

[0049] In some ways, the BERT model based on the Transformer large language model is used to implement the language representation model. The encoder-decoder architecture and the self-attention mechanism are adopted to convert the input operation and maintenance data text into a corresponding text vector representation. The text vector representation contains the context information and semantic information of the operation and maintenance data text. Then, the attention mechanism is used to extract keywords and the scores of keywords from the converted text vector representation. In implementation, the model structure can be adjusted according to the task requirements in the specific application scenario to adapt to the data scale and complexity of the specific task.

[0050] To improve the generalization ability of the language representation model, data augmentation can be achieved by means such as synonym replacement, sentence recombination, and random noise injection, constructing a rich and diverse data sample set, and training the language representation model with the augmented data sample set. In the training stage, the masked language model and / or the next sentence prediction task are used to train the language representation model, and a large number of labeled operation and maintenance data text samples are used for long-term pre-training to enable the model to learn rich context information and domain knowledge.

[0051] Among them, for the masked language model task, some words in the operation and maintenance data text sample are randomly selected for masking, and the selected words are replaced with the special token [MASK]. The masking ratio can be adjusted according to the task requirements, such as a masking ratio of 15%. Select a deep learning framework (such as TensorFlow or PyTorch) to load the pre-trained language representation model, set the training parameters, input the masked text into the model, calculate the prediction result through forward propagation, use the cross-entropy loss function to calculate the difference between the prediction result and the actual result, and update the model parameters through backpropagation. For the next sentence prediction task, sentence pairs are randomly sampled from the operation and maintenance data text sample. One part is used as positive samples (real consecutive sentences), and the other part is used as negative samples (randomly combined non-consecutive sentences). A binary classifier is constructed using a deep learning framework and a pre-trained language representation model. The embedded representation of the sentence pair is used as the input of the classifier, and the probability of the continuity of the sentence pair is output. The cross-entropy loss function is used to calculate the loss of the classification result, and the parameters of the classifier are updated through backpropagation.

[0052] In some ways, the language representation model is trained by combining the masked language model and the next sentence prediction task. By adjusting parameters such as the learning rate, batch size, and number of training epochs of the model, the model performance is optimized. The early stopping method is adopted to prevent overfitting, and cross-validation is used to evaluate the model performance until the optimal effect is achieved.

[0053] S104: Calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance values of the keywords.

[0054] In this embodiment, after obtaining multiple keywords of the operation and maintenance data text and the scores of each keyword by using the language representation model, the TF-IDF value of each keyword is statistically calculated. According to the TF-IDF values of the keywords, the keywords in the operation and maintenance data text that have a greater impact on data classification are selected for subsequent operation and maintenance data classification. Among them, calculating the importance value of each keyword and selecting at least one keyword that meets the importance condition according to the importance values of the keywords includes:

[0055] Statistically calculate the TF-IDF value of each keyword;

[0056] Sort the keywords in descending order according to the TF-IDF values;

[0057] Based on the sorted keywords, select a predetermined number of keywords ranked at the front, or select the keywords with TF-IDF values greater than the preset importance threshold.

[0058] In this embodiment, for each keyword, the term frequency (TF) of the keyword is counted. Based on all the texts in the operation and maintenance data text set, the inverse document frequency (IDF) is counted. The TF-IDF value is calculated according to the term frequency and the inverse document frequency and used as the importance value of the keyword. According to the magnitudes of the TF-IDF values of the keywords, keywords greater than the importance threshold or a predetermined number of keywords with larger TF-IDF values are selected for subsequent data classification.

[0059] Among them, counting the TF-IDF value of each keyword includes:

[0060] Determine the weight of the keyword according to the score of each keyword;

[0061] Determine the importance weighted term frequency according to the weight of the keyword and the calculated term frequency;

[0062] Determine the importance inverse document frequency according to the importance of the operation and maintenance data text and the calculated inverse document frequency;

[0063] Calculate the TF-IDF value according to the importance weighted term frequency and the importance inverse document frequency.

[0064] In this embodiment, when calculating the term frequency, the weight can be set in combination with the score of the keyword. The greater the score of the keyword predicted by the model, the higher the weight set when calculating the term frequency, that is, the importance weighted term frequency considering the importance degree of the keyword is calculated by combining the term frequency and weight of the keyword. When calculating the inverse document frequency, the inverse document frequency is weighted and adjusted according to the particularity and importance of the operation and maintenance data text. For technical terms that frequently appear in the operation and maintenance data text but are rare in non-operation and maintenance data files, the obtained inverse document frequency is relatively large. Thus, the TF-IDF value is calculated according to the importance weighted term frequency and the importance inverse document frequency.

[0065] Table 2 TF-IDF values of some keywords

[0066] Keyword Term Frequency (TF) Inverse Document Frequency (IDF) TF-IDF Value Instance Name 5 2.3 11.5 Running Status 3 2.1 6.3 Resource Utilization Rate 2 1.9 3.8

[0067] S105: Perform feature fusion processing on the key category features and the importance values of at least one keyword to obtain classification features;

[0068] In this embodiment, the key category features are obtained by using a language representation model. After selecting important keywords according to the importance values of the keywords, the key category features and the keywords are subjected to feature fusion by using a preset fusion method for subsequent classification based on the fused classification features. In some ways, methods such as feature splicing, weighted summation, or learning the fusion weight through a neural network can be used to form classification features with both statistical characteristics and deep semantic information.

[0069] S106: Input the classification features into a pre-constructed classifier, and the classifier outputs the corresponding label categories.

[0070] In this embodiment, a pre-constructed classifier is used to classify the classification features, and the classifier outputs the category label to which the classification features belong, that is, to determine the category label to which the operation and maintenance data text belongs, thereby realizing the classification of the operation and maintenance data text. Based on the statistical characteristics and context relationships of the keywords in the operation and maintenance data text, accurate and refined classification can be achieved.

[0071] Table 3 Partial Category Labels

[0072] Knowledge Number Technical Type Label Device Type Label Technical Standard Label C001 Resource Specification Server Load Balancing (SLB) National Technical Standard C002 Resource Utilization Rate Message Queue (MQ) International Technical Standard

[0073] In some ways, specific category labels are determined according to the actual application scenario, and the category labels should accurately reflect the core content of the transportation data. For example, when applied to the power industry, the category labels of the classifier can include multiple dimensions such as technical types, equipment types and specifications, and technical standards and specifications.

[0074] In some embodiments, through joint training of a language representation model, a classifier, and a statistical module for statistically calculating the importance values of keywords, an operation and maintenance data classification model for classifying operation and maintenance data texts is obtained. When training the operation and maintenance data classification model, a suitable loss function is designed, such as cross-entropy loss, and at the same time, the parameters of the language representation model and the fusion weights of the TF-IDF values are optimized. An alternating training strategy is adopted. First, the language representation model and the statistical module are trained separately. After the two parts of the model converge, they are jointly trained and the overall data classification model is fine-tuned to achieve the best performance collaboration. In some ways, methods such as Bayesian optimization and grid search are used for hyperparameter search, and systematic tuning is performed on the learning rate, batch size, number of training epochs, parameters of the fusion layer, etc. At the same time, consider using model pruning and quantization techniques to reduce the model size and computational amount, and improve the response speed and resource efficiency in actual applications.

[0075] In some embodiments, for the trained operation and maintenance data classification model, the model is verified using operation and maintenance data text validation sets from different sources and scales to evaluate the generalization ability and stability of the model in different scenarios. The selected validation sets should cover as many types and scenarios of operation and maintenance knowledge as possible. After verification, compare the performance indicators of the model on different validation sets, such as accuracy, recall rate, F1 score, etc., to evaluate the robustness and consistency of the model. The verified operation and maintenance data classification model can be used to determine the category label of the operation and maintenance data text, that is, to classify the operation and maintenance data text. During the application process, the classification accuracy of the operation and maintenance data can be improved by continuously optimizing and updating the model, the model performance can be enhanced, and the requirements for refined management can be met.

[0076] In some ways, after using a classifier to determine the category label corresponding to the operation and maintenance data text, the operation and maintenance data text and the corresponding category label are saved in a database to form an operation and maintenance knowledge dataset including operation and maintenance data text data and their categories, and a classification index and a search engine are established to improve the retrieval efficiency and accuracy. Among them, the operation and maintenance data text includes an operation and maintenance knowledge number, a generation date, etc. Based on the created operation and maintenance knowledge dataset, retrieval can be provided according to keywords such as the operation and maintenance knowledge number, category label, generation date, etc.

[0077] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.

[0078] It should be noted that the specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0079] As Figure 2 shown, the embodiment of the present application also provides an operation and maintenance data classification device, including:

[0080] An acquisition module, configured to acquire operation and maintenance data text;

[0081] A preprocessing module, configured to preprocess the operation and maintenance data text to obtain the preprocessed operation and maintenance data text;

[0082] An extraction module, configured to extract key category features from the preprocessed operation and maintenance data text; among them, the key category features include multiple keywords;

[0083] A calculation module, configured to calculate the importance degree value of each keyword, and select at least one keyword that meets the importance condition according to the importance degree values of the keywords;

[0084] A fusion module, configured to perform feature fusion processing on the key category features and the importance degree values of at least one keyword to obtain classification features;

[0085] A classification module, configured to input the classification features into a pre-constructed classifier, and the classifier outputs the corresponding label category.

[0086] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0087] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0088] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0089] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0090] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0091] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0092] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to enable communication interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0093] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0094] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0095] The electronic device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0096] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0097] Those of ordinary skill in the art should understand that: the discussion of any above embodiment is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the idea of the present disclosure, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

[0098] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0099] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0100] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present disclosure.

Claims

1. A method for classifying operation and maintenance data, characterized in that: include: Get operation and maintenance data text; Preprocessing the operation and maintenance data text to obtain a preprocessed operation and maintenance data text; Extracting key category features from the preprocessed operation and maintenance data text; wherein the key category features include multiple keywords; Calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance value of each keyword; Performing feature fusion processing on the key category feature and the importance value of at least one keyword to obtain a classification feature; The classification features are input into a pre-built classifier, and the classifier outputs a corresponding label category.

2. The method according to claim 1, characterized in that: Extract key category features from the preprocessed operation and maintenance data text, including: The preprocessed operation and maintenance data text is input into a pre-built language representation model, the preprocessed operation and maintenance data is converted into a text vector using the language representation model, and key category features are extracted from the converted text vector representation.

3. The method according to claim 1, characterized in that Calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance value of each keyword, including: Count the TF-IDF value of each keyword; Sort the keywords in descending order according to the TF-IDF value; Based on the sorted keywords, a predetermined number of keywords that are ranked in the front are selected, or keywords whose TF-IDF values ​​are greater than a preset importance threshold are selected.

4. The method according to claim 3, characterized in that The key category feature includes the score of each keyword; the statistical TF-IDF value of each keyword includes: Determine the weight of the keyword based on the score of each keyword; Determine the importance-weighted word frequency based on the keyword weights and the calculated word frequency; Determine the importance reverse file frequency according to the importance of the operation and maintenance data text and the calculated reverse file frequency; The TF-IDF value is calculated based on the importance-weighted word frequency and the importance-inverse document frequency.

5. An operation and maintenance data classification device, characterized in that: include: Acquisition module, used to obtain operation and maintenance data text; A preprocessing module, used to preprocess the operation and maintenance data text to obtain the preprocessed operation and maintenance data text; An extraction module, used to extract key category features from the preprocessed operation and maintenance data text; wherein the key category features include multiple keywords; A calculation module, used to calculate the importance value of each keyword, and select at least one keyword that meets the importance condition according to the importance value of each keyword; A fusion module, used for performing feature fusion processing on the key category feature and the importance value of at least one keyword to obtain a classification feature; The classification module is used to input the classification features into a pre-built classifier, and the classifier outputs a corresponding label category.

6. The device according to claim 5, characterized in that The extraction module is used to input the preprocessed operation and maintenance data text into a pre-built language representation model, use the language representation model to perform text vector conversion on the preprocessed operation and maintenance data, and extract key category features from the converted text vector representation.

7. The device according to claim 5, characterized in that The calculation module is used to count the TF-IDF value of each keyword; sort the keywords in descending order of TF-IDF value; based on the sorted keywords, select a predetermined number of keywords that are ranked in front, or select keywords whose TF-IDF value is greater than a preset importance threshold.

8. The device according to claim 7, characterized in that The key category features include scores of each keyword; The calculation module is used to determine the weight of the keyword according to the score of each keyword; determine the importance weighted word frequency according to the weight of the keyword and the calculated word frequency; According to the importance of the operation and maintenance data text and the calculated reverse file frequency, the importance reverse file frequency is determined; and according to the importance weighted word frequency and the importance reverse file frequency, the TF-IDF value is calculated.