Information knowledge extraction method based on power system, controller and storage medium

By using the joint method of RoBERTa, BiLSTM and MFM models in the power system for data preprocessing and optimization word segmentation, a mapping rule library for equipment entities was established, and the problem of difficulty in identifying non-structural data equipment entities in the power system was solved, and high-precision and high-efficiency equipment entity recognition and extraction were achieved.

CN120163147AActive Publication Date: 2025-06-17CHINA SOUTHERN POWER GRID ENERGY STORAGE CO LTD INFORMATION & COMM BRANCH

Patent Information

Application Number
CN202510008381.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-17
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and extract equipment entities in non-structured data in power systems, resulting in the limitation of the development of smart grids.

Method used

The combined method of text word segmentation and part-of-speech labeling based on the RoBERTa model, BiLSTM model and MFM model is adopted, and combined with the power system's special text lexicon and the reverse maximum matching algorithm, the power system data elements are preprocessed and optimized to form a standard power system metadatabase, and a mapping rule database between data element information and device entities is established.

Benefits of technology

It improves the accuracy and efficiency of the identification of the entity entity of the power system data element naming equipment, and can effectively identify and extract the categories and models of the equipment entity in the power system data, making up for the inefficiency and inaccuracy of manual sampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163147A_ABST
    Figure CN120163147A_ABST
Patent Text Reader

Abstract

The invention discloses an information knowledge extraction method based on a power system, a controller and a storage medium. The method comprises the following steps: reading non-structural data of an information document of the power system; performing data cleaning on the non-structural data to obtain power system data elements; preprocessing the data elements of the power system to obtain segmented words, and correcting the segmented words by adopting a reverse maximum matching algorithm; based on a standard power system metadatabase, according to the construction direction of a power system, a data meta-information and entity extraction rule is established, and device entity category candidate word extraction is achieved; identifying and filtering the equipment entity category candidate words to obtain high-frequency nouns; performing entity joint extraction on the equipment entity model candidate words to obtain confidence high-frequency vocabularies; generating equipment entity category and model labels according to the high-frequency nouns and the vocabularies with the highest confidence coefficient, and storing the equipment entity category and model labels into an index; and outputting equipment entity knowledge according to the leading rope. According to the method, the entity identification accuracy of the power system data element naming equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power systems, and in particular, to an information knowledge extraction method, a controller, and a storage medium based on a power system. Background Art

[0002] With the construction and development of the energy Internet, smart grid, and ubiquitous power Internet of Things, the long-term production and operation in the field of power systems have accumulated a large amount of power multi-source heterogeneous big data, including data in various links such as power production, transmission, and sales. This big data can provide support for the construction of the smart grid. However, most of the accumulated power multi-source heterogeneous big data is unstructured data. The unstructured data has diverse formats and relatively implicit data meanings, and cannot be stored in a relational database. It can only be stored in different file forms. The massive multi-source heterogeneous data will bring challenges in storage, transmission, information processing, etc. Therefore, how to automatically analyze and mine the unstructured data in the power field has become one of the bottlenecks restricting the development of the smart grid at the present stage.

[0003] In related technologies, for information extraction from unstructured data, it is mainly based on rules and dictionaries, conditional random fields (CRFs), classifiers, etc. Such methods require a large number of manually labeled templates and training data, have relatively high requirements for domain knowledge, and have low scalability. In recent years, with the continuous development of deep learning models, the entity recognition and relationship extraction methods based on deep learning have been improved in performance. However, due to the complex and diverse word formation forms of medical entities contained in the Chinese texts of power systems, they often appear in the form of declarative phrases, such as device entities. Therefore, the Chinese word segmentation and entity recognition tasks often face a large number of nested entities, which results in the frequent occurrence of entity nesting phenomena in the Chinese texts of power systems. However, the current methods for model entity recognition often ignore the unique entity nesting rules of the Chinese texts of power systems and directly adopt the sequence labeling method, resulting in problems such as low accuracy of domain entity recognition and inability to accurately identify entities, and unable to accurately identify and extract the content of power grid data. Summary of the Invention

[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the embodiments of the present application provide an information knowledge extraction method, a controller, and a storage medium based on a power system, which can achieve accurate recognition and extraction of power system data elements, and is beneficial to improving the accuracy of naming device entities of power system data elements.

[0005] In a first aspect, the embodiments of the present application provide an information knowledge extraction method based on a power system, including:

[0006] Obtain the information document of the power system, read the unstructured data of the information document of the power system, and convert the unstructured data into txt format;

[0007] Perform data cleaning, classification, cataloging, and standardization processing on the unstructured data to obtain power system data elements;

[0008] Preprocess the power system data elements through the character embedding layer based on the RoBERTa model, the BiLSTM model, and the tag speculation CRF layer based on the MFM model, combined with the joint method of power system text word segmentation and part-of-speech tagging, to obtain word segments, perform recombination and vectorization processing on the word segments, and use the reverse maximum matching algorithm to correct the word segments to obtain an optimized word segmentation result;

[0009] Train the optimized word segmentation result based on the power system specific text word library to obtain a standard power system metadata database, establish rules for data element information and entity extraction according to the construction direction of the power system, and form a mapping rule library for data element information and device entities to realize the extraction of candidate words for device entity categories;

[0010] Identify and filter the candidate words for device entity categories to obtain high-frequency nouns;

[0011] Perform entity joint extraction on the candidate words for device entity models to obtain high-confidence high-frequency vocabulary;

[0012] Identify the high-confidence high-frequency vocabulary to obtain the vocabulary with the highest confidence;

[0013] Generate device entity categories and model labels corresponding to the power system data elements according to the high-frequency nouns and the vocabulary with the highest confidence and store them in the index;

[0014] Output device entity knowledge according to the index.

[0015] According to some embodiments of the present application, the identifying and filtering the candidate words for device entity categories to obtain high-frequency nouns includes:

[0016] Form a basic rule knowledge base for device entity application scenarios, entity boundaries, and semantics based on power system standards and specification knowledge;

[0017] Based on the basic rule knowledge base for device entity application scenarios, entity boundaries, and semantics, fuse the Chinese entity recognition method for device entity nesting rules, and convert the recognition task of device entities into a joint training task for boundary recognition and boundary start-end relationship recognition of device entities;

[0018] During the decoding process of the joint training task, the recognition and filtering process of the decoding result is combined with the entity nesting rule of the power system, so that the recognition result conforms to the composition rule of the inner and outer layer entity nesting combination in the text;

[0019] Terminologize and standardize the candidate words of the device entity category according to the terminology standard of the power system to obtain high-frequency nouns.

[0020] According to some embodiments of the present application, the preprocessing of the power system data element is performed by combining the character embedding layer of the RoBERTa model, the BiLSTM model, and the tag speculation CRF layer based on the MFM model, and the power system text word segmentation and part-of-speech tagging joint method to obtain word segmentation, including:

[0021] Input the power system data element into the character embedding layer of the RoBERTa model to obtain a word vector reflecting the context semantics and output the word vector;

[0022] Input the word vector into the BiLSTM model, and capture the long-distance dependence relationship and context sequence information of the word vector through the BiLSTM model to discard the useless word vector sequence and extract the useful word vector sequence;

[0023] Input the useful word vector sequence into the tag speculation CRF layer of the MFM model, and jointly model the tag sequence by using the tag speculation CRF layer of the MFM model to calculate the conditional probability of the input useful word vector sequence output in the tag sequence;

[0024] Determine the highest conditional probability of the output, and select the useful word vector sequence corresponding to the highest conditional probability as the word segmentation.

[0025] According to some embodiments of the present application, the inputting the power system data element into the character embedding layer of the RoBERTa model to obtain a word vector reflecting the context semantics and outputting the word vector includes:

[0026] Calculate the word vector of the power system data element with the first parameter matrix, the second parameter matrix, and the third parameter matrix respectively to obtain a first vector, a second vector, and a third vector with the same dimension as the word vector;

[0027] Construct an attention matrix according to the first vector, the second vector, and the third vector;

[0028] Perform mappings of different linear transformations on the first vector, the second vector, and the third vector, and splice them through the attention matrix to obtain a multi-head attention matrix;

[0029] Connect the multi-head attention matrix to a feed-forward neural network, encode the power system data element through linear transformation and the Relu activation function, and perform preliminary word segmentation on the power system data element using a Chinese word segmentation algorithm to obtain a word vector reflecting the context semantics and output the word vector.

[0030] According to some embodiments of the present application, the BiLSTM model includes a forgetting gate, an input gate, and an output gate. Input the word vector into the BiLSTM model, and capture the long-distance dependence relationship and context sequence information of the word vector through the BiLSTM model to discard useless word vector sequences and extract useful word vector sequences, including:

[0031] Input the word vector into the forgetting gate of the BiLSTM model. Determine the information to be discarded in the previous neuron through the forgetting gate. The input is the current word vector and the previous word vector. Map the previous neuron cell state to 0-1, where 0 means completely deleted and 1 means completely retained, to obtain the result of the forgetting gate;

[0032] Determine whether to record a new word vector into the neuron of the BiLSTM model through the input gate. Update the neuron state by determining the word vector to be updated through the Sigmoid activation function of the input layer to determine the memory information to be recorded, and obtain the word vector sequence to be remembered;

[0033] Determine the part of the output neuron state through the Sigmoid activation function of the output gate. Process the neuron state through the tanh function and multiply it by the output of the Sigmoid activation function to obtain a hidden layer state sequence with the same length as the word vector sequence, so as to capture the long-distance dependence relationship and context sequence information of the word vector, discard useless word vector sequences, and extract useful word vector sequences.

[0034] According to some embodiments of the present application, the reverse maximum matching algorithm is used to correct the word segmentation to obtain an optimized word segmentation result, including:

[0035] From right to left, match the characters composed of several consecutive phrases in the string to be re-segmented after the first word segmentation with the domain dictionary;

[0036] If the match is successful, segment a new word to obtain the optimized word segmentation result;

[0037] If the match is unsuccessful, remove the leftmost phrase in the first word segmentation result and then match it with the domain dictionary. If the match is successful, segment a new word to obtain the optimized word segmentation result; if the match is unsuccessful, iterate until only one phrase remains in the first word segmentation result to obtain the optimized word segmentation result.

[0038] According to some embodiments of the present application, the entity joint extraction of the candidate words of the device entity model to obtain high-confidence high-frequency words includes:

[0039] Carry out data core description of the candidate words of the device entity model through data elements to form a fixed feature expression form of power information;

[0040] According to the preset data element sorting and device entity extraction experience, combine the standard document number, system classification number and test report type to determine the power system domain factor, and the power system domain factor carries the data source and classification of the candidate words of the device entity model and the serial number of the device entity extraction rule library;

[0041] Determine the feature vector of the candidate words of the device entity model through bidirectional calculation of the long short-term memory network and the power system domain factor, and use the BiLSTM model to extract the key features for identifying the candidate words of the power system device entity model;

[0042] Introduce the feature vector and the key features of the candidate words of the device entity model for combined calculation to identify the device entity name and weight corresponding to the candidate words of the device entity model;

[0043] Use the label speculation of the MFM model to mark the globally optimal sequence of the CRF layer, and convert the hidden state sequence into the best label sequence to realize the entity joint extraction of the candidate words of the device entity model and obtain high-confidence high-frequency words.

[0044] According to some embodiments of the present application, the identification process of the high-confidence high-frequency words to obtain the word with the highest confidence includes:

[0045] Determine the distance between the high-confidence high-frequency words and the device model category;

[0046] Arrange the distances from small to large to obtain a high-frequency word order table with confidence from high to low;

[0047] Clean the negative words in the high-frequency word order table and output the high-frequency word with the highest confidence in the high-frequency word order table to obtain the word with the highest confidence.

[0048] In a second aspect, an embodiment of the present application provides a controller, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor runs the computer program, it executes the method described in the technical solution of the first aspect above.

[0049] Thirdly, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method described in the technical solution of the first aspect as above.

[0050] The information knowledge extraction method, controller, and storage medium based on a power system provided by the embodiments of the present application at least have one of the following advantages or beneficial effects: obtaining an information document of the power system, uniformly converting the unstructured data of the information document of the power system into the txt format, and then cleaning off spaces, line breaks, invalid characters, etc. through data cleaning to improve the uniformity of the unstructured data; classifying, grading, layering, itemizing, and standardizing the unstructured data to achieve classified itemization and standardization processing, and obtaining power system data elements; preprocessing the power system data elements through a character embedding layer based on the RoBERTa model, a BiLSTM model, and a tag speculation CRF layer based on the MFM model, combined with a joint method of power system text word segmentation and part-of-speech tagging to obtain word segments, performing recombination and vectorization processing on the word segments, realizing the vectorization processing of the power system data elements after splitting them into half-sentences, and using the reverse maximum matching algorithm to correct the word segments to obtain an optimized word segmentation result, reducing the influence of out-of-vocabulary words in the power system field on the word segmentation effect. Then, training the optimized word segmentation result based on a power system-specific text word library to obtain a standard power system metadata database, capable of establishing a standard metadata database with unity, fine granularity, and identifiability, and establishing rules for data element information and entity extraction according to the construction direction of the power system to form a mapping rule library for data element information and device entities, providing a good data foundation for knowledge extraction of the power system, enabling the recognition result of device entities to conform to the composition law of inner and outer entity nesting combinations in the actual text, improving the accuracy and efficiency of multi-modal power system data element naming device entity recognition, and at the same time being able to make up for the inefficiency and inaccuracy of current manual sampling inspection. Identifying and filtering device entity category candidate words to obtain high-frequency nouns, performing entity joint extraction on device entity model candidate words, and performing entity joint extraction according to upper entity and relationship concept extraction and lower entity and relationship concept extraction to obtain high-confidence high-frequency vocabulary, and performing recognition processing on the high-confidence high-frequency vocabulary to obtain the vocabulary with the highest confidence, further improving the accuracy of power system data element naming device entity recognition. Finally, generating device entity categories and model labels corresponding to the power system data elements according to the high-frequency nouns and the vocabulary with the highest confidence and storing them in the index, and outputting device entity knowledge according to the index to improve the accuracy of power system data element naming device entity recognition, thereby realizing the accurate recognition and extraction of power system data.

[0051] Other features and advantages of the present application will be described in the subsequent specification, and in part will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings

[0052] Figure 1 is a flowchart of a method for extracting information knowledge based on a power system provided by an embodiment of the present application;

[0053] Figure 2 is a flowchart of a method for identifying and filtering candidate words of device entity categories to obtain high-frequency nouns provided by an embodiment of the present application;

[0054] Figure 3 is a flowchart of a method for preprocessing power system data elements to obtain word segmentation provided by an embodiment of the present application;

[0055] Figure 4 is a flowchart of a method for inputting power system data elements into the character embedding layer of a RoBERTa model to obtain word vectors reflecting context semantics and outputting word vectors provided by an embodiment of the present application;

[0056] Figure 5 is a flowchart of a method for capturing long-distance dependency relationships and context sequence information of word vectors through a BiLSTM model provided by an embodiment of the present application;

[0057] Figure 6 is a flowchart of a method for correcting word segmentation using a reverse maximum matching algorithm to obtain an optimized word segmentation result provided by an embodiment of the present application;

[0058] Figure 7 is a flowchart of a method for entity joint extraction of candidate words of device entity models to obtain high-confidence frequent words provided by an embodiment of the present application;

[0059] Figure 8 is a flowchart of a method for identifying high-confidence frequent words to obtain the word with the highest confidence provided by an embodiment of the present application;

[0060] Figure 9 is a schematic structural diagram of a controller provided by an embodiment of the present application. Detailed Embodiments

[0061] This section will describe in detail the specific embodiments of the present application. The preferred embodiments of the present application are shown in the accompanying drawings. The function of the accompanying drawings is to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present application. However, it should not be construed as a limitation on the protection scope of the present application.

[0062] In the description of the present application, the meaning of "several" is one or more, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number. "Any one" means one or more, and "at least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0063] It should be noted that words such as "set", "installed", "connected", etc. in the embodiments of the present application should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the embodiments of the present application in combination with the specific content of the technical solution. For example, the term "connected" can be a mechanical connection, an electrical connection, or can communicate with each other; it can be directly connected or indirectly connected through an intermediate medium.

[0064] It should be noted that the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0065] Refer to Figure 1 as shown in Figure 1 is a flowchart of an information knowledge extraction method based on a power system provided by an embodiment of the present application. The information knowledge extraction method based on a power system includes but is not limited to steps S100 to S900. Specifically,

[0066] Step S100: Obtain the information document of the power system, read the unstructured data of the information document of the power system, and convert the unstructured data into txt format;

[0067] Step S200: Perform data cleaning, classification and cataloging, and standardization processing on the unstructured data to obtain power system data elements;

[0068] Step S300: Preprocess the power system data elements through the character embedding layer based on the RoBERTa model, the BiLSTM model, and the token speculation CRF layer based on the MFM model, combined with the joint method of power system text word segmentation and part-of-speech tagging, to obtain word segmentation. Recombine and vectorize the word segmentation, and use the reverse maximum matching algorithm to correct the word segmentation to obtain the optimized word segmentation result;

[0069] Step S400: Train the optimized word segmentation result based on the power system specific text library to obtain the standard power system metadata database. According to the construction direction of the power system, establish the rules for data element information and entity extraction, and form the mapping rule library of data element information and device entities to realize the extraction of candidate words for device entity categories;

[0070] Step S500: Identify and filter the candidate words for device entity categories to obtain high-frequency nouns;

[0071] Step S600: Conduct entity joint extraction on the candidate words for device entity models to obtain high-confidence high-frequency vocabulary;

[0072] Step S700: Identify the high-confidence high-frequency vocabulary to obtain the vocabulary with the highest confidence;

[0073] Step S800: Generate the device entity categories and model labels corresponding to the power system data elements according to the high-frequency nouns and the vocabulary with the highest confidence, and store them in the index;

[0074] Step S900: Output the device entity knowledge according to the index.

[0075] In some embodiments of the present application, the information knowledge extraction method based on the power system includes: obtaining the information documents of the power system, reading the unstructured data of the information documents of the power system, and through document reading, uniformly converting unstructured data such as word, pdf, txt, images, etc. into txt format; then through data cleaning, removing spaces, line breaks, invalid characters, etc., to achieve data cleaning of the unstructured data, improve the uniformity of the unstructured data, and thus play an important role in standardizing data, reducing sharing costs, improving sharing efficiency, supporting efficient data processing and in-depth application; classifying, sorting, layering, itemizing, and standardizing the unstructured data in the power system field to achieve classified sorting, itemizing, and standardizing processing, and obtaining power system data elements; through the character embedding layer based on the RoBERTa model, the BiLSTM model, and the tag speculation CRF layer based on the MFM model, combined with the joint method of power system text word segmentation and part-of-speech tagging, preprocessing the power system data elements to obtain word segments, performing recombination and vectorization processing on the word segments, realizing the vectorization processing of the power system data elements after splitting them into half-sentences, using the reverse maximum matching algorithm to correct the word segments, and obtaining an optimized word segmentation result, reducing the impact of out-of-vocabulary words in the power system field on the word segmentation effect. After that, training the optimized word segmentation result based on the power system-specific text word library to obtain a standard power system metadata database, which can establish a standard metadata database with unity, fine granularity, and identifiability, and according to the construction direction of the power system, establish rules for data element information and entity extraction, forming a mapping rule library of data element information and device entities, providing a good data foundation for the knowledge extraction of the power system, enabling the recognition results of device entities to conform to the composition rules of the inner and outer layer entity nesting combinations in the actual text, improving the accuracy and efficiency of multi-modal power system data element naming device entity recognition, and at the same time being able to make up for the inefficiency and inaccuracy of the current manual sampling inspection. Identifying and filtering the candidate words of the device entity category to obtain high-frequency nouns, performing entity joint extraction on the candidate words of the device entity model, and performing entity joint extraction according to the upper entity and relationship concept extraction, and the lower entity and relationship concept extraction to obtain high-confidence high-frequency vocabulary, and performing recognition processing on the high-confidence high-frequency vocabulary to obtain the vocabulary with the highest confidence, further improving the accuracy of power system data element naming device entity recognition. Finally, generating the device entity category and model labels corresponding to the power system data elements according to the high-frequency nouns and the vocabulary with the highest confidence and storing them in the index, and outputting the device entity knowledge according to the index, improving the accuracy of power system data element naming device entity recognition, so as to achieve the accurate recognition and extraction of power system data.

[0076] In some embodiments of the present application, a thesaurus of the power industry is obtained from industry experts, national standards, Baidu Input Method, Sogou Input Method, etc. After being screened by artificial intelligence technology, the optimized word segmentation results are trained based on the special text thesaurus of the power system to obtain a standard power system meta-database. According to the construction direction of the power system, rules for data element information and entity extraction are established. Then, the data elements of the power system are vectorized after being split into half-sentences and compared with the thesaurus of the power industry. The words with a cosine value of the word vector greater than 0 are obtained to form a mapping rule library of data element information and equipment entities, making up for the inefficiency and inaccuracy of the current manual screening.

[0077] The standard power system meta-database of the present application includes but is not limited to the following information: Power system information resources: classification of power system information resources, name of information resources, provider of information resources, abstract of information resources; Power system standard documents: release date, release agency, standard name, attachment upload, standard number; Power system data elements: information resources to which they belong, definition description, Chinese name of data elements, English name of data elements, name of code set, remarks, number, data element field, provider of data elements, data format, data type, original data type; Power system value range: name of code set, description, coding rule, code name. A standard power system meta-database with unity, fine granularity, and relevance is established, and according to the construction direction of the power system, rules for data element information and entity extraction are established to form a mapping rule library of data element information and equipment entities, providing a good data foundation for knowledge extraction of the power system.

[0078] Facing the massive and unstructured power system data, the complex entities of the power system data lack obvious features and are complex and diverse in syntactic structure and part-of-speech composition. On the other hand, due to the complex structure of the Chinese language itself, it is difficult to identify word boundaries, and word boundaries are generally also the boundaries of named entities. The span-based method widely used in existing models to identify nested entities shows fuzziness in entity boundary detection, affecting the recognition performance. Therefore, there are still many problems in the recognition of Chinese named equipment entities for long texts. The present application provides a method that can improve the accuracy and efficiency of recognizing named equipment entities in multiple power system data.

[0079] Refer to Figure 2 as shown Figure 2 is a flowchart of a method for identifying and filtering candidate words of equipment entity categories to obtain high-frequency nouns provided by an embodiment of the present application. The method for identifying and filtering candidate words of equipment entity categories to obtain high-frequency nouns includes but is not limited to steps S510 to S520. Specifically,

[0080] Step S510: Based on the knowledge of power system standards and specifications, form a basic rule knowledge library for the application scenarios, entity boundaries, and semantics of equipment entities;

[0081] Step S520: Based on the equipment entity application scenario, entity boundary, and semantic basic rule knowledge base, fuse the Chinese entity recognition method for equipment entity nesting rules, and transform the recognition task of equipment entities into a joint training task of equipment entity boundary recognition and boundary head-tail relationship recognition;

[0082] Step S530: During the decoding process of the joint training task, combine the entity nesting rules of the power system to identify and filter the decoding results, so that the recognition results conform to the composition rules of the inner and outer layer entity nesting combinations in the text;

[0083] Step S540: Terminate and standardize the mapping of the equipment entity category candidate words according to the terminology standard of the power system to obtain high-frequency nouns.

[0084] In some embodiments of the present application, for the equipment entities in the power system field, which are mostly nested entities and have characteristics such as long characters and strong context relevance, the collected original power system data elements are cleaned and processed, and multi-source heterogeneous data fusion for power system text entity recognition is carried out. Based on knowledge such as power system standards, expert consensus, and power system specifications, a basic rule knowledge base for equipment entity application scenarios, entity boundaries, and semantics is formed; and a Chinese entity recognition method is designed by fusing equipment entity nesting rules, and the recognition task of equipment entities is transformed into a joint training task of equipment entity boundary recognition and boundary head-tail relationship recognition. During the decoding process, the decoding results are filtered by combining the equipment entity nesting rules summarized from actual power system texts, so that the recognition results can conform to the composition rules of the inner and outer layer entity nesting combinations in the actual text, improving the accuracy and efficiency of power system data element naming equipment entity recognition. At the same time, for the extracted power system data elements, the important fields related to the power system are terminologized and standardized according to the relevant terminology standards in the industry to obtain high-frequency nouns, further achieving the purpose of data structuring, improving the accuracy and efficiency of multi-modal power system data element naming equipment entity recognition, and forming a structured power system data knowledge base applicable to the key features in multiple dimensions of power system operation and maintenance and evaluation.

[0085] In the development process of Chinese word segmentation and part-of-speech tagging in the field of power systems, there are two major problems that have not been completely solved: word segmentation ambiguity and out-of-vocabulary word recognition. 1) The phenomenon of word segmentation ambiguity: In the real text environment, the phenomenon of word segmentation ambiguity is common. The root cause lies in the flexibility of Chinese characters, which can freely appear in any position in the Chinese text, forming diverse phrase combinations, which greatly increases the complexity of lexical boundary definition. If a natural language processing system has insufficient ambiguity resolution ability, it will be unable to cope when processing large-scale corpora, and will even lead to frequent word segmentation errors. Such errors will directly affect the accuracy of Chinese automatic word segmentation and even the accuracy of the entire syntactic analysis. 2) The problem of out-of-vocabulary words: The so-called out-of-vocabulary words refer to those words that appear in the text to be processed but are not included in the existing dictionary. In actual research, we collectively refer to all words that appear in the test corpus but are not involved in the training corpus as out-of-vocabulary words. Such words widely cover various proper nouns such as device names, place names, and organization names, and their processing difficulty cannot be ignored.

[0086] Aiming at the problems of lack of domain corpus and difficulty in correctly segmenting out-of-vocabulary words such as professional nouns and phrases in the power system text segmentation process, this application proposes a method for preprocessing power system data elements to obtain word segmentation.

[0087] Referring to Figure 3 as shown, Figure 3 is a flowchart of a method for preprocessing power system data elements to obtain word segmentation provided by an embodiment of this application. The method for preprocessing power system data elements to obtain word segmentation includes but is not limited to steps S310 to S340. Specifically,

[0088] Step S310: Input the power system data element into the character embedding layer of the RoBERTa model to obtain a word vector reflecting the context semantics and output the word vector;

[0089] Step S320: Input the word vector into the BiLSTM model, and capture the long-distance dependence relationship and context sequence information of the word vector through the BiLSTM model to discard the useless word vector sequence and extract the useful word vector sequence;

[0090] Step S330: Input the useful word vector sequence into the label speculation CRF layer of the MFM model, and use the label speculation CRF layer of the MFM model to jointly model the label sequence, and calculate the conditional probability of the input useful word vector sequence output in the label sequence;

[0091] Step S340: Determine the highest conditional probability of the output, and select the useful word vector sequence corresponding to the highest conditional probability as the word segmentation.

[0092] The embodiment of this application provides a method for jointly segmenting and part-of-speech tagging power system text based on the RoBERTa model, BiLSTM model, and MFM model to preprocess power system data elements to obtain segmented words, including: using the RoBERTa model with transfer learning ability and text feature representation ability as the feature representation layer, inputting the power system data elements into the character embedding layer of the RoBERTa model to obtain word vectors reflecting context semantics and output the word vectors, combining with the BiLSTM model to extract global and local features of the power system data elements for text segmentation, capturing the long-distance dependence relationship and context sequence information of the word vectors through the BiLSTM model to discard useless word vector sequences and extract useful word vector sequences. Then, input the useful word vector sequences into the tag speculation CRF layer of the MFM model. The tag speculation CRF layer of the MFM model has the ability to process overlapping features and long-distance dependencies and can handle the tag bias problem well. Use the tag speculation CRF layer of the MFM model to jointly model the tag sequence, calculate the conditional probability of the input useful word vector sequences output in the tag sequence, determine the highest conditional probability output, and select the useful word vector sequences corresponding to the highest conditional probability as the segmented words to improve the accuracy and efficiency of naming device entity recognition of power system data elements.

[0093] Refer to Figure 4 as shown Figure 4 is a flowchart of a method for inputting power system data elements into the character embedding layer of the RoBERTa model to obtain word vectors reflecting context semantics and output the word vectors. The method for inputting power system data elements into the character embedding layer of the RoBERTa model to obtain word vectors reflecting context semantics and output the word vectors includes but is not limited to steps S311 to S314. Specifically,

[0094] Step S311: Calculate the word vectors of the power system data elements with the first parameter matrix, the second parameter matrix, and the third parameter matrix respectively to obtain a first vector, a second vector, and a third vector with the same dimension as the word vectors;

[0095] Step S312: Construct an attention matrix according to the first vector, the second vector, and the third vector;

[0096] Step S313: Perform mappings of different linear transformations on the first vector, the second vector, and the third vector, and splice them through the attention matrix to obtain a multi-head attention matrix;

[0097] Step S314: Connect the multi-head attention matrix to a feed-forward neural network, perform encoding processing on the power system data elements through linear transformation and the Relu activation function, and perform preliminary segmentation processing on the power system data elements using a Chinese word segmentation algorithm to obtain word vectors reflecting context semantics and output the word vectors.

[0098] The character embedding layer of the RoBERTa model is the core of the RoBERTa model. The main function of this layer is to calculate the correlation between different words in a sentence. The overall framework of the character embedding layer is stacked by multiple encoder structures of Transformer. Each encoder consists of two parts: multi-head attention and a feed-forward neural network. Among them, multi-head attention expands the ability of the model to focus on subspaces. Each multi-head attention structure consists of h self-attention structures. In the embodiments of the present application, power system data elements are input into the character embedding layer of the RoBERTa model. The word vectors of the power system data elements are respectively calculated with the first parameter matrix WQ, the second parameter matrix WK, and the third parameter matrix WV to obtain the first vector Q, the second vector K, and the third vector V with the same dimension as the word vector. An attention matrix is constructed according to the first vector Q, the second vector K, and the third vector V. Then, the attention matrix is mapped through h different linear transformations on the first vector Q, the second vector K, and the third vector V. Connecting the multi-head attention matrix to the feed-forward neural network can effectively increase the non-linear fitting ability of the model. The power system data elements are encoded through linear transformation and the Relu activation function, and the power system data elements are initially segmented using the Chinese word segmentation algorithm to obtain word vectors reflecting the context semantics and output the word vectors. The word vectors are used as the input information of the BiLSTM model for subsequent semantic encoding.

[0099] In some embodiments of the present application, constructing the attention matrix according to the first vector Q, the second vector K, and the third vector V is implemented through the following formula:

[0100]

[0101] Among them, Attention(Q, K, V) is the attention matrix, Q is the first vector, K is the second vector, V is the third vector, and d is the dimension of the word vector of the input power system data element.

[0102] Referring to Figure 5 as shown, Figure 5 is a flowchart of a method for capturing the long-distance dependence relationship and context sequence information of word vectors through a BiLSTM model provided by an embodiment of the present application. The method for capturing the long-distance dependence relationship and context sequence information of word vectors through a BiLSTM model includes but is not limited to steps S321 to S323. Specifically,

[0103] Step S321: Input the word vector into the forget gate of the BiLSTM model. Through the forget gate, determine the information that needs to be discarded in the previous neuron. The input is the current word vector and the previous word vector. Map the cell state of the previous neuron to 0 to 1, where 0 means complete deletion and 1 means complete retention, to obtain the result of the forget gate;

[0104] Step S322: Determine whether to record a new word vector into the neurons of the BiLSTM model through the input gate. Use the Sigmoid activation function of the input layer to determine the word vectors to be updated and update the neuron state to determine the memory information to be recorded, obtaining the word vector sequence to be memorized.

[0105] Step S323: Use the Sigmoid activation function of the output gate to determine the part of the output neuron state. Process the neuron state through the tanh function and multiply it by the output of the Sigmoid activation function to obtain a hidden layer state sequence of the same length as the word vector sequence, so as to capture the long-distance dependencies and context sequence information of the word vectors, discard the useless word vector sequences, and extract the useful word vector sequences.

[0106] The BiLSTM model is an improved Recurrent Neural Network (RNN) model that can be used to capture long-distance dependencies and context sequence information. The BiLSTM model includes a forget gate, an input gate, and an output gate. By introducing a gating mechanism, it overcomes the problem of vanishing or exploding gradients caused by overly long text sequences in traditional RNNs. The forget gate determines which information needs to be discarded, the input gate determines which information is retained by the memory unit, and the output gate determines which information is output and enters the next moment's loop iteration. The algorithm implementation process is as follows:

[0107] First step: Calculate the forget gate and select the information to be forgotten: Input the word vector into the forget gate of the BiLSTM model. Determine the information to be discarded in the previous neuron through the forget gate. The inputs are the current word vector and the previous word vector. Map the previous neuron cell state to 0 - 1, where 0 means complete deletion and 1 means complete retention, obtaining the result of the forget gate.

[0108] The BiLSTM model determines the information to be discarded in the previous neuron through the forget gate. The inputs are the outputs of the current word vector and the previous word vector. Map the previous neuron cell state to 0 - 1, where 0 means complete deletion and 1 means complete retention. The result of the forget gate at time t is:

[0109]

[0110] In the formula, W fx , W fh , W fc are the matrix weights multiplied by the forget gate, the input x t , the intermediate output h t-1 and the neuron state c t-1 of the BiLSTM model at time t - 1 respectively, and b fis the bias term of the forget gate, and σ is the sigmoid activation function. x t is the input data at time t. h t-1 is the output data at time t-1.

[0111] Step 2: Calculate the input gate and select the information to be remembered. Determine whether to record the new word vector into the neurons of the BiLSTM model through the input gate. The Sigmoid activation function of the input layer decides the word vector to be updated to update the neuron state, so as to determine the memory information to be recorded and obtain the word vector sequence to be remembered.

[0112] The input gate decides whether to record new information into the neurons of this BiLSTM model. The input is the output of the current word vector and the previous word vector. The Sigmoid activation function of the input layer decides the value i t to be updated, and the neuron state is updated to obtain c t through the following formula:

[0113] i t = σ(W ix x t + W ih h t-1 + W ic c t-1 + b i );

[0114] c t = f t c t-1 + i t tanh(W cx x t + W ch h t-1 + b c );

[0115] In the formula, W ix , W ih , W ic are the matrix weights of the input gate multiplied by the input x t and the intermediate output h t-1 and the neuron state c t-1 of the LSTM at time t-1 respectively. b f is the bias term of the forget gate. W cx , W ch are the matrix weights of the input node multiplied by the input x t and the intermediate output h t-1 respectively. b c is the bias term of the input node. c t is the neuron state of the BiLSTM model at time t. tanh is the activation function. it is an n-dimensional vector at time t in the input gate, taking a number between 0 and 1.

[0116] Step 3: Calculate the output gate and the hidden layer state at the current moment, and select the output value. Determine the part of the output neuron state through the Sigmoid activation function of the output gate, process the neuron state through the tanh function, and multiply it by the output of the Sigmoid activation function to obtain a hidden layer state sequence of the same length as the word vector sequence, so as to capture the long-distance dependence relationship and context sequence information of the word vector, discard the useless word vector sequence, and extract the useful word vector sequence.

[0117] The output gate is used to determine the output information, and determine the part o of the output neuron state through the Sigmoid activation function of the output layer t , process the neuron state through the tanh function, and multiply it by the output of the Sigmoid gate to obtain the final output h t , finally, a hidden layer state sequence {h1, h2, h3,..., h n} of the same length as the word vector can be obtained.

[0118] o t = σ(W ox x t + W oh h t-1 + W oc c t + b o );

[0119] h t = o t tanh(c t );

[0120] In the formula, W ox , W oh , W oc are the matrix weights of the input gate multiplied by the input x t and the intermediate output h t-1 multiplied by the neuron state c t-1 of the BiLSTM model at time t-1 respectively. b o is the bias matrix of the output gate. o t is the generation matrix of logistic regression after the fully connected layer. h t is the output content of the current state at time t.

[0121] From the working principle of the BiLSTM model, the BiLSTM model uses memory gates, forget gates, and output gates to process information, can discard some useless information, enhances the memory of neurons, is suitable for processing long time relationship data, and solves the long-term dependence problem.

[0122] The word segmentation task is usually regarded as a sequence labeling task. There is a strong dependence between the label results of the front and back characters. In the problems of character-based part-of-speech tagging and Chinese word segmentation, the collocation relationship between adjacent tags must be considered. Therefore, we cannot make tagging decisions independently, but use the MFM model to jointly model the tagging sequence. The MFM model is an undirected graph model used to calculate the conditional probability of the output of random variables given the input random variables. It combines the characteristics of the hidden Markov model and the maximum entropy model, has the ability to handle overlapping features and long-distance dependencies, and can handle the tagging bias problem well.

[0123] In the embodiment of the present application, the useful word vector sequence is input into the tagging inference CRF layer of the MFM model, and the tagging inference CRF layer of the MFM model is used to jointly model the tagging sequence, and calculate the conditional probability of the input useful word vector sequence output in the tagging sequence.

[0124] The calculation process of the tagging inference CRF layer based on the MFM model is as follows: For the output sequence X=(x1, x2,..., x t )(which is also the input sequence of the CRF layer) of the BiLSTM layer of the model, assume that P∈R n×k is the output score matrix of the BiLSTM, where n is the number of words and k is the number of tags, and the matrix element P ij is the output score of the i-th word under the j-th tag. For the useful word vector sequence Y=(y1, y2,..., y t ), the total score of this tag sequence is:

[0125]

[0126] In the formula, A is the transition score matrix, and A i,j represents the transition score from tag i to tag j, P i,j is the output score of the i-th word under the j-th tag, n is the number of words, k is the number of tags, the size of A is k + 2, X=(x1, x2,..., x t ) is the input sequence of the CRF layer, Y=(y1, y2,..., y t ) is the useful word vector sequence, and s(X, Y) represents the scoring of each tagging sequence.

[0127] The probability generated by the useful word vector sequence Y is represented by the following formula:

[0128]

[0129] In the formula, X=(x1, x2,..., x t ) is the input sequence of the CRF layer, Y=(y1, y2,..., y tis a useful word vector sequence, s(X, Y) represents the score of each tag sequence, and Y% is the true tag sequence; Y X is all possible tag sequences, P(Y|X) represents the probability of each labeled tag. The tag sequence with the highest probability is selected as the conditional probability output in the tag sequence to obtain the predicted result output.

[0130] Refer to Figure 6 as shown Figure 6 is a flowchart of a method for correcting word segmentation using the reverse maximum matching algorithm to obtain an optimized word segmentation result provided by an embodiment of the present application. The method for correcting word segmentation using the reverse maximum matching algorithm to obtain an optimized word segmentation result includes but is not limited to steps S350 to S370. Specifically,

[0131] Step S350: From right to left, match the characters composed of several consecutive word groups in the string to be re-segmented after the first word segmentation with the domain dictionary;

[0132] Step S360: If the match is successful, split out a new word to obtain an optimized word segmentation result;

[0133] Step S370: If the match is unsuccessful, remove the leftmost word group in the first word segmentation result and then match it with the domain dictionary. If the match is successful, split out a new word to obtain an optimized word segmentation result; if the match is unsuccessful, iterate until only one word group remains in the first word segmentation result to obtain an optimized word segmentation result.

[0134] The method of the present application for correcting word segmentation using the reverse maximum matching algorithm to obtain an optimized word segmentation result reduces the influence of out-of-vocabulary words in the power system field on the word segmentation effect.

[0135] In some embodiments of the present application, assume that the string obtained from the first word segmentation result is S1, the matching result is S2 (initially S2 is an empty string), and the maximum word length in the domain dictionary D is N. The algorithm steps for reverse maximum matching of S1 are as follows:

[0136] Step 1. Initialize the algorithm parameters and input the string S1 to be re-segmented;

[0137] Step 2. Set the maximum word length of the domain dictionary D to N;

[0138] Step 3. Start from the rightmost of S1 to take the candidate string W.

[0139] Step 4. Match the candidate string W with the domain dictionary D. If the match is successful, go to step 6; if the match is unsuccessful, execute step 5;

[0140] Step 5. Delete the leftmost original phrase of the candidate string W and assign it to S2; determine whether the candidate string W is a single phrase. If it is, go to Step 6; otherwise, return to Step 4.

[0141] Step 6. S2 = "" + W + S2 (where "" is two spaces), and remove the original phrase contained in W from the rightmost of S2;

[0142] Step 7. Output the matching result S2.

[0143] Then, use the word segmentation automatic optimization algorithm to merge and optimize the "scattered words". Regarding the deviation, to avoid the expansion of the deviation, transfer learning is introduced to enhance the adaptability of the system. The embodiment of this application combines the domain corpus annotation data of the power industry special field for domain word segmentation transfer learning. According to the data similarity between the power system domain word segmentation corpus and the target domain corpus, different levels of network transfer are performed on the tagging speculation CRF layer based on the RoBERTa model, BiLSTM model, and MFM model to solve the domain adaptability problem of word segmentation, and finally output the optimized word segmentation result.

[0144] Refer to Figure 7 as shown in Figure 7 is a flowchart of a method for entity joint extraction of device entity model candidate words to obtain confidence high-frequency words provided by the embodiment of this application. The method for entity joint extraction of device entity model candidate words to obtain confidence high-frequency words includes but is not limited to steps S610 to S650. Specifically,

[0145] Step S610: Corely describe the device entity model candidate words through data elements to form a fixed feature expression form of power information;

[0146] Step S620: Determine the power system domain factors according to the preset data element sorting and device entity extraction experience, in combination with the standard document number, system classification number, and test report type. The power system domain factors carry the data source and classification of the device entity model candidate words and the serial number of the device entity extraction rule library;

[0147] Step S630: Determine the feature vector of the device entity model candidate words through bidirectional calculation of the long short-term memory network and the power system domain factors, and use the BiLSTM model to extract the key features for identifying the power system device entity model candidate words;

[0148] Step S640: Introduce the feature vector and the key features of the device entity model candidate words for combined calculation to identify the device entity name and weight corresponding to the device entity model candidate words;

[0149] Step S650: Speculate the globally optimal sequence of CRF layer tags through the tags of the MFM model, and convert the hidden state sequence into the best tag sequence, so as to realize the entity joint extraction of the candidate words of the device entity model and obtain the high-confidence high-frequency words.

[0150] Facing the massive and unstructured power system data, the complex entities of the multimodal power system data lack obvious features, and are complex and diverse in syntactic structure and part-of-speech composition. The embodiment of the present application provides a method for entity joint extraction of candidate words of device entity models to obtain high-confidence high-frequency words, improving the accuracy of device entity recognition in the naming of power system data elements.

[0151] Keywords are identified based on multiple data sources such as data standard documents, information system databases, and power detection reports. At this time, Chinese word segmentation technology for the power system field will be used to remove stop words and other operations, and finally the key equipment entity model candidate words of the data source will be left. After that, the data core description of the equipment entity model candidate words is performed through data elements to form different types and multi-channel fixed feature expressions of power information. At the same time, based on the experience of data element combing and equipment entity extraction, the data element and entity extraction rule base are continuously improved to guide the generation of power system field factors. According to the preset data element combing and equipment entity extraction experience, the power system field factor is determined in combination with the standard document number, system classification number and detection report type. The power system field factor carries the data source and classification of the equipment entity model candidate words and the equipment entity extraction rule base number, which is conducive to the expression of power data information knowledge, efficient identification of abnormal values ​​and sources, and provides a basis for the expression of power data information knowledge; the power system data element and equipment entity extraction rule base number is used to guide the confirmation of entities in the same field during entity extraction, which improves the efficiency of entity extraction. The vector generated by the RoBERTa model at the character embedding layer is generated by superposition of word vectors, sentence vectors and position vectors. The word vector is obtained by converting the input data by querying the word vector table; the sentence vector represents the text information of the sentence and is used to distinguish different sentences; the position vector is obtained by encoding the position information corresponding to each word, which can distinguish the semantic information of the text at different positions. The output layer is the word vector that can reflect the contextual semantics obtained after processing by the encoding layer. These word vector sequences are used as the input information of the BiLSTM model for subsequent semantic encoding. The feature vector of the equipment entity model candidate word is determined by bidirectional calculation between the long short-term memory network and the power system domain factor. The key features of the power system equipment entity model candidate word recognition are extracted using the BiLSTM model. The feature vector and the key features of the equipment entity model candidate word are introduced for combined calculation to identify the equipment entity name and weight corresponding to the equipment entity model candidate word; the global optimal sequence of CRF layer tags is inferred through the MFM model tag, and the hidden state sequence is converted into the optimal tag sequence to realize the entity joint extraction of the equipment entity model candidate word and obtain the confidence high-frequency vocabulary. This application enables the recognition results of equipment entities to conform to the composition rules of nested combinations of inner and outer entities in actual texts, improves the accuracy and efficiency of multimodal power system data element naming equipment entity recognition, and thus realizes accurate recognition and extraction of power system data.

[0152] Various events in the power system have certain scopes and themes. Therefore, the method for extracting equipment entity relationships in this application is different from the traditional method. This application uses the Word2vec model and bootstrapping method to complete the extraction of entity relationships. Word2vec processes text content into a vector space form through model training, so that the similarity between a single word and other words can be obtained. There are two main models of Word2vec, namely the CBOW model and the improved Skip-gram model. The CBOW model predicts the generation probability of the current word based on the words around the current word, and the Skip-gram model predicts the surrounding words based on the current word, thereby extracting relational knowledge. It is mainly divided into upper-level entity and relationship concept extraction, and lower-level entity and relationship concept extraction.

[0153] The method of extracting the concept of a superordinate entity based on mixed lexical rules is as follows:

[0154] When extracting the superordinate entity concept, a hybrid lexical rule is proposed to extract the relation extraction (RE) part. In the sentence containing the relational feature word "is a", the RE part has a simple structure, so it is relatively easy to obtain the superordinate entity concept from RE. The superordinate concept is mainly extracted based on the lexical rule analysis method. The specific extraction rules are as follows:

[0155] 1) If the RE does not contain the character "的", the noun following the relational feature word is directly obtained through lexical analysis;

[0156] 2) If the RE contains the character "的", the noun following the character "的" with the largest position number in the RE is directly extracted through lexical analysis;

[0157] 3) Mixed lexical rules are used to extract RE. If the RE contains "的", the noun following "的" is extracted; otherwise, the noun following the feature word is extracted.

[0158] The method of extracting sub-entity concepts based on mixed syntactic rules is as follows:

[0159] The extraction of subordinate entity concepts is different from that of superordinate entity concepts. By matching the "is a" relationship feature sentence, the LE part of the subordinate entity concept may contain multiple commas. After analysis, most sentences containing multiple commas contain multiple subject-predicate structures. This application analyzes the LE part from a syntactic perspective, thereby summarizing some syntactic relationship features and accurately extracting subordinate entity concepts.

[0160] Reference Figure 8 As shown, Figure 8It is a flowchart of a method provided by an embodiment of the present application for identifying and processing high-frequency confidence words to obtain the word with the highest confidence. The method for identifying and processing high-frequency confidence words to obtain the word with the highest confidence includes but is not limited to steps S710 to S730. Specifically,

[0161] Step S710: Determine the distance between the high-frequency confidence word and the device model category;

[0162] Step S720: Arrange the distances from small to large to obtain a high-frequency word order list with confidence levels from high to low;

[0163] Step S730: Clean the negative words in the high-frequency word order list and output the high-frequency word with the highest confidence in the high-frequency word order list to obtain the word with the highest confidence.

[0164] In some embodiments of the present application, entity joint extraction is performed on the candidate words of the device entity model to obtain high-frequency confidence words, and the high-frequency confidence words are identified and processed to obtain the word with the highest confidence. The method for identifying and processing high-frequency confidence words to obtain the word with the highest confidence includes: First, calculate the semantic distance, similarity score, or other quantitative indicators of correlation between each high-frequency confidence word and the device model category to evaluate the correlation between each high-confidence high-frequency word and the device model. After calculating the distances between all high-frequency confidence words and the device model, sort the high-frequency confidence words according to these distance values, arrange the distances from small to large to obtain a high-frequency word order list with confidence levels from high to low, that is, the high-frequency confidence word with the smallest distance value (i.e., the word most relevant to the device model) will be ranked at the top of the high-frequency word order list, thus forming an order list with confidence levels from high to low. Since negative words may reduce the confidence of words because they usually represent negative or opposite meanings. Therefore, in the present application, negative words are cleaned and removed from the high-frequency word order list. After cleaning, select the high-frequency word with the highest confidence from the remaining words as the final result, and output the high-frequency word with the highest confidence to obtain the word with the highest confidence.

[0165] Through this method, keywords in high-frequency word queries in the power system can be identified more accurately, thereby providing more relevant content or services, improving the accuracy of device entity recognition in power system data element naming, and realizing accurate recognition and extraction of power system data.

[0166] Refer to Figure 9 as shown Figure 9FIG. 0 is a schematic structural diagram of a controller 1000 provided by an embodiment of the present application, including a processor 1001, which may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the method provided by the embodiment of the present application; a memory 1002, which may be implemented in the form of a read-only memory 1002 (ROM), a static storage device, a dynamic storage device, or a random access memory 1002 (RAM), etc. The memory 1002 may store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the embodiments of the present application; an input / output interface 1003, which is used to implement information input and output; a communication interface 1004, which is used to implement communication interaction between this device and other devices, and may implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.); a bus, which transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004); wherein the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus.

[0167] Those of ordinary skill in the art will appreciate that all or some of the steps and systems disclosed above in the methods can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer-readable storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0168] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application may be realized and attained by the structure particularly pointed out in the specification, the claims, as well as the drawings.

Claims

1. A method for extracting information knowledge based on a power system, characterized in that: include: Acquire information documents of the power system, read unstructured data of the information documents of the power system, and convert the unstructured data into a txt format; Cleaning, classifying, cataloguing and standardizing the unstructured data to obtain power system data elements; Preprocess the power system data element through a character embedding layer based on the RoBERTa model, a BiLSTM model, and a tag inference CRF layer based on the MFM model, combined with a power system text segmentation and part-of-speech tagging joint method, to obtain segmentation, reorganize and vectorize the segmentation, and use an inverse maximum matching algorithm to correct the segmentation to obtain an optimized segmentation result; The optimized word segmentation results are trained based on a special text word library for the power system to obtain a standard power system metadata library, and according to the construction direction of the power system, rules for extracting data metadata information and entities are established to form a mapping rule library between data metadata information and equipment entities, so as to realize the extraction of candidate words for equipment entity categories; Identifying and filtering candidate words for the device entity category to obtain high-frequency nouns; Perform entity joint extraction on the candidate words of the device entity model to obtain high-frequency words with high confidence; Performing recognition processing on the high-frequency words with the confidence level to obtain the words with the highest confidence level; Generating equipment entity category and model labels corresponding to the power system data element according to the high-frequency nouns and the words with the highest confidence and storing them in an index; The device entity knowledge is output according to the index.

2. The information knowledge extraction method based on the power system according to claim 1 is characterized in that: The device entity category candidate words are identified and filtered to obtain high-frequency nouns including: Form a knowledge base of basic rules on equipment entity application scenarios, entity boundaries, and semantics based on power system standards and specifications; Based on the equipment entity application scenario and the basic rule knowledge base of entity boundary and semantics, the equipment entity nesting rule Chinese entity recognition method is integrated to transform the equipment entity recognition task into a joint training task of equipment entity boundary recognition and boundary head-tail relationship recognition; In the decoding process of the joint training task, the decoding results are identified and filtered in combination with the entity nesting rules of the power system, so that the recognition results conform to the composition rules of the inner and outer entity nesting combinations in the text; According to the terminology standard of the power system, the candidate words of the equipment entity category are terminologically converted and standardized to obtain high-frequency nouns.

3. The method for extracting information knowledge based on the power system according to claim 1, characterized in that: The power system data element is preprocessed by combining the power system text word segmentation and part-of-speech tagging joint method through the character embedding layer based on the RoBERTa model, the BiLSTM model, and the tag inference CRF layer based on the MFM model to obtain word segmentation, including: Input the power system data element into the character embedding layer of the RoBERTa model to obtain a word vector reflecting contextual semantics and output the word vector; Inputting the word vector into the BiLSTM model, capturing the long-distance dependency and context sequence information of the word vector through the BiLSTM model, so as to discard useless word vector sequences and extract useful word vector sequences; Input the useful word vector sequence into the tag inference CRF layer of the MFM model, use the tag inference CRF layer of the MFM model to jointly model the tag sequence, and calculate the conditional probability of the input useful word vector sequence being output in the tag sequence; The highest conditional probability of the output is determined, and a useful word vector sequence corresponding to the highest conditional probability is selected as the word segmentation.

4. The method for extracting information knowledge based on the power system according to claim 3 is characterized in that: The step of inputting the power system data element into the character embedding layer of the RoBERTa model, obtaining a word vector reflecting contextual semantics and outputting the word vector comprises: Calculating the word vector of the power system data element with the first parameter matrix, the second parameter matrix and the third parameter matrix respectively to obtain a first vector, a second vector and a third vector with the same dimension as the word vector; constructing an attention matrix according to the first vector, the second vector and the third vector; Perform different linear transformation mappings on the first vector, the second vector, and the third vector, and concatenate them through the attention matrix to obtain a multi-head attention matrix; The multi-head attention matrix is ​​connected to a feedforward neural network, the power system data element is encoded through linear transformation and Relu activation function, and the power system data element is preliminarily segmented using a Chinese word segmentation algorithm to obtain a word vector that reflects contextual semantics and output the word vector.

5. The method for extracting information knowledge based on the power system according to claim 3 is characterized in that: The BiLSTM model includes a forget gate, an input gate, and an output gate. The word vector is input into the BiLSTM model, and the long-distance dependency and context sequence information of the word vector are captured by the BiLSTM model to discard useless word vector sequences and extract useful word vector sequences, including: The word vector is input into the forget gate of the BiLSTM model, and the information to be discarded in the previous neuron is determined by the forget gate. The input is the current word vector and the previous word vector, and the previous neuron cell state is mapped to 0-1, where 0 represents complete deletion and 1 represents complete retention, to obtain the result of the forget gate; Determine whether to record a new word vector into the neuron of the BiLSTM model through the input gate, determine the word vector that needs to be updated through the input layer Sigmoid activation function to update the neuron state, so as to determine the memory information that needs to be recorded, and obtain the word vector sequence that needs to be memorized; The part of the output neuron state is determined by the Sigmoid activation function of the output gate, the neuron state is processed by the tanh function, and multiplied by the output of the Sigmoid activation function to obtain a hidden layer state sequence with the same length as the word vector sequence to capture the long-distance dependency and context sequence information of the word vector, so as to discard useless word vector sequences and extract useful word vector sequences.

6. The method for extracting information knowledge based on the power system according to claim 3 is characterized in that: The word segmentation is corrected by using the reverse maximum matching algorithm to obtain an optimized word segmentation result, including: Match the characters consisting of several consecutive phrases in the string to be re-segmented after the first segmentation with the domain dictionary from right to left; If the match is successful, new words are segmented to obtain optimized word segmentation results; If the match is unsuccessful, the leftmost phrase in the first word segmentation result is removed and matched with the domain dictionary. If the match is successful, new words are segmented to obtain an optimized word segmentation result. If the match is unsuccessful, it is iterated until only a single phrase remains in the first word segmentation result to obtain an optimized word segmentation result.

7. The method for extracting information knowledge based on the power system according to claim 1, characterized in that: The entity joint extraction of the device entity model candidate words to obtain high-frequency words with high confidence includes: The candidate words of the equipment entity model are described in the data core through data elements to form a fixed characteristic expression form of power information; According to the preset data element combing and equipment entity extraction experience, the power system domain factor is determined in combination with the standard document number, system classification number and test report type. The power system domain factor carries the data source and classification of the equipment entity model candidate word and the equipment entity extraction rule base number; The feature vector of the candidate words for the equipment entity model is determined through bidirectional calculation between the long short-term memory network and the power system domain factor, and the key features for identifying the candidate words for the equipment entity model of the power system are extracted using the BiLSTM model. The feature vector is introduced and the key features of the device entity model candidate word are combined and calculated to identify the device entity name and weight corresponding to the device entity model candidate word; The global optimal sequence of CRF layer tags is inferred through the tags of the MFM model, and the hidden state sequence is converted into the optimal tag sequence, so as to realize entity joint extraction of the candidate words of the device entity model and obtain high-frequency words with confidence.

8. The method for extracting information knowledge based on the power system according to claim 1, characterized in that: The step of performing recognition processing on the confidence high frequency words to obtain the words with the highest confidence includes: Determine the distance between the confidence high frequency words and the device model category; Arrange the distances from small to large to obtain a high-frequency vocabulary sequence table with confidence levels from high to low; The negative words are cleaned out from the high-frequency word sequence table, and the high-frequency words with the highest confidence in the high-frequency word sequence table are output to obtain the words with the highest confidence.

9. A controller, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Power equipment fault knowledge graph construction method

    CN111737496A

  • Electric power Chinese named entity recognition method in combination with word sequence

    CN114564950A

  • Domain new word discovery method and device

    CN114925675A

  • Electric power standard named entity identification method

    CN118607527A

Cited By

  • Forest fire knowledge modeling method based on named entity recognition and relation extraction

    CN121524352A