Power grid defect detection method, electronic device, and program product based on entity recognition
By extending training data generation and model training methods for grid entity recognition models, the problem that grid entity recognition relies on a large amount of labeled data is solved, and the accuracy of entity recognition and grid equipment defect detection is improved.
Patent Information
- Application Number
- CN202510196303.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-21
AI Technical Summary
In power grid entity recognition, the prior art relies on a large number of labeled data to train models, resulting in low accuracy of entity recognition in the case of lack of training data, which in turn affects the accuracy of knowledge graph and grid equipment defect detection.
By extending the sample data of the entity annotation standard, the required amount of training data is generated and the entity recognition model is trained based on this. The specific steps include extracting text data from the power equipment term specification file, establishing a term library with a hierarchical structure, generating fault paths and prompt words, generating synthetic text using a large language model for entity annotation, and finally training the entity recognition model through training data.
The training efficiency and accuracy of the entity recognition model are improved, the accuracy of the knowledge graph is ensured, and thus the accuracy of the power grid equipment defect detection is improved.
Smart Images

Figure CN119691154B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of power grid equipment defect detection, and in particular to a power grid defect detection method, electronic equipment, and program product based on entity recognition. Background Art
[0002] In the scenario of power grid equipment defect detection, the entities and relationships in the power grid can be organized in the form of a graph through the knowledge graph to realize the detection of power grid equipment defects. In the related technology, in the process of constructing the knowledge graph, entity recognition is mainly performed through named entity recognition technology, and then the knowledge graph is created based on the entity.
[0003] Named entity recognition technology is mainly divided into rule-based methods, machine learning-based methods, and deep learning-based methods. Rule-based methods rely on pre-defined patterns and dictionaries. Although they are simple to implement, they lack flexibility and generalization capabilities. Machine learning-based methods train models through feature engineering and can handle complex entity recognition tasks, but they require a large amount of labeled data. For deep learning-based methods, it is particularly difficult to obtain a large amount of high-quality training data in low-resource fields such as power grids due to the large number of professional terms and high data annotation costs. This challenge limits the application of deep learning models in named entity recognition tasks.
[0004] In the related technology, power grid entity recognition needs to rely on a large amount of labeled data training. In the absence of training data, there is a low accuracy rate of entity recognition model, which leads to low accuracy of knowledge graph, and then leads to low accuracy of power grid equipment defect detection. No effective solution has been proposed so far. Summary of the invention
[0005] The embodiments of the present invention provide a power grid defect detection method, electronic device, and program product based on entity recognition, which at least solve the problem that power grid entity recognition in the related technology needs to rely on a large amount of labeled data training. In the absence of training data, there is a low accuracy rate of entity recognition by the entity recognition model, resulting in low accuracy of the knowledge graph, and further resulting in low accuracy of power grid equipment defect detection.
[0006] According to one aspect of an embodiment of the present invention, a method for detecting power grid defects based on entity recognition is provided, comprising: acquiring power grid data, wherein the power grid data includes operating data of various equipment in the power grid; inputting the power grid data into a trained entity recognition model, and outputting corresponding entities from the entity recognition model, wherein the entity recognition model is obtained by training the entity recognition model by obtaining training data after expanding sample data with a large language model; creating a corresponding knowledge graph of power grid equipment defects according to the entities; and detecting the power grid data according to the knowledge graph to determine whether there are power grid equipment defects.
[0007] As an optional solution, before the power grid data is input into a trained entity recognition model and the entity recognition model outputs the corresponding entity, the method also includes: extracting text data from a power equipment terminology specification file to establish a hierarchical term library, wherein each level represents a defect of a different defect logic level; generating a fault path by randomly selecting entities from the term library layer by layer; generating prompt words according to the fault path, wherein the entities in the prompt words are marked; inputting the prompt words into a large language model, and generating a preset number of synthetic texts by the large language model; performing entity annotation on the synthetic text to generate training data for training the entity recognition model; and training the entity recognition model using the training data until the model converges.
[0008] As an optional solution, text data is extracted from the power equipment terminology specification file to establish a term library with a hierarchical structure, including: extracting text data from the power equipment terminology specification file and organizing it into a table form; converting the text data in the table form into a multi-level nested dictionary with a hierarchical structure by formatting and structural processing, wherein each level represents equipment type, equipment category, component, component category, location and specific defects; and using the nested dictionary as the term library.
[0009] As an optional solution, the synthetic text is entity labeled to generate training data for training the entity recognition model, including: preprocessing the synthetic text to identify and label entities through regular expressions; comparing the labeled entities with the entities in the corresponding fault paths to see whether the labeling method is accurate, and whether the labeled entities are consistent with the corresponding entities in the fault paths; when all entity labels in the synthetic text are accurate and consistent with the entities in the corresponding fault paths, labeling different texts of the entities according to the text positions of the entities in the synthetic text, wherein the labeling is used to characterize the position of the text of the entity in the entity; and converting the labeled synthetic text into a format to obtain training data in a target format.
[0010] As an optional solution, the entity recognition model is trained by the training data until the model converges, including: preprocessing the training data, converting the text into multiple tokens, and adding corresponding start tags and end tags at the first position of the token, and generating a corresponding attention mask for each token to obtain an input sequence; processing the input sequence according to the target pre-trained language model to obtain the word vector of the corresponding vocabulary in the input sequence, wherein the target pre-trained language model is a bidirectional encoder representation model based on the transformer architecture; determining the context information of the word vector in the input sequence according to the word vector through a bidirectional long short-term memory network; establishing a label transfer matrix according to the context information through a conditional random field, and entity labeling the vocabulary in the input sequence to obtain an output label sequence, wherein the label sequence is used to characterize the entities in the input synthetic text; comparing the label sequence with the real entity label of the synthetic text, calculating the training index parameter, and determining that the entity recognition model training is completed when the training index parameter reaches the corresponding threshold condition; wherein the entity recognition model includes the target pre-trained language model, the bidirectional long short-term memory network, and the conditional random field.
[0011] As an optional solution, the input sequence is processed according to the target pre-trained language model to obtain the word vector of the corresponding phrase in the input sequence, including: calculating the attention score of the vocabulary in the input sequence to other vocabulary, wherein the calculation formula of the attention score Attention (Q, K, V) is as follows:
[0012]
[0013] Where Q, K, and V represent query, key, and value, respectively, and d k The dimension of the word vector representing the vocabulary, T represents the transposed matrix; according to the attention score, the word vector of the corresponding vocabulary is determined, wherein the word vector is used to represent the information of the corresponding vocabulary itself and related vocabulary information.
[0014] As an optional solution, the context information of the word vector in the input sequence is determined based on the word vector through a bidirectional long short-term memory network, including: processing the word vector of the input sequence according to the forward long short-term memory network to obtain the forward state information of the word vector; processing the word vector of the input sequence according to the backward long short-term memory network to obtain the backward state information of the word vector; merging the forward state information and the backward state information to obtain the context information of the word vector in the input sequence.
[0015] As an optional solution, a label transfer matrix is established according to the context information by using a conditional random field (CRF), and entity annotation is performed on the words in the input sequence to obtain an output label sequence, including: the conditional random field (CRF) establishes a label transfer matrix according to the context information; and calculates the transition probability between different entity labels in the label transfer matrix, wherein the calculation formula of the transition probability is as follows:
[0016]
[0017] Where P(y|X) is the transition probability of entity label y in input sequence X, T yi-1,yi For entity label y from y i-1 State to y i The transition fraction of the state, S yi,i is the position i and state y i ; performing entity labeling on the words in the input sequence according to the transition probability to obtain an output label sequence.
[0018] As an optional scheme, the training index parameters include at least one of the following: precision, recall rate, F1 value; before preprocessing the training data, the method also includes: determining the training set and verification set of the training cycle according to the training data through a cross-validation method, wherein the training set is used to train the entity recognition model, and the verification set is used to calculate the training index parameters; after multiple training cycles are completed, the entity recognition model of the corresponding training cycle is selected according to the training index parameters as the entity recognition model that has finally been trained.
[0019] According to another aspect of the embodiment of the present invention, there is also provided an electronic device, comprising: a processor, and a memory for storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the above-mentioned power grid defect detection method based on entity recognition.
[0020] According to another aspect of the present invention, a computer program product is provided, including a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the above-mentioned power grid defect detection method based on entity recognition is implemented.
[0021] The power grid defect detection method based on entity recognition provided by the embodiment of the present invention expands the sample data of the entity annotation standard to quickly obtain the required amount of training data, and trains the entity recognition model based on this. The entity model is effectively trained based on a small amount of accurate standard sample data, thereby ensuring the efficiency and accuracy of entity recognition model training, and then ensuring the accuracy of creating a knowledge graph based on the entity, thereby improving the accuracy of power grid equipment defect detection based on the knowledge graph, and solving the problem that power grid entity recognition in related technologies needs to rely on a large amount of labeled data training. In the absence of training data, the accuracy of the entity recognition model in identifying entities is low, resulting in low accuracy of the knowledge graph, and then low accuracy of power grid equipment defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.
[0023] Figure 1 It is a flow chart of a power grid defect detection method based on entity recognition according to an embodiment of the present invention.
[0024] Figure 2 It is a schematic diagram of the training process of the grid defective equipment entity recognition model in a low-resource scenario of an embodiment created by the present invention.
[0025] Figure 3 It is a structural schematic diagram of the electronic device created by the present invention. DETAILED DESCRIPTION
[0026] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0027] Grid defect handling is a key link in power system maintenance, which involves a series of complex processes such as real-time monitoring of grid equipment, fault diagnosis, defect identification and maintenance decision-making. With the development of smart grid technology, the amount of data generated by grid equipment has increased dramatically. How to effectively process and analyze this data and improve the efficiency and accuracy of defect handling has become an urgent problem to be solved.
[0028] As a structured semantic knowledge base, the knowledge graph can formally represent entities such as power grid equipment, defects, maintenance strategies and their relationships, providing strong knowledge support for power grid defect handling. Through the knowledge graph, the entities and relationships in the power grid can be organized in the form of a graph to achieve comprehensive perception and intelligent analysis of the power grid status. The combination of the power grid and the knowledge graph can not only improve the accuracy of defect identification, but also optimize the maintenance decision-making process and enhance the level of intelligence of power grid management.
[0029] In the process of building a knowledge graph, named entity recognition technology plays a vital role. Named entity recognition technology aims to identify specific entities from unstructured text, such as names of people, places, organizations, and specific entities in professional fields. These entities are the basic elements for building a knowledge graph.
[0030] In the field of power grids, named entity recognition technology can extract key power grid equipment defect information from the text of power grid anomaly analysis reports, such as equipment type, components, and locations, providing necessary data support for the construction of knowledge graphs. At present, named entity recognition technology is mainly divided into rule-based methods, machine learning-based methods, and deep learning-based methods.
[0031] Rule-based methods rely on predefined patterns and dictionaries. Although they are simple to implement, they lack flexibility and generalization. Machine learning-based methods train models through feature engineering and can handle complex entity recognition tasks, but they require a large amount of labeled data. Deep learning-based methods, such as BERT (Bidirectional Encoder Representations from Transformers), LSTM (Long - Short Term Memory), BiLSTM (Bidirectional LSTM), and CRF (Conditional Random Field) models, can better capture contextual information in text and improve the accuracy of entity recognition, but they also face the challenge of large data requirements.
[0032] In low-resource fields such as power grids, it is particularly difficult to obtain large amounts of high-quality training data due to the large number of professional terms and high data annotation costs. This challenge limits the application of deep learning models in named entity recognition tasks.
[0033] In order to solve the problem of insufficient training data, the methods based on heuristic rules and generative models are currently mainly used. The heuristic rule-based method modifies or replaces the original samples according to preset rules to generate new training samples. This type of method is simple and easy to implement, but the diversity of the synthesized data is poor and it is too dependent on the original samples, which limits the improvement of model performance.
[0034] Generative model-based methods use the generative capabilities of pre-trained language models to synthesize new training data. Such methods can train language models on large-scale corpora and generate new samples that are semantically similar to the original samples. However, these methods still rely on the original sample data and have limitations in generating data with complex structures and rich semantics. In addition, the synthesized data may contain erroneous or inconsistent information, which requires further cleaning and verification to ensure data quality.
[0035] In order to solve the above technical problems, the embodiment of the present invention provides a power grid defect detection method based on entity recognition. Figure 1 As shown, Figure 1 The flowchart of a method for detecting a power grid defect based on entity recognition according to an embodiment of the present invention is as follows. The method for detecting a power grid defect based on entity recognition according to an embodiment of the present invention comprises the following steps:
[0036] Step S101, acquiring power grid data, wherein the power grid data includes operation data of various devices in the power grid;
[0037] Step S102, inputting the power grid data into a trained entity recognition model, and the entity recognition model outputs the corresponding entity, wherein the entity recognition model is obtained by training the entity recognition model by obtaining training data after expanding the sample data through a large language model;
[0038] Step S103, creating a corresponding power grid equipment defect knowledge graph according to the entity;
[0039] Step S104, detect the power grid data according to the knowledge graph to determine whether there are any defects in the power grid equipment.
[0040] The above-mentioned power grid defect detection method based on entity recognition provided by the embodiment created by the present invention expands the sample data of the entity labeling standard, quickly obtains the required amount of training data, and trains the entity recognition model based on this, and effectively trains the entity model based on a small amount of accurate standard sample data, thereby ensuring the efficiency and accuracy of entity recognition model training, and then ensuring the accuracy of creating a knowledge graph based on the entity, thereby achieving the technical effect of improving the accuracy of power grid equipment defect detection based on the knowledge graph.
[0041] The execution subject of the above steps can be the controller or server of the power grid monitoring system. It is connected with the power equipment and detection equipment installed in various places in the power grid to obtain power grid data. Since the power grid system is very large and the power equipment involved is very complex, the above operation data can be understood as mainly data related to equipment and power operation. For example, current, voltage, power, temperature, etc.
[0042] The power grid data is then input into the trained entity recognition model, which outputs the corresponding entity. Entities can be understood as units for creating knowledge graphs, such as equipment type, equipment category, component, component category, location, and specific defect name.
[0043] The above-mentioned execution entities can be connected with the power grid data management system, which may include power grid monitoring system, equipment maintenance management system, fault repair system, etc. Various types of power grid data are collected from these systems to ensure the integrity and accuracy of the data.
[0044] Focus on equipment defect data, which may exist in various forms, such as equipment failure reports (including failure time, name of failed equipment, failure phenomenon, etc.), equipment inspection records (recording abnormal equipment operating status), etc.
[0045] After the grid data is acquired, data cleaning and preprocessing can be performed. The acquired grid data is cleaned to remove noise data, such as error records, duplicate data, etc. For example, if there are obvious input errors in the equipment failure report, such as spelling errors in the equipment name or illogical failure time (such as the failure time is at a time point when the equipment has not been installed), it needs to be corrected or deleted.
[0046] Standardize the data to make the data formats of different sources uniform. For example, unify the description of device names in different systems to avoid problems in subsequent entity recognition due to inconsistent names. At the same time, classify and organize the data to facilitate subsequent input operations.
[0047] The above entity recognition model is obtained by expanding the sample data with a large language model to obtain training data, and then training the entity recognition model with the training data. This is because the number of relevant operation and maintenance analysis reports in the named entity recognition task in the field of power grid equipment operation and maintenance is small, and the number of professional terms makes the data annotation cost high, and there is a serious shortage of training data.
[0048] Therefore, this embodiment expands the sample data to obtain the required amount of training data, and then performs model training using the training data to obtain an entity recognition model with high accuracy.
[0049] The entity recognition model is also improved to enable it to have a better recognition effect on entities with defects in power equipment. The specific details will be explained later.
[0050] According to the entities, the corresponding knowledge graph of power grid equipment defects is created. There are mainly the following ways to build the knowledge graph of power defects according to the power equipment defect entities: top-down construction method; bottom-up construction method; joint extraction construction method; ontology-based construction method; combined with graph database construction method, etc. The bottom-up construction method and the joint extraction-based construction method are more commonly used.
[0051] Taking the joint extraction construction method as an example, when establishing the knowledge graph of power equipment defects, a joint annotation system is first constructed to integrate the entity annotation and entity relationship annotation of the identified power equipment defects, and design a unified annotation set containing entity labels and relationship labels. For example, the BIO annotation method is used to annotate entities, and the relationship between entities is represented by specific labels. Specifically, in entity annotation, "B" represents the beginning of the entity, "I" represents the middle part of the entity, and "O" represents the non-entity part; in terms of entity relationship annotation, the entity type can be represented by voltage level, equipment, location and phenomenon, and the relationship category can be represented by belonging relationship and defect description.
[0052] Then, for data annotation and review, professionals are organized to annotate the power equipment defect data according to the joint annotation system. Annotators need to have an in-depth understanding of the field of power equipment defects and be familiar with various equipment, defect types and related terms. After the annotation is completed, the annotation results are reviewed and corrected. A multi-person cross-examination method can be used to ensure the accuracy and consistency of the annotation. For controversial or uncertain annotation results, experts are organized to discuss and make decisions.
[0053] Model training and extraction, use labeled data to train joint extraction models, such as the BERT+BiLSTM+CRF model based on deep learning. First, input the labeled data into the model for training. During the training process, the model will automatically learn the semantic information and annotation patterns in the text and adjust the parameters of the model to achieve the best extraction effect. Then, the trained model automatically extracts the power equipment defect entities and relations in the text and converts them into triples in the knowledge graph, i.e. <subject, predicate, object>, where the subject and object are usually entities and the predicate is the relationship between the two.
[0054] Knowledge graph construction and optimization: Import the extracted triple data into the graph database or other knowledge graph construction tools to construct the knowledge graph of power equipment defects. During the construction process, the knowledge graph may need to be further optimized and improved, such as removing duplicate triples, checking and repairing relationship errors between entities, etc. At the same time, the knowledge graph can also be manually reviewed and supplemented based on the knowledge and experience of domain experts to ensure the accuracy and completeness of the knowledge graph.
[0055] In addition, the knowledge graph can be updated and maintained. With the continuous update and accumulation of power equipment defect data, the knowledge graph needs to be updated and maintained regularly to ensure its timeliness and accuracy. Regular update tasks can be set, such as extracting and integrating newly generated equipment defect data into the knowledge graph every month or quarter. At the same time, when there are major changes in the type, operating environment, defect characteristics, etc. of power equipment, the structure and content of the knowledge graph should be adjusted and updated in a timely manner.
[0056] There are many ways to detect power grid data based on the knowledge graph to determine whether there are defects in the power grid equipment, including rule-based reasoning detection, detection based on association analysis, or detection methods based on machine learning.
[0057] The detection based on association analysis is taken as an example for explanation.
[0058] First, perform entity and relationship analysis: In the knowledge graph, analyze the associations between power grid equipment entities and the correlations between equipment attributes and defects. For example, frequent tripping of certain types of switchgear (through event association) may be associated with voltage fluctuations in the power grid (another entity's attribute), which may indicate potential problems with voltage stability in the power grid or faults in the equipment itself.
[0059] Then, data association mining is performed: data mining techniques, such as frequent item set mining and association rule mining, are used to process power grid data to find frequently occurring association patterns between equipment entities and between equipment attributes and potential defects. For example, a high association pattern between "a specific type of insulator" and "a flashover event under specific weather conditions" is mined. When new data conforms to this pattern, it is possible to focus on whether the insulator has defects.
[0060] Then perform similarity calculation and detection: Calculate the similarity between the new power grid data and the known defect-related data in the knowledge graph. You can use a graph similarity algorithm, such as a path-based similarity calculation method or a subgraph matching algorithm. For example, compare the new equipment failure data with the typical failure case subgraph in the knowledge graph. If the similarity is high, there may be similar equipment defects.
[0061] Association analysis can discover complex relationships hidden in data and is suitable for mining potential defect patterns that have not yet been clearly defined by rules. However, it requires a large amount of data for effective association mining and the computational complexity may be high, especially when dealing with large-scale knowledge graphs and massive power grid data.
[0062] As an optional solution, before the power grid data is input into the trained entity recognition model and the entity recognition model outputs the corresponding entity, the method also includes: extracting text data from the power equipment terminology specification file to establish a hierarchical terminology library, wherein each level represents defects of different defect logic levels; generating a fault path by randomly selecting entities from the terminology library layer by layer; generating prompt words according to the fault path, wherein the entities in the prompt words are marked; inputting the prompt words into a large language model, and generating a preset number of synthetic texts from the large language model; annotating the synthetic texts to generate training data for training the entity recognition model; and training the entity recognition model through the training data until the model converges.
[0063] When training the entity recognition model, we first extract text data from the power equipment terminology specification file to build a hierarchical terminology library, where each level represents defects of different defect logic levels. This structured multi-level data structure allows the system to accurately follow the actual composition of power grid equipment and the logical relationship of potential defects when sampling paths, thus facilitating subsequent training.
[0064] By randomly selecting entities from the term base layer by layer, a fault path is generated; each entity path contains complete equipment fault information, from equipment type to specific defect description, forming a complete fault scenario. This fault path not only reflects the possible equipment fault scenario, but also provides specific and accurate input conditions for subsequent text generation.
[0065] Generate prompt words according to the fault path, wherein the entities in the prompt words are marked. Prompt words of a large language model can be generated according to the entities of the fault path by using a small amount of prompt technology, a thought chain prompt technology, etc.
[0066] The prompt words are input into the large language model, and the large language model generates a preset number of synthetic texts; the large language model can expand the sample data according to the prompt words to generate corresponding synthetic texts. The main content of the synthetic text may include the language expression and technical details of the fault report when the power equipment fails.
[0067] The synthetic text is annotated with entities to generate training data for the entity recognition model, and the entity recognition model is trained with the training data until the model converges. In this way, the entity recognition model can be trained with less sample data, so that the entity recognition model has a higher entity recognition accuracy.
[0068] As an optional solution, text data is extracted from the power equipment terminology specification file to establish a hierarchical terminology library, including: extracting text data from the power equipment terminology specification file and organizing it into a table form; converting the tabular text data into a hierarchical multi-level nested dictionary by formatting and structuralizing the text data, wherein each level represents equipment type, equipment category, component, component category, location and specific defects; and using the nested dictionary as a terminology library.
[0069] The above-mentioned terminology library is a multi-level data structure, so that when the fault path sampling is performed subsequently, it can accurately follow the actual composition of the power grid equipment and the logical relationship between the potential defects, improve the authenticity and rationality of the fault scenario corresponding to the fault path, thereby facilitating subsequent training and improving the accuracy and authenticity of the synthetic text generated by the large language model.
[0070] As an optional solution, synthetic text is annotated with entities to generate training data for training an entity recognition model, including: preprocessing the synthetic text to identify and mark entities through regular expressions; comparing the marking method of the marked entities with the entities in the corresponding fault paths to see whether it is accurate, and whether the marked entities are consistent with the corresponding entities in the fault paths; when all entity markings in the synthetic text are accurate and consistent with the entities in the corresponding fault paths, annotating different entity texts according to the text positions of the entities in the synthetic text, wherein the annotation is used to represent the position of the text in the entity; and converting the annotated synthetic text into a format to obtain training data in a target format.
[0071] The above preprocesses the synthetic text and identifies and marks entities through regular expressions. Regular expression matching can be used to identify and extract entity tags, and custom entity tags can be used to locate entities in the text, such as "&& circuit breaker##".
[0072] The system compares each generated synthetic text with its corresponding entity path one by one, checking whether the entities in the text are correctly embedded and consistent with the input entity path, ensuring that the generated text content is semantically consistent with the input data and preventing entity tagging errors or omissions. The system also checks whether the marked entities are accurately marked and consistent with the corresponding entities in the fault path.
[0073] According to the text positions of entities in the synthetic text, different texts of entities can be annotated, and the BIO annotation method can be adopted. Specifically, in the BIO tagging method, the "B-" tag indicates the start of an entity, the "I-" tag indicates the part within an entity, and "O" indicates the non-entity part. Each character or word is correspondingly annotated according to its type and position belonging to the entity. Adjust the annotation sequence according to the start and end positions of the entity.
[0074] This step involves deleting the entity markers and applying the entity annotations to the actual text content. For example, if an entity is marked as "&&reactor##" in the text, then "reactor" is annotated as "B-DETYPE", "I-DETYPE", "I-DETYPE". The data after annotation is converted into JSON format to prepare a formatted training dataset for model training.
[0075] As an alternative solution, the entity recognition model is trained with training data until the model converges, including: preprocessing the training data, converting the text into multiple tokens, adding corresponding start and end markers at the beginning and end positions of the tokens respectively, and generating corresponding attention masks for each token to obtain an input sequence; processing the input sequence according to the target pre-trained language model to obtain the word vectors of the corresponding vocabulary in the input sequence, where the target pre-trained language model is a bidirectional encoder representation model based on the transformer architecture; determining the context information of the word vectors in the input sequence through a bidirectional long short-term memory network; establishing a label transition matrix through a conditional random field according to the context information to perform entity annotation on the vocabulary in the input sequence to obtain an output label sequence, where the label sequence is used to represent the entities in the input synthetic text; comparing the label sequence with the true entity labels of the synthetic text to calculate the training metric parameters, and determining that the entity recognition model training is completed when the training metric parameters reach the corresponding threshold conditions.
[0076] The entity recognition model of this embodiment is the target pre-trained language model BERT, the bidirectional long short-term memory network BiLSTM, and the conditional random field CRF. It can be a BERT-BiLSTM-CRF model.
[0077] As a pre-trained model, BERT can learn rich semantic knowledge and context information through pre-training on a large-scale corpus, providing a good foundation for subsequent tasks.
[0078] BiLSTM can further capture the context information in the sequence, model the text from both the front and back directions, enhance the understanding of long-distance dependencies, and better understand the semantic and syntactic structures of the text.
[0079] CRF can take into account the dependencies between labels, select the optimal label sequence through global optimization, and avoid the problem of local optimal solution, thereby improving the accuracy and coherence of sequence labeling. It is particularly suitable for tasks such as named entity recognition and part-of-speech tagging, making the labeling results more accurate and reasonable.
[0080] As an optional solution, the input sequence is processed according to the target pre-trained language model to obtain the word vector of the corresponding phrase in the input sequence, including: calculating the attention score of the words in the input sequence to other words, where the calculation formula of the attention score Attention (Q, K, V) is as follows:
[0081]
[0082] Where Q, K, and V represent query, key, and value, respectively, and d k Represents the dimension of the word vector of the vocabulary, and T represents the transposed matrix; according to the attention score, the word vector of the corresponding vocabulary is determined, where the word vector is used to represent the information of the corresponding vocabulary itself and related vocabulary information.
[0083] Through the above calculation formula, the assistant scores of different words in the input sequence to other words can be accurately calculated, indicating the correlation between the word and the current fault path.
[0084] In the BERT model, the attention score is used to calculate the influence of each word on other words, thereby generating a context-dependent representation of each word. Specifically, the attention mechanism determines the importance of each word in different contexts by calculating the relationship between the query (Q), key (K), and value (V).
[0085] Through the attention mechanism, BERT can generate a vector representation for each word that contains the surrounding context information. This vector representation is the word vector, which contains not only the information of the word itself, but also the information of other related words, so as to better capture the semantics of the word.
[0086] As an optional solution, the contextual information of the word vector in the input sequence is determined based on the word vector through a bidirectional long short-term memory network, including: processing the word vector of the input sequence according to the forward long short-term memory network to obtain the forward state information of the word vector; processing the word vector of the input sequence according to the backward long short-term memory network to obtain the backward state information of the word vector; merging the forward state information and the backward state information to obtain the contextual information of the word vector in the input sequence.
[0087] BiLSTM captures contextual information in a sequence through two LSTM networks, forward and backward. The forward LSTM processes information flow from left to right, while the backward LSTM processes information flow from right to left. The hidden states of the forward and backward LSTMs are merged (usually concatenated or added) at each time step to form a comprehensive representation. This merged hidden state can contain both forward and backward information, thereby better representing the context of each word in the sequence.
[0088] As an optional solution, a label transfer matrix is established according to the context information through the conditional random field CRF, and the words in the input sequence are entity labeled to obtain an output label sequence, including: the conditional random field CRF establishes a label transfer matrix according to the context information; and calculates the transition probability between different entity labels in the label transfer matrix, wherein the calculation formula of the transition probability is as follows:
[0089]
[0090] Where P(y|X) is the transition probability of entity label y in input sequence X, T yi-1,yi For entity label y from y i-1 State to y i The transition fraction of the state, S yi,i is the position i and state y i The emission score of ; the words in the input sequence are entity labeled according to the transition probability to obtain the output label sequence.
[0091] The output of the CRF layer is a label sequence, which is a label for each position in the input sequence. This label sequence is calculated based on the feature vector output by the BiLSTM layer and the label transition probability learned by the CRF layer. The above calculation formula can accurately calculate the transition probability of the entity label, thereby improving the accuracy of the label sequence output by the CRF layer.
[0092] As an optional solution, the training index parameters include at least one of the following: precision, recall rate, F1 value; before preprocessing the training data, the method also includes: determining the training set and validation set of the training cycle based on the training data through a cross-validation method, wherein the training set is used to train the entity recognition model and the validation set is used to calculate the training index parameters; after multiple training cycles are completed, the entity recognition model of the corresponding training cycle is selected according to the training index parameters as the entity recognition model that has been finally trained.
[0093] The above cross-validation is a model evaluation method, which is mainly used to evaluate the performance of machine learning models, especially when the amount of data is limited, so that data can be used more effectively for model training and verification. It divides the data set into multiple subsets, and trains and verifies on different subset combinations, so as to obtain a more comprehensive and reliable model performance evaluation.
[0094] Specific verification methods can be K-fold cross validation, leave-one-out cross validation, and stratified cross validation.
[0095] It should be noted that this embodiment also provides an optional implementation, which is described in detail below.
[0096] This embodiment provides a method for detecting defects in power grid equipment based on entity recognition, which mainly focuses on the recognition of power grid defect entities enhanced by synthetic data based on a generative language model in low-resource scenarios, and the detection of power grid equipment defects based on the recognized entities. It aims to solve the problem of a small number of relevant operation and maintenance analysis reports in the named entity recognition task in the field of power grid equipment operation and maintenance, and the high cost of data annotation due to the large number of professional terms, which leads to a serious shortage of training data.
[0097] This problem limits the effective application of deep learning models (such as BERT-BiLSTM-CRF, etc.) that require a large amount of labeled data to train to achieve high accuracy in this field. The purpose of this embodiment is to overcome this limitation in low-resource scenarios by synthesizing grid defective equipment training data based on a large language model, thereby improving the accuracy and robustness of grid defective equipment entity recognition.
[0098] This implementation can effectively improve the training effect and recognition accuracy of the power grid entity extraction model in the absence of sufficient training data. By intelligently generating samples from the existing term library, this implementation uses synthetic data to train the sequence annotation model, significantly improving the performance of the model in real scenarios.
[0099] Figure 2 Schematic diagram of the training process of the power grid defective equipment entity recognition model in a low resource scenario according to an embodiment of the present invention, such as Figure 2 As shown, the entity identification process of power grid defective equipment in the above low-resource scenario includes the following steps:
[0100] Step 1: Analyze the corporate standards of State Grid Corporation of China, extract terms related to power grid equipment and its defects, and build a hierarchical term library;
[0101] Step 2: Use the constructed term library to sample entity paths to simulate the association between real equipment and defects;
[0102] Step 3: Based on the sampled entity paths, construct prompt words for the generative large model;
[0103] Step 4: Use the generative big model to generate synthetic text describing the power grid equipment and its defects based on the constructed prompt words;
[0104] Step 5: Annotate the generated synthetic text and convert the data into a training set using the BIO tagging method;
[0105] Step 6: Train the BERT-BiLSTM-CRF sequence labeling model and calculate the precision, recall, and F1 value of the model on the synthetic training data;
[0106] Step 7: Perform entity recognition on a manually annotated test dataset to evaluate the performance of the model.
[0107] Specifically, in the above step 1, based on the standard document "Specification of Defect Terms for Power Transmission and Transformation Equipment", the complete equipment defect description standard is extracted by analyzing the equipment defect description terms in detail. The table covers the hierarchical structure from equipment type to equipment category, component, component category, location, and the final defect description.
[0108] The data in the document was identified using OCR (Optical Character Recognition) technology and initially organized into an Excel table. The data was further formatted and structured using Python's Pandas library. The data was converted into a hierarchical JSON format and a multi-level nested dictionary was constructed, in which each level represented the equipment type, equipment category, component, component category, location, and specific defects.
[0109] In step 2, first, entities are randomly selected layer by layer from the nested dictionary that has been structured in step 1, starting from the equipment type, and then selecting equipment type, component, component type, location, and finally defect description, thereby generating a complete equipment failure path.
[0110] In step 3, based on the entity paths generated by the hierarchical terminology library in step 2, prompt words are constructed to guide the generative large model. Based on the sampled entity paths, the system constructs prompt words for input into the generative large model. These prompt words are presented in the form of natural language, in which specific symbols (such as "&&entity##") are explicitly used to mark each power grid entity to ensure that these entities are accurately reflected when generating text.
[0111] In step 4, a large generative model is used to generate synthetic text describing power grid equipment and its defects based on carefully constructed prompt words. This step specifically uses OpenAI's GPT-4o model, a pre-trained language model based on the Transformer architecture that can generate highly coherent and contextually relevant text.
[0112] The prompt words input into the model are generated from the physical paths of power grid equipment, which detail the complete information chain from equipment type to specific defects. The model generates conditional text based on these prompt words, and each generation is designed to simulate the language expressions and technical details that may appear in power grid equipment fault reports.
[0113] This step generates a total of 5,000 power grid equipment defect texts, which cover various possible fault scenarios and enhance the diversity of the data set.
[0114] In step 5, the process of entity labeling of the generated synthetic text and converting it into a training set using the BIO notation method involves the following key steps:
[0115] Step S51: perform text preprocessing. The synthesized text is first identified and extracted through regular expression matching. This step uses custom entity tags (such as "&& circuit breaker##") to locate entities in the text.
[0116] Step S52: The generated labeled data is quality controlled and verified before being saved to ensure the accuracy and consistency of the data and avoid errors and inconsistencies in the training data.
[0117] Step S53: Entity extraction, labeling and processing. The extracted entities and their location information in the text are used to apply the BIO tagging method.
[0118] In step 6, the present invention uses synthetic data to train the BERT-BiLSTM-CRF sequence labeling model to identify different entities of power grid equipment. This process involves the following key steps:
[0119] Step S61: The first step is data preprocessing to prepare a format suitable for BERT model input. This includes converting the text into token IDs and generating a corresponding attention mask for each token. Special tags such as [CLS] and [SEP] are also added to distinguish the beginning and end of a sentence.
[0120] Step S62: Next, the pre-trained BERT model is used to extract deep semantic features of the text in the input sequence. The BERT layer calculates the attention scores of all word pairs in the input sequence through the self-attention mechanism. Through this mechanism, the BERT layer generates a vector representation for each word that contains the surrounding context information, providing rich semantic input for subsequent layers.
[0121] Step S63: The BiLSTM layer receives the word vector output by BERT and captures the context information of the sequence data through its bidirectional structure. The goal of this layer is to enhance the model's understanding of semantics through forward and backward LSTM networks. The forward LSTM processes information flow from left to right, while the backward LSTM processes information flow from right to left. The output of each time step is obtained by merging the hidden states of the forward and backward.
[0122] Step S64: The CRF layer is used to optimize the prediction of the label sequence in sequence labeling. It ensures the consistency and accuracy of the predicted sequence by learning the optimal label transition probability. The CRF layer calculates the conditional probability of the entire label sequence, and the optimization goal is to maximize the probability of the correct label sequence.
[0123] Step S65: During the model training process, the model parameters are optimized through the cross-validation method to adapt to the characteristics of the training data. After the model is trained using synthetic data, the precision, recall and F1 value of the model on the validation set are calculated. These indicators evaluate the performance of the model on the entity recognition task.
[0124] Precision refers to the ratio of correctly identified entities to the total number of identified entities, recall refers to the ratio of correctly identified entities to the total number of true entities, and the F1 value is the harmonic mean of precision and recall, which is used to measure the overall performance of the model.
[0125] In step 7, the BERT-BiLSTM-CRF sequence labeling model trained on synthetic data is used to evaluate the performance in real scenarios. The test is conducted on a total of 200 manually labeled power grid anomaly analysis texts.
[0126] Table 1 is a statistical table of training index parameters for different entity categories of test data. As shown in Table 1, the test results show that the model exhibits extremely high accuracy and recall rate in different entity categories. Combining all test data, an F1 value of 0.9810 for the overall data of all categories is obtained. These performance indicators verify the effectiveness and reliability of this implementation in low-resource scenarios of power grid defective equipment entity recognition by means of synthetic training data.
[0127] Table 1 Statistics of training index parameters for different entity categories of test data
[0128] Entity Class Accuracy Recall F1 value defect 0.9927 0.9945 0.9936 Location 0.9177 0.9667 0.9416 Equipment Type 1.0000 1.0000 1.0000 Device Type 0.9848 0.9375 0.9606 part 0.9903 0.9855 0.9879 Parts Type 0.9310 1.0000 0.9643 All Categories 0.9802 0.9817 0.9810
[0129] The synthetic data enhancement method for identifying defective power grid equipment entities in low-resource scenarios provided in this embodiment has the following significant beneficial effects:
[0130] Efficient solution for low-resource scenarios: In response to the common data scarcity problem in the power grid field, this paper significantly expands the data set that can be used for model training by building a hierarchical terminology library from limited power grid standard documents and generating synthetic text. This method is particularly suitable for low-resource environments in the power grid field where data acquisition is difficult, effectively improving the diversity and coverage of data, so that model training is no longer limited by the quantity and quality of actual data.
[0131] Enhanced model accuracy and generalization: By constructing a hierarchical term library and sampling entity paths based on it, the present invention is able to generate high-quality synthetic training data. This not only expands the coverage of the dataset, but also ensures that the dataset can simulate various power grid equipment defect scenarios in the real world, thereby significantly improving the effectiveness and practicality of model training. The model can therefore more accurately identify and classify a wide range of power grid equipment entities and their defects, enhancing the model's generalization ability to unseen data.
[0132] Improve the efficiency and reduce the cost of entity recognition: The generative large model used in this invention automatically generates synthetic text based on structured prompt words. This automated process significantly improves the speed and efficiency of entity annotation. Compared with traditional manual annotation methods, this method can generate a large amount of training data in a short time, accelerate the development and deployment cycle of the model, and significantly reduce the cost of data preparation and model training.
[0133] This implementation method constructs a hierarchical terminology library containing multi-level entities by analyzing power grid standard documents, and uses a random sampling method for the terminology library to provide structured input for generating synthetic text. The generative large model is used to automatically generate text describing power grid equipment and its defects according to the hierarchical entity path, which effectively expands the amount of data for model training and solves the problem of actual data scarcity. The generated text is converted into the BIO annotation format through an automated annotation system, which improves the efficiency and quality of training data preparation.
[0134] Based on the above-mentioned entity recognition-based power grid defect detection method provided by the embodiment of the invention, the embodiment of the invention also provides an entity recognition-based power grid defect detection device, which is applied to power grid equipment defects. The device includes: a data acquisition module, an entity recognition module, and a graph creation module. The device is described in detail below.
[0135] A data acquisition module is used to acquire power grid data, wherein the power grid data includes operation data of various devices in the power grid;
[0136] An entity recognition module is connected to the data acquisition module and is used to input the power grid data into a trained entity recognition model, and the entity recognition model outputs the corresponding entity, wherein the entity recognition model is obtained by training the entity recognition model by expanding the sample data with a large language model to obtain training data;
[0137] A graph creation module, connected to the above-mentioned entity recognition module, is used to create a corresponding power grid equipment defect knowledge graph according to the entity;
[0138] The defect detection module is connected to the above-mentioned graph creation module and is used to detect the power grid data based on the knowledge graph to determine whether there are any defects in the power grid equipment.
[0139] The above-mentioned power grid defect detection device based on entity recognition provided by the embodiment created by the present invention expands the sample data of the entity annotation standard, quickly obtains the required amount of training data, and trains the entity recognition model based on this, and effectively trains the entity model based on a small amount of accurate standard sample data, thereby ensuring the efficiency and accuracy of entity recognition model training, and further ensuring the accuracy of creating a knowledge graph based on the entity, thereby achieving the technical effect of improving the accuracy of power grid equipment defect detection based on the knowledge graph.
[0140] The embodiments of the present invention also provide a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.
[0141] The embodiments of the present invention further provide a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.
[0142] The embodiment of the invention also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform the method of the embodiment of the invention when executed by the at least one processor.
[0143] refer to Figure 3, a block diagram of an electronic device that can be used as a server or client of an embodiment of the invention will now be described, which is an example of hardware devices that can be applied to various aspects of the invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or required herein.
[0144] like Figure 3 As shown, the electronic device includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory ROM 302 or a computer program loaded from a storage unit 308 into a random access memory RAM 303. In RAM 303, various programs and data required for the operation of the electronic device can also be stored. The computing unit 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An input / output I / O interface 305 is also connected to the bus 304.
[0145] Multiple components in the electronic device are connected to the I / O interface 305, including: an input unit 306, an output unit 307, a storage unit 308, and a communication unit 309. The input unit 306 can be any type of device that can input information to the electronic device, and the input unit 306 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device. The output unit 307 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 308 can include but is not limited to a disk, an optical disk. The communication unit 309 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0146] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a CPU, a graphics processing unit GPU, various dedicated artificial intelligence AI computing units, various computing units running machine learning model algorithms, a digital signal processor DSP, and any appropriate processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above. For example, in some embodiments, the method embodiments created by the present invention may be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via a ROM 302 and / or a communication unit 309. In some embodiments, the computing unit 301 may be configured to perform the above method in any other appropriate manner (e.g., by means of firmware).
[0147] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the computer program, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of the invention, the machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memory RAM, read-only memory ROM, erasable programmable read-only memory EPROM or flash memory, optical fibers, portable compact disk read-only memory CD-ROM, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0149] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0150] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0151] The various steps described in the method implementation methods provided by the embodiments of the present invention can be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.
[0152] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same and similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.
[0153] The above-described embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of protection. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the attached claims.
Claims
1. A power grid defect detection method based on entity recognition, characterized in that: include: Extract text data from the power equipment terminology specification file and establish a hierarchical terminology library, where each level represents defects of different defect logic levels; Generate a fault path by randomly selecting entities layer by layer from the term base; generating a prompt word according to the fault path, wherein entities in the prompt word are marked; Inputting the prompt word into a large language model, and generating a preset number of synthetic texts by the large language model; Performing entity annotation on the synthetic text to generate training data for training the entity recognition model; Training the entity recognition model using the training data until the model converges; Acquiring power grid data, wherein the power grid data includes operation data of various devices in the power grid; Input the power grid data into a trained entity recognition model, and the entity recognition model outputs the corresponding entity, wherein the entity recognition model is obtained by training the entity recognition model by expanding the sample data with a large language model to obtain training data; Creating a corresponding power grid equipment defect knowledge graph according to the entity; The power grid data is detected according to the knowledge graph to determine whether there are any power grid equipment defects.
2. The method according to claim 1, characterized in that Extract text data from the power equipment terminology specification file and establish a hierarchical terminology database, including: Extracting text data from the power equipment terminology specification document and arranging it into a table form; By formatting and structuralizing the text data in the table form, the text data is converted into a multi-level nested dictionary with a hierarchical structure, wherein each level represents the equipment type, equipment category, component, component category, location and specific defect; The nested dictionary is used as the term base.
3. The method according to claim 1, characterized in that Performing entity labeling on the synthetic text to generate training data for training the entity recognition model includes: Preprocessing the synthetic text to identify and mark entities using regular expressions; Check whether the marked entities are accurately marked and whether the marked entities are consistent with the corresponding entities in the fault path; When all entity labels in the synthetic text are accurate and consistent with the entities in the corresponding fault path, the different texts of the entity are annotated according to the text position of the entity in the synthetic text, wherein the annotation is used to represent the position of the text of the entity in the entity; The format of the annotated synthetic text is converted to obtain training data in the target format.
4. The method according to claim 1, characterized in that: The entity recognition model is trained using the training data until the model converges, including: Preprocessing the training data, converting the text into multiple tokens, adding corresponding start tags and end tags at the first position of the tokens, and generating a corresponding attention mask for each token to obtain an input sequence; Processing the input sequence according to a target pre-trained language model to obtain word vectors of corresponding words in the input sequence, wherein the target pre-trained language model is a bidirectional encoder representation model based on a transformer architecture; Determine context information of the word vector in the input sequence based on the word vector through a bidirectional long short-term memory network; Establishing a label transfer matrix according to the context information through a conditional random field, performing entity labeling on the words in the input sequence, and obtaining an output label sequence, wherein the label sequence is used to represent entities in the input synthetic text; By comparing the label sequence with the real entity label of the synthetic text, a training index parameter is calculated, and when the training index parameter reaches a corresponding threshold condition, it is determined that the entity recognition model training is completed; The entity recognition model includes the target pre-trained language model, the bidirectional long short-term memory network, and the conditional random field.
5. The method according to claim 4, characterized in that The input sequence is processed according to the target pre-trained language model to obtain word vectors corresponding to the phrases in the input sequence, including: Calculate the attention score of the words in the input sequence to other words, where the calculation formula of the attention score Attention (Q, K, V) is as follows: ; Where Q, K, and V represent query, key, and value, respectively. d k Represents the dimension of the word vector of the vocabulary, and T represents the transposed matrix; According to the attention score, a word vector of the corresponding vocabulary is determined, wherein the word vector is used to represent information of the corresponding vocabulary itself and related vocabulary information.
6. The method according to claim 4, characterized in that Determining context information of the word vector in the input sequence through a bidirectional long short-term memory network according to the word vector, including: Processing the word vectors of the input sequence according to a forward long short-term memory network to obtain forward state information of the word vectors; Processing the word vectors of the input sequence according to a backward long short-term memory network to obtain backward state information of the word vectors; The forward state information and the backward state information are merged to obtain the context information of the word vector in the input sequence.
7. The method according to claim 4, characterized in that A label transfer matrix is established according to the context information through the conditional random field CRF, and the words in the input sequence are entity labeled to obtain an output label sequence, including: The conditional random field CRF establishes a label transfer matrix according to the context information; The transition probabilities between different entity labels in the label transfer matrix are calculated, wherein the calculation formula of the transition probability is as follows: ; Where P(y|X) is the transition probability of entity label y in input sequence X, T yi-1,yi For entity label y from y i-1 State to y i The transition fraction of the state, S yi,i is the position i and state y i The emission fraction; The words in the input sequence are entity labeled according to the transition probability to obtain an output label sequence.
8. The method according to claim 4, characterized in that The training index parameters include at least one of the following: precision, recall rate, F1 value; Before preprocessing the training data, the method further includes: Determine a training set and a validation set of a training cycle according to the training data by a cross-validation method, wherein the training set is used to train the entity recognition model, and the validation set is used to calculate the training indicator parameters; After multiple training cycles are completed, the entity recognition model of the corresponding training cycle is selected according to the training indicator parameters as the entity recognition model that has been finally trained.
9. An electronic device, comprising: A processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Power equipment operation and maintenance method and device based on knowledge graph
CN118378194A
Named entity recognition method and equipment based on large language model
CN119129596A