Power equipment fault knowledge graph construction method based on cooperation of large and small models
Patent Information
- Application Number
- CN202510777832.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies face the problems of high cost and low efficiency when constructing knowledge graphs for power equipment faults, especially in obtaining triple information in the power field and the "hallucination" problem of large language models in the power field, which makes it impossible to effectively distinguish the subtle differences in terms in the power field.
A method based on the collaboration of large and small models is adopted. By combining a large language model with a traditional deep learning small model, data preprocessing, named entity recognition, relationship extraction and graph storage are used to construct a knowledge graph of power equipment faults. This includes steps such as data preprocessing, building prompt word templates, hierarchical verification and entity alignment, which reduces manual labeling costs and improves recognition accuracy.
It achieves efficient and accurate construction of the knowledge graph of power equipment failure, reduces the cost of manual labeling, improves the accuracy of named entity recognition and relationship extraction, and ensures the integrity and accuracy of the knowledge graph of power equipment failure.
Smart Images

Figure CN120670602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method and system for constructing a knowledge graph of power equipment faults based on the collaboration of large and small models. Background Art
[0002] The rapid development of information technology is reshaping the operational models of various industries, driving their transformation and upgrading towards intelligent and digital capabilities. As a key branch of artificial intelligence, knowledge graphs construct complex semantic networks through the organic combination of entities, relationships, and attributes. They structure knowledge in the form of "subject-predicate-object" triples, providing powerful tools for efficient knowledge management and application.
[0003] In the field of power equipment maintenance, under the traditional management model, equipment maintenance mainly relies on expert experience and manual records, which is inefficient and prone to production interruptions. Building a high-quality knowledge graph can systematically integrate historical maintenance data, intelligently identify defects, predict potential risks, and automatically generate maintenance plans to achieve more efficient equipment management. At present, the technical methods for constructing power equipment fault knowledge graphs are mainly divided into rule-based methods, statistical machine learning-based methods, and deep learning small model-based methods. Rule-based methods rely on clear logical rules and professional knowledge. The construction process is very complex and difficult to cope with the ever-expanding data. Statistical machine learning small models and deep learning small models require a large amount of labeled training data. The labeling process is time-consuming and expensive, significantly increasing the cost and time of knowledge graph construction.
[0004] In recent years, the outstanding performance of large language models in natural language processing tasks has attracted widespread attention. The latest research has gradually focused on combining large language models with knowledge graphs, using their powerful language understanding and contextual reasoning capabilities to extract knowledge, in order to address the challenges faced in the current knowledge graph construction, such as the difficulty in acquiring domain knowledge and the high cost of data annotation. However, large language models still have the "hallucination" problem in the power field, which may produce meaningless output and cannot effectively distinguish the subtle differences in terms in the power field. In summary, the construction of high-quality power equipment fault knowledge graphs mainly faces the following challenges: (1) How to efficiently and accurately obtain triple information in the power field at a low labor cost; (2) How to give full play to the complementary advantages of large language models and traditional methods in the whole process of constructing knowledge graphs to achieve the efficient construction of high-quality power equipment fault knowledge graphs. Summary of the Invention
[0005] 1. Technical problems to be solved:
[0006] In response to the above technical problems, the present invention provides a method for constructing a knowledge graph of power equipment faults based on collaboration between large and small models.
[0007] 2. Technical solution:
[0008] The method for constructing a knowledge graph of power equipment faults based on collaboration between large and small models is characterized by:
[0009] Step 1: Data preprocessing: This involves obtaining text and image information related to power equipment troubleshooting and preprocessing the information. The preprocessed image information is stored in a cloud database for easy access, and the preprocessed text information is semantically segmented using a large language model to form independent text blocks. The text information that meets preset quality requirements is then filtered through topic classification.
[0010] Step 2: Use the large language model to analyze and integrate the power equipment fault maintenance text information in the pre-processed data to obtain core entity types, and combine the domain expert knowledge to construct the domain ontology of the knowledge graph model layer;
[0011] Step 3: Construct a prompt word template and use the large language model to pre-label 10% of the text information in step 1. Use rule matching and power experts to correct the pre-labeling results to obtain a training data set for the subsequent traditional deep learning small model for named entity recognition; Based on the characteristics of the fault field data, optimize the traditional deep learning small model for named entity recognition and use the training data set to train the model. After training, recognize the remaining 90% of the text information in step 1; Design a hierarchical verification strategy to stratify the large language model based on cost. Select a low-cost large language model and apply the pre-labeled prompt word template to recognize the remaining 90% of the text information in step 1. Compare its recognition results with those of the traditional deep learning small model, and eliminate erroneous entities to obtain an entity set;
[0012] Step 4: Based on the characteristics of power equipment fault data and the domain ontology, different entity relationship type combinations are designed, and corresponding prompt word templates are constructed for each entity relationship type combination. For the relationship type in the entity relationship type combination, the entity type connected to it is found from the ontology constructed in Step 2. These entity types are used to filter the corresponding entities from the entity set. For each entity relationship type combination, the respective prompt word templates and the filtered entities are used to extract the entity relationships using the large language model, obtaining structured triple data containing entity-relationship-entity.
[0013] Step 5: Select entities from the training dataset from Step 3. By designing a prompt word model, the large language model is guided to generate synonymous expressions for entities in the domain. Power experts correct the results to obtain similar entity pairs, and then shuffle them to obtain dissimilar entity pairs. The SBERT model is then fine-tuned using these similar and dissimilar entity pairs. For the entity set, the fine-tuned SBERT model and the depth-first search algorithm are used to obtain entity groups with similar semantic entities. The large language model then combines contextual information to achieve entity alignment.
[0014] Step 6: Store the entities and relationships extracted through the above steps into the graph database Neo4j, and preliminarily generate several initial entity nodes and several initial relationship edges; then, create the fault repair case content and the text block content cut from the fault repair case as case nodes and text block nodes respectively and store them in Neo4j, and establish connections between the case nodes and the initial entity nodes they contain, as well as between the text block nodes and the initial entity nodes they contain, ultimately generating several nodes and several edges, and then linking the image URL links generated by the cloud database to the corresponding nodes in the form of attributes.
[0015] Furthermore, the preprocessing in step one includes using professional document conversion tools to convert the power equipment fault inspection and repair PDF document into editable text information and image information, and the errors that occur during the conversion process are initially corrected with the help of a large language model and then manually verified.
[0016] Furthermore, the preprocessed text information described in step one is semantically segmented using a large language model to form independent text blocks, and text information that meets preset quality conditions is screened out through topic classification, specifically including: applying a large language model to segment each case file in the text information based on semantic integrity; classifying the segmented text blocks based on the topic classification method of the large language model, identifying and retaining the core defect information text blocks of the power equipment, and eliminating non-core defect information text blocks; at the same time, the text blocks of each case are standardized and numbered, and finally the rule matching script is applied to complete the classification.
[0017] Furthermore, step 2 specifically includes: Randomly extract several texts and provide m groups of different data samples to the large language model Get a separate set of analysis results Then, the large language model is used again to analyze and integrate the results, and the core entity type set ε and its relationship type set with high confidence are extracted. Constructing the initial body Then, electric power experts are introduced to review and improve the automatically generated entity types and their relationship types, and finally form an ontology that conforms to the characteristics of the electric power field.
[0018] Furthermore, the pre-labeling task in step 3 includes:
[0019] S31: The constructed prompt word template is refined into prompt = {Role, Entity Definition, Entity Recognition Rules, Task, Examples, Restriction, Text}; Role provides professional role positioning for the large language model; Entity Definition defines the meaning and boundaries of various entities in detail; Entity Recognition Rules customize recognition rules based on the characteristics of domain entities; in the Task section, CoT prompt word optimization technology is applied to design a progressive recognition process from simple to complex entity recognition, using XML tags for position marking; the Examples section provides input and output pairing examples; the Restriction section constrains the output format to ensure the consistency and processability of the annotation results; Text contains the input data to be processed;
[0020] S32: Design a correction script based on difference detection and rule matching, using the sequence alignment algorithm of the Python standard library difflib to correct the difference characters between the original text and the pre-annotated text. The correction process establishes the correspondence between the original text and the pre-annotated text by tracking the position index of the characters in the text, and analyzes the continuous matching areas between the two texts. While retaining the original annotation information, it corrects the mismatched parts of the text.
[0021] S33: After correction, the knowledge of power experts is used to systematically correct semantic deviations, inaccurate entity boundaries, and classification errors in the pre-labeling results.
[0022] Furthermore, based on the characteristics of fault domain data, the traditional deep learning small model for named entity recognition is optimized, specifically including the following steps:
[0023] S34: The traditional deep learning small model for named entities is a W2NER model, and the input layer of the W2NER model is replaced by the RoBERTa model instead of the original BERT model of the W2NER model;
[0024] S35: In the loss function part of the W2NER model, the Focal loss function is used to replace the cross entropy loss function used by the W2NER model:
[0025]
[0026] in, is the weight factor, Represents a set of predefined relationship types between characters. Represents a character pair (x i , x j ) is predicted as a relationship The probability score, Indicates whether it is a true label, represented by a binary vector, and N represents the number of characters in the sentence.
[0027] Furthermore, in step three, based on a tiered verification strategy, the large language models are stratified based on cost. A low-cost large language model is selected to apply the pre-labeled prompt word template to recognize the remaining 90% of the text information in step one. The recognition results are compared with those of the traditional deep learning small model to eliminate incorrect entities. Specifically:
[0028] S36: Utilize a pre-set, low-cost large language model, apply a pre-labeled prompt word template, and a trained traditional deep learning model to recognize the remaining text information. Analyze the recognition results. If the entities recognized by both are assigned a high confidence score, no additional verification is required. If there are discrepancies between the two entities, the higher-cost large language model with better performance is used for verification. The prompt word template used for verification is refined to prompt = {Role, EntityDefinition, Task, EntityInstances, Text}, where Role and EntityDefinition provide role positioning and entity type definition, respectively. In Task, a multi-dimensional evaluation framework is designed based on CoT prompt word optimization technology to evaluate the correct extraction of entities from three dimensions: "whether the entity corresponds to the given type," "whether the entity semantics are correct," and "whether the entity complies with the recognition rules." The EntityInstances section displays instances of each type of entity and their contextual information. The Text section represents the entity to be verified. Each verification entity is provided in the form of [e; c(e)], where e represents the entity and c(e) represents the entity's contextual information, consisting of the text block containing the entity and the group of text blocks before and after it.
[0029] Furthermore, step four specifically includes: extracting relationships based on the large language model and verified entities, performing differentiated processing on different types of relationships, forming entity relationship type combinations for relationship types with strong connections, and allowing the large language model to extract different entity relationship type combinations separately; customizing a specific extraction strategy for each entity relationship type combination, and the prompt word template of the extraction strategy is refined to prompt = {Role, Entity Definition, Task, Examples, Text}, where Role and Entity Definition provide role positioning and entity type definition respectively, Task provides specific extraction rules for different relationship combinations, the Examples section provides input and output examples, and Text contains the input data to be processed.
[0030] Furthermore, step five specifically includes:
[0031] S51: Build a dataset for SBERT fine-tuning based on the large language model. Select entities from the training dataset in step 3 and design a prompt strategy to guide the large language model to generate synonymous expressions of entities in the domain. The prompt word template for the synonymous expression is refined into prompt = {Role, Task, Text}, where Role provides role positioning; Task's specific requirements include synonyms or similar words, sentence structure adjustment, or expression modification for rewriting; Text contains the input data to be processed. Then, power experts conduct a secondary verification to obtain a fine-tuning dataset for fine-tuning the SBERT model.
[0032] S52: For the entity set, run the fine-tuned SBERT model to output entity pairs that exceed the specified cosine similarity threshold. Use the text block containing the entity as the entity's context information, and each entity as a node. Establish edge connections between entity pairs whose cosine similarity exceeds the specified threshold. Use the depth-first search algorithm to identify all connected subgraphs. Each subgraph represents a group of potentially related entity groups. Use the large language model to analyze each entity group to determine whether it is a similar entity and whether entity alignment is required. The corresponding prompt word template is refined as prompt = {Role, Task, Restriction, Examples, Text}, where Role provides role positioning; the Task part provides task requirements; the Restriction part specifies the output format; and Text contains the entity group to be processed and its context information.
[0033] Furthermore, a system for constructing a knowledge graph of power equipment faults includes:
[0034] Data preprocessing model: This involves acquiring text and image information related to power equipment troubleshooting and preprocessing the acquired information. The preprocessed image information is stored in a cloud database for easy access. The preprocessed text information is semantically segmented using a large language model to form independent text blocks. The text information that meets preset quality requirements is then filtered through topic classification.
[0035] Ontology construction module: This module uses a large language model to analyze and integrate the power equipment fault maintenance text information in the pre-processed data to obtain core entity types. This module then combines domain expert knowledge to complete the construction of the domain ontology, i.e., the knowledge graph model layer.
[0036] Named Entity Recognition Module: Construct a prompt word template and use the large language model to pre-label 10% of the text information in step one. Use rule matching and power experts to correct the pre-labeling results to obtain a training data set for the subsequent named entity recognition small model. Based on the characteristics of fault field data, optimize the traditional deep learning small model for named entity recognition and use the training data set for model training. After training, recognize the remaining 90% of the text information in step one. Design a hierarchical verification strategy to stratify the large language model based on cost. Select a low-cost large language model and apply the pre-labeled prompt word template to recognize the remaining 90% of the text information in step one. Compare its recognition results with those of the traditional deep learning small model to eliminate erroneous entities.
[0037] Relationship extraction module: Based on the characteristics of power equipment fault data and domain ontology, different entity relationship type combinations are designed. Different prompt word templates are constructed for different entity relationship type combinations. For the relationship types in the entity relationship type combinations, the entity types connected to them are found from the ontology constructed by the ontology construction model. These entity types are used to filter the corresponding entities from the entity set. For different entity relationship type combinations, the respective prompt word templates and the filtered entities are used to extract the relationship between the entities using a large language model, obtaining structured triple data containing entity-relationship-entity.
[0038] Knowledge fusion module: Select entities from the training dataset in step 3, and guide the large language model to generate synonymous expressions of entities in the field by designing a prompt word model. Power experts correct the results to obtain similar entity pairs, and shuffle them to obtain dissimilar entity pairs. The SBERT model is fine-tuned and trained using similar and dissimilar entity pairs. For entity sets, the fine-tuned SBERT model and the depth-first search algorithm are used to obtain entity groups with similar semantic entities. The large language model then combines contextual information to achieve entity alignment.
[0039] Graph Storage Module: After the data preprocessing model, ontology construction module, named entity recognition module, relationship extraction module, and knowledge fusion module, the extracted entities and relationships are stored in the graph database Neo4j, initially generating several initial entity nodes and initial relationship edges. Subsequently, the troubleshooting case content and the text blocks segmented from the troubleshooting case are created as case nodes and text block nodes, respectively, and stored in Neo4j. Connections are established between the case nodes and the initial entity nodes they contain, as well as between the text block nodes and the initial entity nodes they contain. Ultimately, several nodes and edges are generated. The image URL links generated by the cloud database are then linked to the corresponding nodes as attributes.
[0040] 3.Beneficial effects:
[0041] (1) In the method for constructing a knowledge graph of power equipment faults based on the collaboration of large and small models provided by the present invention, the long text of power equipment fault inspection and maintenance is processed into short texts with clear semantic connections from the two dimensions of text segmentation and text classification based on the large language model, providing a high-quality data source for tasks such as named entity recognition.
[0042] (2) The present invention provides a method for constructing a knowledge graph of power equipment faults based on the collaboration of large and small models. It designs an initial ontology construction method based on a large language model to guide the construction of the final ontology and reduce the influence of subjective factors in the ontology construction process.
[0043] (3) The method for constructing a knowledge graph of power equipment failure based on the collaboration of large and small models provided by the present invention integrates a large language model, rule matching and power expert knowledge to achieve efficient pre-labeling of named entity recognition model training data, reduce the cost of manual labeling, and realize the adaptive optimization of the W2NER model in the field of power equipment failure. In addition, without relying on an external knowledge base, the entity recognition results are verified based on a hierarchical verification strategy and a large language model, reducing the error propagation of recognition results to downstream tasks. In the process of relationship extraction, the correlation between different relationships is fully considered, and the strongly correlated relationships are integrated and extracted, effectively improving the accuracy of triple extraction of power equipment fault text.
[0044] (4) The method for constructing a knowledge graph of power equipment faults based on the collaboration of large and small models provided by the present invention considers the verification and matching mechanism of entity context based on the large language model and SBERT model design, avoiding the possible erroneous fusion caused by simply relying on the specified cosine similarity threshold.
[0045] (5) The method for constructing a knowledge graph of power equipment faults based on the collaboration of large and small models provided by the present invention carries out the construction of several prompt word templates for the entire process of constructing the knowledge graph of power equipment faults, and constructs a framework based on the collaboration of a large language model and a traditional deep learning small model. This framework ensures the integrity and accuracy of the knowledge graph of power equipment faults. The empirical application in the field of power transformer faults proves the significant advantages of this method over the existing technology, and provides a reusable technical solution for the construction of other knowledge graphs of power equipment faults. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is the overall flow chart for constructing a knowledge graph of power equipment faults according to the present invention;
[0047] Figure 2 A schematic diagram of text segmentation during data preprocessing in a specific embodiment;
[0048] Figure 3 This is a diagram showing text segmentation results during data preprocessing in a specific embodiment;
[0049] Figure 4 A schematic diagram of ontology construction based on a large language model and expert knowledge embedding in a specific embodiment;
[0050] Figure 5 This is a schematic diagram of constructing a knowledge graph ontology for power transformer faults in a specific embodiment;
[0051] Figure 6 This is an example of a partial prompt word template pre-labeled based on data of a large language model in step three of the specific embodiment;
[0052] Figure 7 This is an example of a partial prompt word template for entity verification based on a large language model in step three of the specific embodiment;
[0053] Figure 8 The experimental results of different large language model entities verification in step 3 of the specific embodiment;
[0054] Figure 9 The experimental results of entity verification based on the layered verification strategy in step 3 of the specific embodiment are as follows;
[0055] Figure 10 This is an example of a partial prompt word template for relation extraction based on a large language model in a specific embodiment;
[0056] Figure 11 This is an example of a prompt word template for entity alignment based on a large language model in a specific embodiment;
[0057] Figure 12 This is an example of a defect diagram in the power transformer fault knowledge graph in a specific embodiment;
[0058] Figure 13 This is an example of knowledge fusion in the knowledge fusion module in a specific embodiment;
[0059] Figure 14 This is an example of the quantity distribution of each node and edge in the power transformer fault knowledge graph before and after knowledge fusion in a specific embodiment. DETAILED DESCRIPTION
[0060] The present invention will be described in detail below with reference to specific embodiments.
[0061] As attached Figure 1 As shown, the method for constructing a knowledge graph of power equipment faults based on collaboration between large and small models is characterized by:
[0062] Step 1: Data preprocessing: This involves obtaining text and image information related to power equipment troubleshooting and preprocessing the information. The preprocessed image information is stored in a cloud database for easy access, and the preprocessed text information is semantically segmented using a large language model to form independent text blocks. The text information that meets preset quality requirements is then filtered through topic classification.
[0063] Step 2: Use the large language model to analyze and integrate the power equipment fault maintenance text information in the pre-processed data to obtain core entity types, and combine the domain expert knowledge to construct the domain ontology of the knowledge graph model layer;
[0064] Step 3: Construct a prompt word template and use the large language model to pre-label 10% of the text information in step 1. Use rule matching and power experts to correct the pre-labeling results to obtain a training data set for the subsequent traditional deep learning small model for named entity recognition; Based on the characteristics of the fault field data, optimize the traditional deep learning small model for named entity recognition and use the training data set for model training. After training, recognize the remaining 90% of the text information in step 1; Design a hierarchical verification strategy, stratify the large language model based on cost, select a low-cost large language model and apply the pre-labeled prompt word template to recognize the remaining 90% of the text information in step 1, compare its recognition results with those of the traditional deep learning small model, and eliminate erroneous entities;
[0065] Step 4: Based on the characteristics of power equipment fault data and the domain ontology, different entity relationship type combinations are designed, and corresponding prompt word templates are constructed for different entity relationship type combinations. For the relationship types in the entity relationship type combinations, the entity types connected to them are found from the ontology constructed in Step 2. These entity types are used to filter the corresponding entities from the entity set obtained in Step 3. For different entity relationship type combinations, the respective prompt word templates and the filtered entities are used to extract the entity relationships using the large language model, obtaining structured triple data containing entity-relationship-entity.
[0066] Step 5: Select entities from the training dataset from Step 3. By designing a prompt word model, the large language model is guided to generate synonymous expressions for entities in the domain. Power experts correct the results to obtain similar entity pairs, and then shuffle them to obtain dissimilar entity pairs. The SBERT model is then fine-tuned using these similar and dissimilar entity pairs. For the entity set, the fine-tuned SBERT model and the depth-first search algorithm are used to obtain entity groups with similar semantic entities. The large language model then combines contextual information to achieve entity alignment.
[0067] Step 6: Store the entities and relationships extracted through the above steps into the graph database Neo4j, and preliminarily generate several initial entity nodes and several initial relationship edges; then, create the fault repair case content and the text block content cut from the fault repair case as case nodes and text block nodes respectively and store them in Neo4j, and establish connections between the case nodes and the initial entity nodes they contain, as well as between the text block nodes and the initial entity nodes they contain, ultimately generating several nodes and several edges, and then linking the image URL links generated by the cloud database to the corresponding nodes in the form of attributes. Specific embodiment 1
[0069] like Figure 2 The figure shows a schematic diagram of the data preprocessing process, i.e., text segmentation and text block classification based on a large language model in step 1. In this embodiment, taking the field of power transformers as an example, a professional document conversion tool was used to process the PDF electronic documents "Analysis of Typical Transformer Fault Cases" and "Typical Problems and Analysis of Transformer Maintenance Processes" to obtain initial data. The obtained image information was systematically stored using an image server, and an image URL address was generated for subsequent access, complementing the text content. For the obtained text information, taking into account the inevitable errors that occur during document OCR recognition and conversion, prompts such as "Correct typos, missing words, and incorrect punctuation in the text" were constructed. The large language model was used to perform preliminary corrections for the errors, followed by manual verification.
[0070] From a text structure perspective, the fault case texts generated through the above process can be divided into two categories: complex case texts and short case texts. Complex case texts contain diverse content, such as defect descriptions, equipment information, and cause analysis, interspersed with a large amount of redundant information unrelated to the core defect. This increases the difficulty of subsequent triple extraction and also impacts extraction efficiency and accuracy. Therefore, a large language model is first applied to intelligently segment each case text based on semantic integrity. Then, drawing on the concept of topic classification, the segmented text blocks are classified, identifying and retaining core power transformer defect information blocks while eliminating non-core defect information blocks. The prompt template for text segmentation and block classification is refined into prompt = {Role, Task, Examples, Text}. Role defines the role of the large language model, Task incorporates chained thinking (CoT) prompt technology to present specific task steps, Examples provides standardized examples, and Text contains the input data to be processed. Given that large language models are essentially generative models rather than classification models, to ensure the reliability of the classification results, both text segmentation and text block classification prompts require that the text blocks of each case be standardized and numbered. These numbers are retained in the output results so that the rule matching script can be applied later to complete the classification.
[0071] Figure 3 The figure below shows the results of text segmentation based on a large language model during data preprocessing. 123 case studies, totaling 70,853 characters, were collected from the articles "Analysis of Typical Transformer Fault Cases" and "Typical Problems and Solutions in Transformer Maintenance Processes." After text segmentation in the data preprocessing module, the original corpus was divided into 501 initial text blocks. Further text block classification eliminated blocks containing no valid information, resulting in 431 text blocks. Of these, 74.72% of the text blocks contained fewer than 200 characters, and 95.82% contained fewer than 500 characters. This distribution strongly demonstrates that the text segmentation strategy achieved the desired results and laid a good foundation for subsequent triple extraction. For the few long text blocks exceeding 500 characters, given the model's limited extraction capabilities for long text, a Python script was designed to perform secondary processing: using natural paragraphs (line breaks) as the basic segmentation unit, adjacent paragraphs were grouped together to ensure that each combined text block did not exceed the 500-character threshold. When the cumulative number of characters in a paragraph exceeds a threshold, it automatically undergoes block processing, ensuring that each newly generated text block maintains semantic coherence while also meeting the model's processing capabilities. Ultimately, 455 text blocks were obtained, providing the data foundation for subsequent named entity recognition and relation extraction experiments.
[0072] In the data preprocessing module, the experimental results of text block classification based on the large language model are shown in Table 1, and comparative experiments are conducted on multiple large language models.
[0073] Table 1 Experimental results of text block classification based on large language model
[0074]
[0075]
[0076] Claude 3.7Sonnet performed best, achieving an accuracy of 95.61%. Models Deepseek-R1 and GLM-4-Plus also performed well. The accuracy of the four models all reached above 90%, verifying the effectiveness and feasibility of the text processing prompt strategy constructed by the present invention. Specific embodiment 2
[0078] like Figure 4 The figure shows the schematic diagram of ontology construction based on large language model and expert knowledge embedding in the ontology construction process. A number of texts are randomly sampled from the dataset and fed into the large language model to complete the analysis and summary of the domain entity types and their relationship types. Considering the inherent randomness of the large language model and the differences in the coverage of entity types and their relationship types in different input texts, a single analysis will have bias. To reduce the impact of this bias, a multi-sampling analysis strategy is adopted - providing m groups of different data samples to the large language model. Get a separate set of analysis results Then, the large language model is used again to analyze and integrate these results, and the core entity type set ε and its relationship type set with high confidence are extracted. Constructing the initial body Finally, domain experts are introduced to review and improve the automatically generated entity types and their relationship types, and finally form an ontology that conforms to the characteristics of the domain.
[0079] Specifically, in this embodiment, for the power transformer defect text, 5 groups of different data samples are provided to the large language model, and the result sets are obtained by independent analysis. Then, these five sets of analysis results are conceptually integrated through a large language model to obtain the initial ontology. Finally, the initial ontology is modified by combining expert knowledge to obtain During this process, after expert evaluation, the entity types generated by the large language model, "transformer equipment, transformer components, fault phenomena, maintenance measures, fault causes, and detection methods" are conceptually highly reliable, but require professional adjustments in specific term definitions and relationship expressions. For example, in the actual description text, there are potential problems that have not yet caused complete failure of the equipment but have already appeared. There is a semantic deviation in summarizing them with the word "fault", so a more accurate expression should be "defect". After systematic construction and revision, 8 core entity types were finally established, and 11 relationship types were defined: "possess", "include", "appear", "occur", "implement", "carry out", "detect", "test inspection result legend is", "lead to", "take" and "defect legend is", such as Figure 5 shown. Specific embodiment 3:
[0081] like Figure 6 As shown in the figure, this is an example of a prompt word template for pre-labeling data based on a large language model in the named entity recognition module. The prompt word template is refined as prompt = {Role, Entity Definition, Entity RecognitionRules, Task, Examples, Restriction, Text}, where Role provides professional role positioning for the large language model, Entity Definition defines the meaning and boundaries of various entities in detail, and Entity Recognition Recognition rules must be customized based on the characteristics of domain entities. For power transformer defect text, recognition rules must be designed for its unique nested entity structure. Indiscriminately recognizing all nested entities will result in information redundancy. For example, in the text segment "Cracks and ruptures in the lower cone of the bushing cause oil leakage," "Cracks and ruptures in the lower cone of the bushing" should be labeled as a "defect cause" type entity. The "Cracks and ruptures in the lower cone of the bushing" (part entity) and "Cracks and ruptures" (defect phenomenon entity) contained within this entity should also be recognized. Furthermore, the defect legend entity "Broken aluminum tube" also contains part and defect phenomenon type entities, which have already been identified. Re-labeling them would waste computational resources and potentially introduce errors. Therefore, the Entity Recognition Rules are configured to ignore some nested entities in figure and table annotations. In the Task section, CoT word optimization technology is applied to design a progressive recognition process from simple to complex recognized entities. Examples of input and output pairings are provided in the Examples section. In the Restriction section, you can constrain the output format to ensure the consistency and processability of the annotation results. The Text section contains the content to be processed.
[0082] In the named entity recognition module, the pre-annotation experimental results based on the large language model are shown in Table 2. Four well-known large language models in the industry: Claude 3.7Sonnet, Gemini 2.0Flash, GPT-4o and the open source model DeepSeek-V3 were selected for experimental analysis. At the same time, two different entity annotation forms were compared: XML tag form " <label>、< / label> ” and the position index format “[Head_Index, Tail_Index]”. Because the output of the large language model may contain problems such as incorrect addition, omission, or replacement of characters compared to the original text, the output content is corrected using a rule matching script after the large language model completes pre-annotation.
[0083] Table 2 Pre-labeling experimental results based on large language model
[0084]
[0085]
[0086] Labeling significantly outperforms position-indexed annotation. Position-indexed annotation requires not only that the model recognize the corresponding entity but also that it accurately calculate character positions. This hybrid task poses a greater challenge to large language models, making positioning errors more likely to occur when processing long texts. Labeling, on the other hand, embeds tags directly within the text, enabling the model to more naturally identify entity boundaries and reducing position calculation errors. Furthermore, the four large language models all achieved P, R, and F1 values exceeding 50%, which can significantly reduce manual annotation time and validate the feasibility of pre-annotation based on large language models.
[0087] In the named entity recognition module, the named entity recognition experimental results are shown in Table 3, which compares and analyzes the commonly used BERT-BiLSTM-CRF model, PURE-NER model and multiple W2NER variant models. Among them, BERT-BiLSTM-CRF is a commonly used model in named entity recognition tasks, and PURE-NER is the named entity recognition part of the PURE model. In the W2NER variant model, W2NER (RoBERTa) means replacing the input layer with the RoBERTa model; W2NER (Focal) means replacing the model's loss function with Focal loss; W2NER (RoBERTa+Focal) is the complete optimization model proposed by this invention, which integrates two aspects of optimization at the same time. The main parameters set are Epoch 40, Optimizer Adam, batch_size 2, and Learning_rate 1x10 -3 .
[0088] Table 3 Experimental results of named entity recognition
[0089]
[0090]
[0091] The W2NER model's comprehensive performance metric, F1, improved by 10.16% and 34.98% compared to PURE-NER and BERT-BiLSTM-CRF, respectively, demonstrating significant advantages. Regarding W2NER model optimization, replacing the input layer with the RoBERTa model increased the F1 value by 0.31%, demonstrating the effectiveness of RoBERTa's output of word-level character vectors. Replacing the loss function with the Focal loss increased the F1 value by 0.7%, validating Focal loss's advantages in addressing imbalanced classification. When both the RoBERTa model and Focal loss were used, the model achieved an F1 value of 79.86%, a 2.24% improvement over the original W2NER model. This result demonstrates a synergistic effect between the two optimization strategies, and their combined application further amplifies the overall performance improvement of the model.
[0092] like Figure 7 The figure shows an example of a prompt template for the entity verification portion of the named entity recognition module based on a large language model. The prompt template is refined as prompt = {Role, Entity Definition, Task, Entity Instances, Text}, where Role and Entity Definition represent role positioning and entity type definition, respectively. Within Task, a multi-dimensional evaluation framework is designed based on CoT prompt optimization technology to assess the correct extraction of entities based on three dimensions: "whether the entity corresponds to the given type," "whether the entity semantics are correct," and "whether the entity conforms to the recognition rules." The EntityInstances section displays instances of each entity type and their contextual information. The Text section represents the entity to be verified. Each verification entity is provided in the form of [e; c(e)], where e represents the entity and c(e) represents the entity's contextual information, consisting of the text block containing the entity and the preceding and following text blocks. This allows the large language model to be evaluated in real-world contexts.
[0093] like Figure 8Figure 2 shows the experimental results of entity verification using different large language models in the named entity recognition module. Using all entities output by the W2NER (RoBERTa+Focal) model as the base dataset, these entities were verified using multiple large language models, with Baseline representing the original output of the W2NER (RoBERTa+Focal) model. The results show that, with the exception of DeepSeek-V3, all other large language models improved their accuracy after verification, indicating that they successfully eliminated some erroneous entities. This result has a positive impact on downstream tasks such as relation extraction and can effectively reduce error propagation. From the perspective of the comprehensive performance indicator F1 value, only Claude 3.7Sonnet achieved an improvement in the F1 value after verification. Based on this result, Claude 3.7Sonnet was selected as the entity verification model in the subsequent layered verification strategy.
[0094] like Figure 9The figure shows the experimental results of entity verification based on the hierarchical verification strategy in the named entity recognition module. From a cost perspective, the input cost of Gemini 2.0Flash is $0.1 per million tokens and the output is $0.4, which saves $2.9 and $14.6 respectively compared with Claude3.7Sonnet. DeepSeek-V3, as an open source model, can further reduce costs. To reduce the verification cost, Gemini 2.0Flash and DeepSeek-V3 are used in the hierarchical verification strategy to perform entity recognition using the pre-labeled prompt word templates. The results are compared with the output of the W2NER (RoBERTa+Focal) model. After filtering out inconsistent entities, Claude 3.7Sonnet is used to perform entity verification experiments. Gemini2.0Flash-recognition and DeepSeek-V3-recognition represent the results of entity verification after screening based on Gemini 2.0Flash and DeepSeek-V3, respectively. Claude 3.7Sonnet represents the results of verification for all entities. Baseline represents the raw output of the W2NER (RoBERTa+Focal) model. The results show that Gemini 2.0Flash-recognition and DeepSeek-V3-recognition achieve precision improvements of 1.57% and 1.45%, respectively, compared to the baseline, demonstrating that this method can effectively identify and remove erroneous entities. Although recall decreases by 0.42% and 1.16%, respectively, the F1 value increases by 0.48% and 0.04%, respectively, demonstrating that the improvement in precision from removing erroneous entities outweighs the loss in recall caused by mistakenly deleting correct entities. Among them, Gemini 2.0Flash-recognition improves the F1 value by 0.24% compared with the method of verifying all entities, which means that the layered verification strategy can significantly reduce the API call cost while maintaining the verification effect; the recall rate improves by 0.23% compared with the method of verifying all entities, indicating that the use of Gemini 2.0Flash for recognition filters out a small number of entities that are difficult for Claude 3.7Sonnet to judge. This method effectively integrates the knowledge advantages of different models. Specific embodiment 4:
[0096] like Figure 10The following figure shows an example of a prompt word template for relation extraction based on a large language model during the relation extraction process. In the power transformer domain, eight entity relationship type combinations were designed. Extraction strategies were customized based on the characteristics of different entity relationship type combinations. The prompt word template was refined into prompt = {Role, Entity Definition, Task, Examples, Text}. Role and Entity Definition represent role positioning and entity type definition, respectively. Task incorporates domain-specific extraction rules. For example, different "part" entities have a hierarchical relationship, connected by a "contains" relationship. The "power transformer" entity is connected to the "part" entity through an "owns" relationship. To avoid redundancy, the "power transformer" entity only needs to establish an "owns" relationship with the highest-level "part" entity and does not need to be connected to part entities at all levels. In this case, the "owns" and "contains" relationships are strongly correlated and need to be extracted simultaneously. Furthermore, the definitions of "owns" and "contains" must clearly indicate that "part" type entities are hierarchical, and the highest-level "part" entity should be directly connected to the "power transformer" type entity. The Examples section provides input and output pairing examples, and the Text section contains the content to be extracted.
[0097] In the relation extraction module, the relation extraction experimental results are shown in Table 4, which compares and analyzes the commonly used deep learning models: BERT, PURE-RE and BERT-BiGRU-Attention, where PURE-RE represents the relation extraction part of the PURE model. In addition, in the named entity recognition task, existing studies mainly use three forms to represent position information: XML tag marking, position index and no output of position information. In order to verify the effectiveness of position information in relation extraction based on large language models, three position information representation schemes are designed for comparative experiments: (1) providing the entity's position index "[Head_Index, Tail_Index]" as additional information; (2) using the XML tag " <label>、< / label> "Directly mark the entity location in the original text; (3) only provide the entity content without including the location information.
[0098] Table 4. Relationship extraction experimental results
[0099] Model or method F1(%) P(%) R(%) BERT 51.83 58.12 46.77 BERT-BiGRU-Attention 33.99 39.29 29.95 PURE-RE 57.98 80.11 45.43 Based on a large language model (no position information provided) 78.67 72.53 85.95 Based on a large language model (marked with XML tags) 79.99 73.59 87.61 Based on a large language model (using position indexing) 80.05 74.04 87.13
[0100] The relationship extraction method based on a large language model in this paper achieves F1 improvements of 26.89%, 20.13%, and 35.76% compared to BERT, PURE-RE, and BERT-BiGRU-Attention, respectively. This method outperforms commonly used deep learning models and demonstrates certain effectiveness. Regarding location information, methods using location indexes and XML tagging outperformed methods that did not provide location information in terms of F1, precision, and recall, demonstrating the positive impact of location information on relationship extraction tasks.
[0101] In the relation extraction module, the experimental results of relation extraction with or without considering the correlation between relations are shown in Table 5, which compares the effects of 11 relations under the two paradigms of independent extraction and considering correlation.
[0102] Table 5. Experimental results of relation extraction with and without considering the correlation between relations.
[0103] method F1(%) P(%) R(%) Based on large language model (independent extraction) 71.18 58.99 89.72 Based on large language model (taking relevance into account) 80.05 74.04 87.13
[0104] Although the independent extraction method has a recall rate 2.59% higher than the method considering correlation, its precision is too low, 15.05% lower than the latter. This shows that the independent extraction method tends to mechanically identify more entity pairs as having relationships, lacks the ability to accurately judge potential relationships, and leads to a large number of incorrect predictions. This combination of high recall and low precision reflects that the model over-predicts relationships rather than truly understanding the semantic associations between entities. In contrast, when the correlation between relationships is taken into account, the F1 value is significantly improved by 8.87%, achieving a better balance between precision and recall, proving that for relationship types with strong correlation, the simultaneous extraction strategy in the present invention is reasonable, can effectively capture the complementary information and constraints between relationships, and improve the accuracy of extraction. Specific embodiment 5:
[0106] Figure 11 As shown in the figure, an example of the prompt word template for the entity alignment part based on the large language model in the knowledge fusion process is shown. The prompt word template is refined as prompt = {Role, Task, Restriction, Examples, Text}, where Role provides role positioning, Task provides task requirements, Restriction specifies the output format, Examples provides input and output examples, and Text contains the entity group to be processed and its context information.
[0107] In the knowledge fusion module, the experimental results of the proportion of entities correctly expressed in different methods are shown in Table 6. An entity alignment strategy based on SBERT is designed: entity pairs are processed in order from low to high according to the cosine similarity ranking. When fusing two entities in an entity pair, one of the entities is randomly selected as the alignment object. Considering that the expression form after the initial fusion still needs to be further integrated, it is necessary to perform a secondary cosine similarity calculation on these entities and then align the entities.
[0108] Table 6 Experimental results of the proportion of entities correctly expressed by different methods
[0109]
[0110]
[0111] After SBERT-based entity alignment, 359 entities received correct representations. Considering that the representations after the initial fusion still needed further integration, a second cosine similarity calculation was performed on these entities, generating another 60 entity pairs. After further fusion, the representations of 51 entities were successfully updated, but only 336 entities ultimately received correct representations. This indicates that after the second fusion, some entities had their correct representations transformed into incorrect representations due to error propagation. A total of 164 entity groups were obtained using a depth-first search algorithm. After using a large language model to judge these entities, 364 entities received correct representations after entity alignment, achieving more accurate representations than fusion methods based solely on SBERT. Specific embodiment 6:
[0113] like Figure 12The figure below shows an example of a defect legend in a power transformer fault knowledge graph. After data preprocessing, ontology construction, named entity recognition, and relationship extraction, the extracted information was stored in the Neo4j graph database, generating 1,959 initial entity nodes and 2,284 initial relationship edges. Subsequently, the case content and the text blocks separated from the case were created as case nodes and text block nodes, respectively, and imported into Neo4j. Connections were established between the case nodes and the initial entity nodes they contained, as well as between the text block nodes and the initial entity nodes they contained. This ultimately generated 2,533 nodes and 8,123 edges. Among them, the introduction of text block nodes facilitates subsequent entity alignment tasks, while the setting of case nodes enhances the application value of the power transformer fault knowledge graph. The image URL link generated by the image bed is imported into the corresponding "defect legend" type node or "test inspection result legend" type node in the form of attributes. This multimodal knowledge representation method makes up for the limitations of pure text descriptions in expressing technical details, so that complex equipment defect characteristics, abnormal conditions, test results and other physical phenomena that are difficult to accurately describe in words can be presented through intuitive visual information, providing downstream knowledge services with visual explanation capabilities and improving the performance of the power transformer fault knowledge graph in technical consulting scenarios.
[0114] like Figure 13 The figure below shows an example of knowledge fusion in the knowledge fusion module. The connections between text block nodes and the entity nodes they contain are used to quickly obtain contextual information about the entities. After entity alignment based on a large language model, domain experts validated the results. The Neo4j APOC plugin then merged related nodes and edges to complete the knowledge fusion. "Low voltage short circuit test" and "low voltage short circuit test" were identified as synonymous entities and fused. Because the expression "low voltage short circuit test" better aligns with professional standards, it was retained. During the fusion process, connections between commonly associated nodes in the text block "2.14.2.2 Low voltage short...", "110kV transformer," and "Case 7" were preserved. All original connections in "low voltage short circuit test" were transferred to the "low voltage short circuit test" node to ensure knowledge integrity and consistency. The resulting power transformer fault knowledge graph was optimized to 2,368 nodes and 7,921 edges.
[0115] like Figure 14Figure 1 shows an example of the distribution of nodes and edges in the power transformer fault knowledge graph before and after knowledge fusion. (a) shows the node distribution, (b) shows the edge distribution of the relationship types generated during ontology construction, (c) shows the edge distribution of text block nodes, and (d) shows the edge distribution of case nodes. Among edge types, edges of the form "text block_part" represent the edge connecting the initial entity node in each text block to the text block node, while edges of the form "case_part" represent the edge connecting the initial entity node in each case to the case node.
[0116] Although the present invention has been disclosed above in terms of preferred embodiments, they are not intended to limit the present invention. Anyone skilled in the art can make various changes or modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection defined by the claims of this application.
Claims
1. A method for constructing a knowledge graph of power equipment faults based on collaboration between large and small models, characterized by: Step 1: Data preprocessing; It includes obtaining text information and image information of power equipment fault inspection and maintenance, and pre-processing the obtained information; The pre-processed image information is stored in the cloud database for easy access. The pre-processed text information is semantically segmented using a large language model to form independent text blocks, and the text information that meets the preset quality conditions is screened through topic classification. Step 2: Use the large language model to analyze and integrate the power equipment fault maintenance text information in the pre-processed data to obtain core entity types, and combine the domain expert knowledge to construct the domain ontology of the knowledge graph model layer; Step 3: Construct a prompt word template and use the large language model to pre-label 10% of the text information from Step 1. Use rule matching and power experts to correct the pre-labeled results to obtain a training dataset for the subsequent traditional deep learning small model for named entity recognition. Based on the characteristics of fault domain data, the traditional deep learning small model for named entity recognition was optimized and trained using a training dataset. After training, the model was used to identify the remaining 90% of the text information in step one. A tiered verification strategy was designed to stratify the large language models based on cost. A low-cost large language model was selected and used with a pre-labeled prompt word template to identify the remaining 90% of the text information in step one. The recognition results were compared with those of the traditional deep learning small model, and the incorrect entities were eliminated to obtain the entity set. Step 4: Based on the characteristics of power equipment fault data and the domain ontology, different entity relationship type combinations are designed, and corresponding prompt word templates are constructed for each entity relationship type combination. For the relationship type in the entity relationship type combination, the entity type connected to it is found from the ontology constructed in Step 2. These entity types are used to filter the corresponding entities from the entity set. For each entity relationship type combination, the respective prompt word templates and the filtered entities are used to extract the entity relationships using the large language model, obtaining structured triple data containing entity-relationship-entity. Step 5: Select entities from the training dataset from step 3 and design a prompt word model to guide the large language model to generate synonymous expressions for entities in the domain. Power experts correct the results to obtain similar entity pairs and shuffle them to obtain dissimilar entity pairs. The similar and dissimilar entity pairs are used to fine-tune the SBERT model. For entity sets, we use a fine-tuned SBERT model and a depth-first search algorithm to obtain entity groups with similar semantic entities. Then, a large language model combines contextual information to achieve entity alignment. Step 6: Store the entities and relationships extracted through the above steps into the graph database Neo4j, and preliminarily generate several initial entity nodes and several initial relationship edges; Subsequently, the troubleshooting case content and the text block content cut from the troubleshooting case are created as case nodes and text block nodes respectively and stored in Neo4j. At the same time, connections are established between the case nodes and the initial entity nodes they contain, as well as between the text block nodes and the initial entity nodes they contain. Finally, several nodes and edges are generated, and the image URL links generated by the cloud database are linked to the corresponding nodes in the form of attributes.
2. The method for constructing a knowledge graph of power equipment faults according to claim 1, characterized in that: The preprocessing in step one includes using professional document conversion tools to convert the power equipment fault inspection and repair PDF document into editable text information and image information, and using a large language model to initially correct errors that occur during the conversion process, followed by manual fine-tuning.
3. The method for constructing a knowledge graph of power equipment faults according to claim 2, characterized in that: The preprocessed text information described in step 1 is semantically segmented using a large language model to form independent text blocks, and text information that meets preset quality conditions is screened out through topic classification, specifically including: applying a large language model to segment each case file in the text information based on semantic integrity; classifying the segmented text blocks based on the topic classification method of the large language model, identifying and retaining the core defect information text blocks of the power equipment, and eliminating non-core defect information text blocks; at the same time, the text blocks of each case are standardized and numbered, and finally the rule matching script is applied to complete the classification.
4. The method for constructing a knowledge graph of power equipment faults according to claim 3, characterized in that: Step 2 specifically includes: Randomly extract several texts and provide m groups of different data samples to the large language model Get a separate set of analysis results Then, the large language model is used again to analyze and integrate the results, and the core entity type set ε and its relationship type set with high confidence are extracted. Constructing the initial body Then, electric power experts are introduced to review and improve the automatically generated entity types and their relationship types, and finally form an ontology that conforms to the characteristics of the electric power field.
5. The method for constructing a knowledge graph of power equipment faults according to claim 4, characterized in that: The pre-labeling tasks completed in step 3 include: S31: Refine the prompt word template into prompt = {Role, Entity Definition, Entity Recognition Rules, Task, Examples, Restriction, Text}; Role provides professional role positioning for the large language model; Entity Definition defines the meaning and boundaries of various entities in detail; Entity Recognition Rules customize recognition rules based on the characteristics of domain entities; In the Task section, apply CoT prompt word optimization technology to design a progressive recognition process from simple to complex entity recognition, using XML tags for position marking; The Examples section provides input and output pairing examples; The Restriction section constrains the output format to ensure the consistency and processability of the annotation results; Text contains the input data to be processed; S32: Design a correction script based on difference detection and rule matching, using the sequence alignment algorithm of the Python standard library difflib to correct the difference characters between the original text and the pre-annotated text. The correction process establishes the correspondence between the original text and the pre-annotated text by tracking the position index of the characters in the text, and analyzes the continuous matching areas between the two texts. While retaining the original annotation information, it corrects the mismatched parts of the text. S33: After correction, the knowledge of power experts is used to systematically correct semantic deviations, inaccurate entity boundaries, and classification errors in the pre-labeling results.
6. The method for constructing a knowledge graph of power equipment faults according to claim 5, characterized in that: Based on the characteristics of fault domain data, the traditional deep learning small model for named entity recognition is optimized. The specific steps include: S34: The traditional deep learning small model for named entities is a W2NER model, and the input layer of the W2NER model is replaced by the RoBERTa model instead of the original BERT model of the W2NER model; S35: In the loss function part of the W2NER model, the Focal loss function is used to replace the cross entropy loss function used by the W2NER model: in, is the weight factor, Represents a set of predefined relationship types between characters. Represents a character pair (x i , x j ) is predicted as a relationship The probability score, Indicates whether it is a true label, represented by a binary vector, and N represents the number of characters in the sentence.
7. The method for constructing a knowledge graph of power equipment faults according to claim 6, characterized in that: In step three, based on a tiered verification strategy, the large language models are stratified based on cost. A low-cost large language model is selected and used to recognize the remaining 90% of the text information from step one using a pre-labeled prompt word template. The recognition results are compared with those of a traditional deep learning small model, and the incorrect entities are eliminated to obtain an entity set. Specifically: S36: Using a pre-annotated prompt word template and a trained traditional deep learning model, the remaining text information is recognized. The recognition results are analyzed. If the entities identified by both are assigned a high confidence score, no additional verification is required. If there are discrepancies between the two entities, the higher-performance, higher-cost large language model is used for verification. The prompt word template used for verification is refined to prompt = {Role, Entity Definition, Task, Entity Instances, Text}, where Role and Entity Definition provide role positioning and entity type definition, respectively. In Task, a multi-dimensional evaluation framework is designed based on CoT prompt word optimization technology to evaluate the correct extraction of entities based on three dimensions: "whether the entity corresponds to the given type," "whether the entity semantics are correct," and "whether the entity complies with the recognition rules." The Entity Instances section displays instances of each type of entity and their contextual information. The Text section represents the entity to be verified. Each verification entity is provided in the form of [e; c(e)], where e represents the entity and c(e) represents the entity's contextual information, which is composed of the text block containing the entity and the group of text blocks before and after it.
8. The method for constructing a knowledge graph of power equipment faults according to claim 7, characterized in that: Step four specifically includes: extracting relationships based on the large language model and verified entities, performing differentiated processing on different types of relationships, forming entity relationship type combinations for relationship types with strong connections, and allowing the large language model to extract different entity relationship type combinations separately; customizing a specific extraction strategy for each entity relationship type combination, and refining the prompt word template of the extraction strategy as prompt = {Role, Entity Definition, Task, Examples, Text}, where Role and Entity Definition provide role positioning and entity type definition respectively, Task provides specific extraction rules for different relationship combinations, the Examples section provides input and output examples, and Text contains the input data to be processed.
9. The method for constructing a knowledge graph of power equipment faults according to claim 8, characterized in that: Step 5 specifically includes: S51: Build a dataset for SBERT fine-tuning based on the large language model. Select entities from the training dataset in step 3 and design a prompt strategy to guide the large language model to generate synonymous expressions of entities in the domain. The prompt word template for the synonymous expression is refined into prompt = {Role, Task, Text}, where Role provides role positioning; Task's specific requirements include synonyms or similar words, sentence structure adjustment, or expression modification for rewriting; Text contains the input data to be processed. Then, power experts conduct a secondary verification to obtain a fine-tuning dataset for fine-tuning the SBERT model. S52: For the entity set, run the fine-tuned SBERT model to output entity pairs that exceed the specified cosine similarity threshold. Use the text block containing the entity as the entity's context information, and each entity as a node. Establish edge connections between entity pairs whose cosine similarity exceeds the specified threshold. Use the depth-first search algorithm to identify all connected subgraphs. Each subgraph represents a group of potentially related entity groups. Use the large language model to analyze each entity group to determine whether it is a similar entity and whether entity alignment is required. The corresponding prompt word template is refined as prompt = {Role, Task, Restriction, Examples, Text}, where Role provides role positioning; the Task part provides task requirements; the Restriction part specifies the output format; and Text contains the entity group to be processed and its context information.
10. The method for constructing a knowledge graph of power equipment faults according to claim 9, characterized in that: A system for constructing a knowledge graph of power equipment faults includes: Data preprocessing model: including obtaining text information and image information of power equipment fault maintenance, and preprocessing the obtained information; the preprocessed image information is stored in the cloud database for easy access, the preprocessed text information is semantically segmented using the large language model to form independent text blocks, and the text information that meets the preset quality conditions is screened out through subject classification; ontology construction module: using the large language model to analyze and integrate the power equipment fault maintenance text information in the preprocessed data to obtain the core entity type, and combining the domain expert knowledge to complete the construction of the domain ontology, that is, the knowledge graph model layer; named entity recognition module: constructing the prompt word template and using the large language model to identify 10% of the 10% in step 1 The text information is pre-labeled, and the pre-labeling results are corrected using rule matching and power experts to obtain a training data set for the subsequent named entity recognition small model; based on the characteristics of fault field data, the traditional deep learning small model for named entity recognition is optimized, and the model is trained using the training data set. After training, the remaining 90% of the text information in step one is recognized; a hierarchical verification strategy is designed to stratify the large language model based on its cost. A low-cost large language model is selected to apply the pre-labeled prompt word template to recognize the remaining 90% of the text information in step one. The recognition results are compared with the recognition results of the traditional deep learning small model to eliminate erroneous entities; Relationship extraction module: Based on the characteristics of power equipment fault data and domain ontology, different entity relationship type combinations are designed. Different prompt word templates are constructed for different entity relationship type combinations. For the relationship types in the entity relationship type combinations, the entity types connected to them are found from the ontology constructed by the ontology construction module. The corresponding entities are filtered from the entity set. For different entity relationship type combinations, the respective prompt word templates and the filtered entities are used to extract the relationship between the entities using a large language model, obtaining structured triple data containing entity-relationship-entity. Knowledge Fusion Module: Entities in the training dataset from step 3 are selected. A prompt word model is designed to guide the large language model to generate synonymous expressions for entities within the domain. Power experts then modify the results to obtain similar entity pairs, which are then shuffled to obtain dissimilar entity pairs. The SBERT model is then fine-tuned using these similar and dissimilar entity pairs. For entity sets, the fine-tuned SBERT model and a depth-first search algorithm are used to obtain entity groups with similar semantics. The large language model then combines contextual information to achieve entity alignment. Graph storage module: After the data preprocessing model, ontology construction module, named entity recognition module, relationship extraction module and knowledge fusion module, the extracted entities and relationships are stored in the graph database Neo4j, and several initial entity nodes and several initial relationship edges are preliminarily generated; then, the fault repair case content and the text block content cut from the fault repair case are created as case nodes and text block nodes respectively and stored in Neo4j, and at the same time, connections are established between the case nodes and the initial entity nodes they contain, as well as between the text block nodes and the initial entity nodes they contain. Finally, several nodes and several edges are generated, and the image URL links generated by the cloud database are linked to the corresponding nodes in the form of attributes.
Citation Information
Cited By
Entity pair guided scientific and technical literature document level relation extraction method and system
CN121501984A
Mine pressure prediction method based on mixing of large and small models
CN121542593A
Data annotation method and system for professional knowledge base of large port facility management and maintenance model
CN121561464A
Vertical field data construction method based on large model
CN121581251A
Ship key equipment fault knowledge graph construction method and system, medium and terminal
CN122311393A