Electric power industry-oriented AI intelligent question answering and professional report generation method

By constructing a multi-layered knowledge graph and deep learning model for the power industry, the problem of dynamic changes in knowledge units within the power industry has been solved, enabling intelligent question answering and professional report generation, and improving information retrieval and decision support capabilities.

CN120875016APending Publication Date: 2025-10-31CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510758303.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the power industry, the relationships between knowledge units are intricate and dynamic. Traditional tree structures are difficult to effectively cope with the needs of multi-granularity knowledge representation and tracing, which affects the performance of AI intelligent question answering systems.

Method used

A multi-level knowledge graph for the power industry is constructed, using high-level nodes to represent concept-level knowledge units and low-level nodes to represent data-level knowledge units. Natural language processing is then combined with a deep learning model to achieve adaptive knowledge tracing and intelligent question answering.

Benefits of technology

It enables intelligent management and application of knowledge in the power sector, improves information retrieval and decision support capabilities, generates professional reports, and builds an intelligent question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875016A_ABST
    Figure CN120875016A_ABST
Patent Text Reader

Abstract

The invention provides an electric power industry-oriented A I intelligent question answering and professional report generation method, which comprises the steps of extracting key information from a constructed knowledge graph to form a preliminary electric power knowledge base structure, and the extraction process comprises the steps of identifying high-frequency nodes, analyzing connectivity among the nodes and evaluating importance weights of the nodes; based on the preliminary knowledge base structure, integrating structured and unstructured data in the power field to form a complete power knowledge base, and keeping synchronous updating with the knowledge graph; based on the electric power knowledge base, training a natural language processing model in the electric power field, which comprises named entity recognition and relation extraction by adopting a deep learning model in sequence; and performing semantic annotation on nodes and edges in the knowledge graph by utilizing a trained natural language processing model, and integrating new semantic information into the knowledge graph through entity alignment and relation mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to an AI-powered intelligent question answering and professional report generation method for the power industry. Background Technology

[0002] In the power industry, the relationships between knowledge units are intricate and dynamically change over time, making traditional tree-like structures ineffective in meeting the needs of multi-granularity knowledge representation and tracing. This challenge is particularly pronounced when building AI-powered question-answering systems, as these systems not only need to understand and process these complex multi-directional and circular relationships but also need to capture and adapt to changes in knowledge over time. Therefore, to improve the efficiency of AI-powered question-answering systems in the power industry, it is urgent to explore and develop new knowledge representation methods to accurately capture and manage this dynamically changing multi-dimensional information, thereby better serving knowledge retrieval and tracing in practical applications. Summary of the Invention

[0003] This invention provides a method for AI-powered intelligent question answering and professional report generation in the power industry, mainly including:

[0004] A knowledge graph for the power industry is constructed. The knowledge graph contains nodes representing knowledge units and edges representing the relationships between knowledge units. A multi-level node representation method is adopted in the knowledge graph. High-level nodes represent concept-level knowledge units, and low-level nodes represent data-level knowledge units. Power industry knowledge concepts are abstracted into high-level nodes, and various types of power data are organized into low-level nodes. A mapping relationship between the high-level nodes and the low-level nodes is established.

[0005] Key information is extracted from the constructed knowledge graph to form a preliminary power knowledge base structure. The extraction process includes identifying high-frequency nodes, analyzing the connectivity between nodes, and evaluating the importance weight of nodes.

[0006] Based on the preliminary knowledge base structure, structured and unstructured data in the power sector are integrated to form a complete power knowledge base, which is kept updated synchronously with the knowledge graph.

[0007] Based on the power knowledge base, a natural language processing model for the power field is trained, including sequentially using a deep learning model for named entity recognition and relation extraction.

[0008] Using a trained natural language processing model, semantic annotation is performed on the nodes and edges in the knowledge graph. Through entity alignment and relation mapping, new semantic information is integrated into the knowledge graph.

[0009] Adaptive knowledge tracing is performed based on the knowledge graph. Starting from a given target knowledge unit, the knowledge graph is traversed using a breadth-first search strategy. The depth and breadth parameters of traversing target knowledge units of different granularities are dynamically adjusted to obtain other knowledge units that are directly or indirectly related to the target knowledge unit. The adjustment is based on the hierarchy, relevance, and time attributes of the knowledge units.

[0010] In the process of knowledge tracing, when the target knowledge unit is a concept-level node, higher-level nodes are traversed first; when the target knowledge unit is a data-level node, the search is focused on the lower-level nodes. At the same time, the relationship between knowledge units at different time points is dynamically weighted.

[0011] Based on the knowledge tracing results, the tracing data is analyzed and organized, and the type distribution, correlation strength, and time distribution characteristics of the tracing knowledge units are statistically analyzed. The data is then organized according to a predefined power professional report template, the report chapters are filled in, and a complete power professional report is generated.

[0012] By leveraging the power industry knowledge graph and power knowledge base, an intelligent question-answering system for the power industry is constructed. Through natural language understanding methods, user questions are mapped to nodes and relationships in the knowledge graph. Combined with semantic search and reasoning methods, the optimal answer to the question is obtained, forming an intelligent human-computer interaction.

[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0014] This invention discloses an AI-powered intelligent question-answering and professional report generation method for the power industry. The method constructs a multi-layered power industry knowledge graph, including concept-level and data-level nodes, with edges possessing weight and temporal attributes. A preliminary knowledge base structure is formed by extracting key information, and a complete power knowledge base is constructed by integrating industry data. A deep learning model is used to train natural language processing capabilities for semantic annotation of the knowledge graph. Based on this, adaptive knowledge tracing is implemented, dynamically adjusting search parameters to acquire relevant knowledge units. Professional reports are generated based on the tracing results, and an intelligent question-answering system is constructed. This invention, through knowledge graph and natural language processing technologies, achieves intelligent management and application of knowledge in the power field, improving industry information retrieval and decision support capabilities. Attached Figure Description

[0015] Figure 1 This is a flowchart of an AI-powered intelligent question answering and professional report generation method for the power industry, based on the present invention.

[0016] Figure 2 This is a schematic diagram of an AI-powered intelligent question answering and professional report generation method for the power industry according to the present invention.

[0017] Figure 3This is another schematic diagram of an AI-powered intelligent question answering and professional report generation method for the power industry according to the present invention. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0019] like Figure 1-3 This embodiment of a method for generating AI-powered intelligent question answering and professional reports for the power industry may specifically include:

[0020] Step S101: Construct a knowledge graph for the power industry. The knowledge graph includes nodes representing knowledge units and edges representing the relationships between knowledge units. A multi-level node representation method is adopted in the knowledge graph. High-level nodes represent concept-level knowledge units, and low-level nodes represent data-level knowledge units. Power industry knowledge concepts are abstracted into high-level nodes, and various types of power data are organized into low-level nodes. A mapping relationship is established between the high-level nodes and the low-level nodes.

[0021] The process involves acquiring power industry concepts and data, abstracting the power industry concepts into high-level nodes, and organizing the power data into low-level nodes. Semantic analysis is performed on the high-level and low-level nodes to determine the mapping relationships between them. The strength of associations between nodes is calculated to obtain the weight attributes of the association edges. Changes to the association edges are recorded using timestamps and version numbers to obtain the time attributes of the association edges. New data streams are acquired, and new knowledge units are identified from these streams. The semantic similarity between the new knowledge units and existing nodes is calculated; if the semantic similarity is higher than a preset threshold, an association is established. Reasoning is then performed on these associations to obtain the association analysis results for explicit knowledge.

[0022] Specifically, a multi-level knowledge unit node structure is constructed, abstracting power industry concepts into high-level nodes and power data organization into low-level nodes. Semantic analysis is used to determine the mapping relationships between nodes, and an ontology model is used to define node attributes and relationship types. Natural language processing (NLP) techniques are used to extract entities and relationships from unstructured text, forming an initial knowledge graph structure. Weight and time attributes are designed for the edges of relationships. The PageRank algorithm is used to calculate the association strength between nodes for the weight attributes, and the time attributes combine timestamps and version numbers to record relationship changes. The timestamps are accurate to milliseconds, and the version numbers use incrementing integers. A time-series database is used to store the historical evolution information of edges, and a graph database is used to achieve efficient graph structure storage and querying. A B-tree index is used to optimize the query efficiency of the time attribute, enabling time backtracking and allowing queries of the knowledge graph state at any point in time. To address the dynamic nature of the knowledge graph, incremental updates are performed. New data streams are obtained from the real-time data interface of the power system, and NLP techniques are used to identify new knowledge units. The TransE algorithm is used to calculate the semantic similarity between new and old nodes; if the similarity is higher than 0.8, an association is established. The Git version control system is used to manage multiple historical versions of the knowledge graph, enabling rapid switching and comparison between versions. Based on the constructed power industry knowledge graph, multi-dimensional knowledge reasoning capabilities are employed, combined with rule-based and statistical reasoning methods, to perform correlation analysis and expansion on explicit knowledge within the graph. Graph convolutional networks are used to capture high-order relationships between nodes, deriving high-level conceptual knowledge from lower-level data nodes. The reasoning results are combined with incremental update methods to continuously optimize the semantic understanding capabilities of the knowledge graph, forming a dynamic update and reasoning closed loop. When constructing the power industry knowledge graph, an ontology model is first used to define node attributes and relationship types. For example, high-level conceptual nodes such as "power plant," "transmission line," and "substation" are defined, along with lower-level data nodes such as "power generation," "line load," and "transformer temperature." Through semantic analysis, mapping relationships between nodes are determined, such as establishing a "production" relationship between "power plant" and "power generation." Natural language processing techniques are used to extract entities and relationships from power industry documents to construct the initial knowledge graph. When designing relational edges, the PageRank algorithm is used to calculate the correlation strength between nodes. For time attributes, a composite structure is used for storage, with timestamps accurate to milliseconds and version numbers using incrementing integers. For example, a load change record for a transmission line is "2024-10-12 14:30:25.789, version number 25". A time-series database is used to store the historical evolution information of edges, and efficient querying is achieved through the graph database Neo4j. A B+ tree index is used to optimize time attribute queries, enabling millisecond-level time backtracking. An incremental update mechanism is developed to address the dynamic nature of the knowledge graph. New data streams are obtained every 5 minutes from the real-time data interface of the power SCADA system, and named entity recognition technology is used to identify newly added knowledge units.The TransE algorithm is used to calculate the semantic similarity between new and old nodes, with a threshold of 0.8. If the similarity is higher than the threshold, an association is established. A Git version control system is used to manage the knowledge graph versions, generating a new version daily and retaining historical versions from the last 30 days. Based on the constructed knowledge graph, a multi-dimensional knowledge reasoning function is designed. Rule-based reasoning and statistical reasoning methods are combined to perform association analysis on explicit knowledge in the graph. For example, reasoning is performed based on the rule "If the input power of substation A is greater than its output power, then substation A may have equipment failure." A graph convolutional network is used to capture high-order relationships between nodes, deriving the high-level concept of "transformer failure" from low-level data nodes such as "abnormal transformer temperature" and "abnormal oil chromatography data." The reasoning results are combined with an incremental update method, updating the knowledge graph hourly to continuously optimize semantic understanding capabilities, forming a dynamic update and reasoning closed loop.

[0023] Step S102: Extract key information from the constructed knowledge graph to form a preliminary power knowledge base structure. The extraction process includes identifying high-frequency nodes, analyzing the connectivity between nodes, and evaluating the importance weight of nodes.

[0024] A depth-first search algorithm is used to traverse the knowledge graph to obtain the frequency of each node. A hash table is used to record the mapping relationship between node ID and its frequency. Based on the mapping relationship, the connectivity of each node in the knowledge graph is calculated, and a weighted adjacency list structure is used to represent the connections between nodes. After the weighted adjacency list structure is represented, multi-dimensional feature vectors of the nodes are obtained based on their frequency, connectivity, node type, and number of attributes. The importance weight of the nodes is determined using these multi-dimensional feature vectors. Based on the importance weights of the nodes, key nodes and their related edges are selected to construct a subgraph structure. After the subgraph structure is constructed, the Louvain algorithm is used to perform community detection on the subgraph to determine the preliminary hierarchical structure of the power knowledge base.

[0025] Specifically, a depth-first search algorithm is used to traverse the knowledge graph, counting the frequency of each node. A hash table is used to record the mapping relationship between node IDs and their occurrence counts, resulting in a set of high-frequency nodes. During the traversal, a predefined power industry ontology dictionary is used to assign extra weights to matching nodes such as "substation," "transmission line," and "load," increasing their frequency counts. The connectivity of each node in the knowledge graph is calculated, and a weighted adjacency list structure is used to represent the connections between nodes. The edge weights are determined according to a predefined relation importance table; for example, the "power supply" relation has a higher weight than the "geographical location" relation. By traversing the adjacency list, the weighted sum of the in-degree and out-degree of each node is calculated to determine the node's connectivity. A multi-dimensional feature vector is constructed based on features such as node frequency, connectivity, node type, and number of attributes. Power industry-specific indicators are incorporated, such as the equipment capacity represented by the node and the number of users covered. The PageRank algorithm is used to calculate node importance, taking into account the power network topology. Based on the calculation results, the nodes are ranked by importance to obtain their weight values. Based on node importance weights, key nodes and their related edges are selected to construct a subgraph structure. The Louvain algorithm is used to perform community discovery on the subgraph, forming a preliminary hierarchical structure for the power knowledge base. According to the physical structure of the power system, such as voltage levels and regional divisions, the communities are adjusted and named, ultimately forming a hierarchical power knowledge base structure. In constructing the power industry knowledge graph, a depth-first search algorithm is first used to traverse the entire graph structure. During the traversal, a hash table is used to record the mapping relationship between node IDs and their occurrence frequency, while a predefined power industry ontology dictionary is used to weight specific nodes. For example, the weight of the "substation" node is set to 1.5, the weight of the "transmission line" node is set to 1.3, and the weight of ordinary nodes is 1. After the traversal, in the high-frequency node set, the "220kV substation" node appears most frequently, 500 times. Subsequently, a weighted adjacency list structure is used to represent the connections between nodes, and the edge weights are determined according to a predefined relationship importance table. For example, the weight of the "power supply" relationship is 2, and the weight of the "geographical location" relationship is 1. By traversing the adjacency list and calculating the weighted connectivity of each node, the "main transformer" node was found to have the highest weighted connectivity of 250. A multi-dimensional feature vector was constructed based on node frequency, connectivity, and other characteristics, incorporating power industry-specific indicators such as equipment capacity and the number of users covered. The PageRank algorithm was used to calculate node importance, converging after 10 iterations, revealing the "500kV substation" node as having the highest importance, with a PageRank value of 0.085. Based on the calculation results, the top 100 important nodes and their associated edges were selected to construct a subgraph structure. The Louvain algorithm was used for community detection in the subgraph, with a resolution parameter of 1.0. After 3 iterations, 10 communities were identified.Finally, the community was adjusted based on the physical structure of the power system, such as grouping nodes with similar voltage levels together to form a hierarchical knowledge base structure for "power generation," "transmission," and "distribution." From the initial knowledge graph containing 100,000 nodes and 500,000 edges, a core structure of the power knowledge base containing 1,000 key nodes and 5,000 important relationships was extracted.

[0026] Step S103: Based on the preliminary knowledge base structure, integrate structured and unstructured data in the power field to form a complete power knowledge base, and keep it updated synchronously with the knowledge graph.

[0027] Structured data is obtained from power equipment databases, SCADA systems, and energy management systems. This structured data undergoes unit unification and naming standardization to obtain standardized power equipment data. The BERT model is used to process power industry documents, operation manuals, and fault reports to identify power equipment names, parameters, and operating actions. A relationship extraction model is constructed based on the identification results to extract topological relationships and operational flow information between devices. A predefined power domain ontology model is obtained. Based on the power domain ontology model, the standardized power equipment data and extracted device relationship information are semantically aligned. Power domain word vectors are trained using the Word2Vec algorithm. The similarity between concepts is calculated based on the word vectors. If the similarity exceeds a preset threshold, the corresponding concepts are merged. A bidirectional synchronous update mechanism is established between the knowledge graph and the knowledge base. If a knowledge graph update message is detected, an incremental update procedure for the knowledge base is triggered. Based on the real-time data stream processing results, corresponding knowledge base update operations are executed.

[0028] Specifically, Apache NiFi is used as the ETL tool to extract structured data from power equipment databases, SCADA systems, and energy management systems. Unified units and naming conventions are established for equipment data, outlier detection and completion are performed on operational data, and time-series alignment is applied to energy management data. Processed data is converted into knowledge base entities and relations using predefined mapping rules and stored in a graph database as the foundational structure of the knowledge base. The BERT model is used to perform named entity recognition from unstructured data such as power industry documents, operation manuals, and fault reports, identifying power equipment names, parameters, and operational actions. A remote supervision method is used to construct a relation extraction model, extracting topological relationships and operational procedures between devices, and converting the extracted information into knowledge triples. Based on a predefined power domain ontology model, the obtained knowledge units are semantically aligned and integrated. Word2Vec is used to train power domain word vectors, and ontology matching is performed based on word vector similarity. A similarity threshold of 0.8 is set, and concepts exceeding the threshold are merged. A rule base is used to resolve concept conflicts, such as naming rules for equipment at different voltage levels. Concept redundancy and contradictions are eliminated, forming a unified power knowledge base structure. A bidirectional synchronous update mechanism was established between the knowledge graph and the knowledge base. Apache Kafka was used as the message queue, and knowledge graph update topics were set. When an update message was detected, the incremental update procedure of the knowledge base was triggered. The update cycle was set to 5 minutes. Through real-time data stream processing, when the knowledge graph changed, the corresponding knowledge base update operation was executed to maintain consistency between the two. In the process of building the power knowledge base, Apache NiFi was first used to design a data processing flow to extract structured data from the power equipment database, SCADA system, and energy management system. For equipment data, the voltage unit was unified to kV and the power unit was unified to MW; outlier detection was performed on the operating data, with a threshold set at the average value ± 3 times the standard deviation, and data exceeding the range was interpolated; time series alignment was performed on the energy management data, with a unified sampling interval of 5 minutes. The processed data was converted into nodes and relationships in the Neo4j graph database through predefined mapping rules. Then, the BERT model was used to perform named entity recognition on power industry documents, achieving a recognition accuracy of 95%. A relation extraction model was constructed using the distant super vision method, extracting 500,000 topological relationships and operational process information between devices from 100,000 documents. Then, a 100-dimensional word vector model for the power industry was trained using the Word2Vec algorithm, and ontology matching was performed based on cosine similarity, with a similarity threshold set at 0.8. Concepts exceeding the threshold were merged, such as merging "transformer" and "main transformer" into a unified concept. A rule base containing 500 rules was used to resolve concept conflicts, such as standardizing the naming of substations of different voltage levels into the format "voltage value + kV + substation".Finally, an Apache Kafka knowledge graph update topic was set up. When an update message is detected, an incremental update procedure for the knowledge base is triggered. The update cycle is set to 5 minutes, and Kafka's message queue mechanism enables near real-time synchronization between the knowledge graph and the knowledge base, ensuring data consistency. This ultimately resulted in a unified power knowledge base containing 1 million entities and 5 million relationships, covering core knowledge in three major areas: power equipment, grid operation, and energy management.

[0029] Step S104: Based on the power knowledge base, train a natural language processing model for the power field, including sequentially using a deep learning model for named entity recognition and relation extraction.

[0030] The received power industry standard documents, equipment manuals, and operation reports are preprocessed to obtain a power domain dictionary. Based on the power domain dictionary and the preprocessed data, a named entity recognition (NAME) task is trained on the preprocessed data to obtain a NAME recognition model. Using the entity information identified by the NAME recognition model and known relationships in the knowledge base, unannotated text is annotated to obtain annotated text. For the annotated text, a recurrent neural network is used to construct a relation extraction model, which determines the relationships between entity pairs by inputting contextual representations of the entity pairs.

[0031] Specifically, the NLTK library in Python was used to perform word segmentation, part-of-speech tagging, and syntactic analysis on standard documents, equipment manuals, and operation reports in the power industry. When constructing the power domain dictionary, 10,000 domain-specific terms were selected by combining word frequency statistics and expert knowledge. A sequence labeling method based on conditional random fields, combined with rule matching, was used to perform preliminary entity labeling on the text. Through these processes, a preliminary labeled corpus and power domain dictionary were formed. Based on the preprocessed data, a pre-trained BERT-based model was loaded using the Transformers library of the HuggingFace open-source platform for training the named entity recognition task. The learning rate was set to 2e-5, the batch size to 32, and the number of training epochs to 3. The AdamW optimizer was used, with weight decay applied. 10-fold cross-validation was used to evaluate model performance, and entity types such as power equipment, parameters, and operational actions were identified by fine-tuning the pre-trained model. Unlabeled text was automatically labeled using the identified entity information and known relations in the knowledge base. The sliding window size was set to 3 sentences; if the window contained entity pairs from the knowledge base, they were labeled as the corresponding relations. Based on dependency parsing results, feature vectors for entity pairs are constructed. A recurrent neural network is used to build a relation extraction model, with the context representation of entity pairs as input. The extracted relations are evaluated, and precision, recall, and F1 score are calculated. Named entity recognition and relation extraction tasks are jointly trained, and a shared BERT encoding layer is designed, with task-specific layers for named entity recognition and relation extraction connected to the upper layers respectively. Task-specific loss functions are used, and the losses of multiple tasks are combined through a weighted sum. An alternating training strategy is adopted, randomly selecting one task for optimization in each batch to improve the model's generalization ability and robustness. In the training process of the natural language processing model in the power industry, 1 million words of power industry documents are preprocessed using Python's NLTK library. The word segmentation accuracy reaches 98%, and the part-of-speech tagging accuracy is 96%. Through word frequency statistics and knowledge screening, a power industry dictionary containing 12,500 proprietary words is constructed. Conditional random fields are used for sequence labeling, combined with 500 manually defined rules, to perform preliminary entity labeling on the text, achieving an F1 score of 85%. Subsequently, the BERT-base model was loaded using HuggingFace's Transformers library and fine-tuned on 500,000 preprocessed sentences. With a learning rate of 2e-5 and a batch size of 32, after three training epochs, the F1 score for named entity recognition improved to 92%. In the relation extraction task, 100,000 known relation pairs from the knowledge base were used to automatically annotate 2 million unannotated text sentences. A sliding window of three sentences was set, successfully annotating 150,000 entity relation pairs. A relation extraction model was built using a bidirectional LSTM network, with 300-dimensional entity pair context representations as input, achieving an F1 score of 88% on the validation set.Finally, a multi-task learning framework was designed, sharing a 12-layer BERT encoding layer, with two task-specific layers for named entity recognition and relation extraction connected to the upper layers respectively. The loss function weights were set to 1:1, and an alternating training strategy was adopted, switching tasks every 100 batches. After 50,000 training steps, the joint model achieved a comprehensive F1 score of 94% on the test set, a 2 percentage point improvement over individual training. This resulted in a natural language processing model for the power sector capable of accurately identifying power equipment, parameters, and operational actions, and extracting the complex relationships between them.

[0032] Step S105: Using the trained natural language processing model, semantic annotation is performed on the nodes and edges in the knowledge graph. Through entity alignment and relation mapping, new semantic information is integrated into the knowledge graph.

[0033] A trained natural language processing model is used to acquire node information from a knowledge graph. Semantic annotation is performed on the node information to obtain the semantic type and attributes of the nodes. Based on the semantic type and attributes of the nodes, the Word2Vec algorithm is used to generate semantic vector representations of the nodes. Based on the semantic vector representations of the nodes, entity alignment is performed using a cosine similarity calculation method, which includes setting a similarity threshold. If the cosine similarity is higher than the preset threshold, entity attributes are compared. If the overlap of the entity attributes is lower than a preset value, a hierarchical clustering algorithm is used to cluster the entity attributes and determine the splitting boundary. Relationships are extracted from the edge information in the knowledge graph, and a convolutional neural network is used to semantically annotate the edge information. Based on the semantic annotation results, the semantic type of the edges is mapped to a predefined relation ontology. The mapped semantic information is integrated with the original knowledge graph to obtain an integrated knowledge graph. Using the batch update function of the graph database, the attributes of the nodes and edges are updated according to the integrated knowledge graph to obtain an updated knowledge graph.

[0034] Specifically, a trained natural language processing model is used to semantically annotate nodes in the knowledge graph. Entity recognition is performed using conditional random fields, achieving an accuracy of 95%. Support vector machines are used for entity classification, categorizing entities into devices, parameters, operations, etc., with a classification accuracy of 92%. The Word2Vec algorithm is used to generate 300-dimensional semantic vector representations of nodes, assigning semantic type and attributes to each node. Based on the semantic vectors of nodes, cosine similarity is used for entity alignment, with a similarity threshold of 0.85. For entity pairs with similarity higher than the threshold, attribute comparison is performed. If the attribute overlap is less than 60%, a hierarchical clustering algorithm is used to cluster entity attributes, determining splitting boundaries and splitting an entity into multiple concepts, thus addressing entity redundancy and splitting issues in the knowledge graph. Relationships are extracted and semantically annotated from edges in the knowledge graph, and relation classification is performed using an attention-based convolutional neural network, achieving an accuracy of 90%. A relation ontology containing 100 relation types is constructed, and WordNet is used for semantic expansion. A graph matching algorithm is used for relation mapping, mapping the semantic types of edges to predefined relation ontologs to unify the representation of relations, achieving a matching accuracy of 88%. New semantic information is integrated with the existing knowledge graph, and a conflict resolution strategy based on timestamps and credibility is designed. For different attribute values ​​of the same entity, the information with the latest timestamp and highest credibility is retained. The batch update function of the graph database Neo4j is used, updating 10,000 nodes at a time. Through a transaction processing mechanism, the attributes of nodes and edges are updated in batches, while maintaining the consistency and integrity of the graph. In the semantic annotation and integration process of the power knowledge graph, a graph containing 1 million nodes is first processed. A conditional random field model is used for entity recognition, achieving an F1 score of 95.5% on 100,000 test data points. Subsequently, a support vector machine is used for entity classification, dividing entities into 15 categories, including transformers, switches, and lines, with a classification accuracy of 92.3%. A 300-dimensional node semantic vector is trained using the Word2Vec algorithm, with a vocabulary size of 500,000 words and a window size of 5. For entity alignment, cosine similarity between nodes was calculated, with a threshold of 0.85, resulting in 250,000 pairs of similar entities. Attribute comparison identified 30,000 pairs of entities requiring splitting. Using the Ward hierarchical clustering algorithm with an inter-cluster distance threshold of 0.6, these entities were split into 75,000 refined concepts. For relation processing, a convolutional neural network with an attention mechanism was used for classification, achieving 90.2% accuracy across 50 relation types. The constructed relation ontology contains 100 core relations and 500 extended relations, with semantic expansion using WordNet, improving coverage by 15%. The Hungarian algorithm was employed for graph matching, achieving a relation mapping accuracy of 88.5%.During the knowledge integration phase, a conflict resolution strategy was designed, assigning a credibility score of 0-1 to each attribute. Combining this with the modification timestamp, the attribute with the highest and most recent score was selected. Using Neo4j's batch update function, 10,000 nodes were processed per batch, with a transaction timeout of 30 seconds, achieving a success rate of 99.9%. The resulting semantically rich and structurally consistent power knowledge graph contains 1.1 million nodes and 3 million relationships, covering knowledge from multiple fields such as power equipment, operational status, and energy management.

[0035] Step S106: Based on the knowledge graph, perform adaptive knowledge tracing. Starting from a given target knowledge unit, traverse the knowledge graph using a breadth-first search strategy, dynamically adjust the depth and breadth parameters of traversing target knowledge units of different granularities, and obtain other knowledge units that are directly or indirectly related to the target knowledge unit. The adjustment criteria include the hierarchy, relevance, and time attributes of the knowledge units.

[0036] The system receives a search request carrying a target knowledge unit identifier, initializes a breadth-first search queue based on the search request, and sets initial search depth and breadth parameters; calculates search weights, which are determined by the node level, with the target knowledge unit's level having a weight of 1; assigns search priorities to nodes at different levels based on the search weights; obtains the semantic vectors of the target knowledge unit and the currently traversed nodes, and calculates the similarity score of the semantic vectors; learns the node's embedding representation, which is used to calculate the correlation score between nodes; combines the similarity score, correlation score, and edge weight attributes to obtain a comprehensive correlation score; receives the node's time attribute information; calculates a time decay coefficient based on the time attribute information; multiplies the time decay coefficient by the node's importance to obtain a node priority that incorporates the time factor; if the number of traversed nodes reaches a preset threshold, calculates the average correlation and standard deviation of the traversed nodes; updates the preset threshold based on the average correlation and standard deviation; and dynamically adjusts the search depth and breadth parameters using the updated threshold.

[0037] Specifically, starting from a given target knowledge unit, a breadth-first search queue is initialized, setting initial search depth and breadth parameters. Search weights are assigned using an exponential decay function w = 0.8^(d-1), where d is the node level. For the target knowledge unit's level, d = 1, with a weight of 1. The weight decreases with each upward or downward level, assigning different search weights to nodes at different levels based on the hierarchical relationship of the knowledge units. During the traversal of the knowledge graph, the correlation between the current node and the target knowledge unit is calculated. A cosine similarity algorithm is used to calculate the semantic vectors between nodes, while a node-2vec algorithm is used to learn the node's embedding representation, setting the window size to 5 and the embedding dimension to 128. Combining the similarity obtained from node-2vec with the edge weight attributes, a comprehensive correlation score is obtained. Based on the node's time attribute, an exponential decay function f(t) = e^(-0.1t) is used to calculate the time decay coefficient, where t is the time difference (calculated in days), and e is the base of the natural logarithm. Multiplying the time decay coefficient by the node importance yields a node priority considering the time factor. A decay coefficient is assigned to knowledge units from historical versions, incorporating the time factor into node importance evaluation and dynamically adjusting node priority during the search process. The depth and breadth parameters of the search are dynamically adjusted based on the quantity and quality of traversed nodes. An initial threshold θ = 0.5 is set. Every 100 nodes traversed, the average correlation μ and standard deviation σ of these nodes are calculated. The threshold θnew = μ + 1.5 * σ is updated. The search depth and breadth are dynamically adjusted based on the new threshold. By setting an adaptive threshold, the traversal range of knowledge units of different granularities is controlled, optimizing the tracing path. In the adaptive knowledge tracing process of the power knowledge graph, starting from the target knowledge unit "500kV transformer," a breadth-first search queue is initialized, with an initial search depth of 5 and a breadth parameter of 10. The search weights are allocated using an exponential decay function w = 0.8^(d-1), with the target unit having a weight of 1, adjacent levels having a weight of 0.8, and the third level having a weight of 0.64. During the traversal, the cosine similarity algorithm was used to calculate the semantic similarity between nodes, while the node 2vec algorithm was used to learn the 128-dimensional node embedding, with a window size of 5. For the "transformer winding" node, the calculated cosine similarity was 0.85, the node 2vec similarity was 0.78, the edge weight was 0.9, and the overall relevance score was 0.84. In the time attribute processing, for historical version information from 30 days ago, a decay coefficient of f(t) = e^(-0.130) ≈ 0.05 was used, which was multiplied by the node importance of 0.9 to obtain a priority of 0.045. After traversing 100 nodes, the average relevance μ = 0.6, the standard deviation σ = 0.15, and the update threshold θnew = 0.6 + 1.50.15 = 0.825. Based on the new threshold, the search depth was adjusted to 4, the breadth parameter was increased to 15, and the traversal range of different granularity knowledge units such as "power equipment", "operating status", and "fault diagnosis" was optimized.Ultimately, 2,000 nodes closely related to the target knowledge unit were selected from 1 million nodes, forming a comprehensive and accurate knowledge tracing network.

[0038] In step S107, during the knowledge tracing process, when the target knowledge unit is a concept-level node, higher-level nodes are traversed first; when the target knowledge unit is a data-level node, the search is focused on the lower-level nodes, and the dynamic weight adjustment is performed on the knowledge unit relationships at different time nodes.

[0039] The target knowledge unit is acquired, and a support vector machine classifier is used to identify its type, classifying it into concept-level nodes or data-level nodes. Based on the type of the target knowledge unit, the search parameters are dynamically adjusted: if the target knowledge unit is a concept-level node, the search breadth parameter is increased from the default value to a preset value, and the depth limit is decreased from the preset value to a specified value; if the target knowledge unit is a data-level node, the search depth limit is increased from the preset value to a specified value, and the breadth parameter is decreased from the default value to a preset value. The timestamp information of the target knowledge unit is parsed, and its time weight is calculated. Combining the type, level information, and time weight of the target knowledge unit, the relationship strength between the target knowledge unit and other knowledge units is calculated. Based on the relationship strength, the traversal queue is prioritized, and the traversal order and depth are dynamically adjusted to optimize the knowledge tracing path.

[0040] Specifically, the target knowledge units are type-identified using a support vector machine classifier, classifying them into concept-level nodes or data-level nodes. The feature vectors include 10 dimensions, such as out-degree, in-degree, and hierarchy depth. The training dataset contains 1000 labeled samples, with 60% being concept-level nodes and 40% being data-level nodes, determining the subsequent traversal strategy. The search algorithm parameters are dynamically adjusted based on the node type. When the target is a concept-level node, the search breadth parameter is increased from the default value of 10 to 20, and the depth limit is reduced from 5 to 3, increasing the traversal weight of higher-level nodes and expanding the search breadth. When the target is a data-level node, the search depth limit is increased from 5 to 8, and the breadth parameter is reduced from 10 to 5, increasing the traversal depth of lower-level nodes and narrowing the search range. During the traversal, the timestamp information of each knowledge unit is parsed, and the time weight is calculated using the exponential decay function w(t) = e^(-0.1t), where t is the age of the knowledge unit (calculated in years), and e is the base of the natural logarithm. For knowledge units in the current year, the weight is 1, decreasing by approximately 9.5% each year thereafter, with newer knowledge units assigned higher weights. The traversal order is initially adjusted based on the time weight, prioritizing nodes with higher weights. Combining node type, hierarchical information, and time weight, the relationship strength between knowledge units is calculated as S = 0.4*T + 0.3*L + 0.3*W, where T is the time weight, L is the hierarchical similarity, and W is the edge weight between nodes. Based on the calculated relationship strength, the traversal queue is prioritized, and the traversal order and depth are dynamically adjusted to optimize the knowledge tracing path. In the adaptive knowledge tracing process of the power knowledge graph, the target knowledge unit "transformer fault diagnosis" is first identified. A support vector machine classifier is used, with feature vectors containing 10 dimensions, including node out-degree (15), in-degree (8), and hierarchical depth (3). In 1000 labeled samples, this unit is identified as a concept-level node with a confidence level of 0.92. Based on this classification result, the search algorithm parameters are dynamically adjusted, with the breadth parameter increasing from 10 to 20 and the depth limit decreasing from 5 to 3. During the traversal, the timestamps of knowledge units are parsed. For example, the timestamp of "Transformer Insulation Aging Model" is 2022, and its time weight is calculated using the exponential decay function w(t) = e^(-0.11) ≈ 0.905. For "Oil Chromatography Analysis Method" in 2020, the time weight is e^(-0.13) ≈ 0.741. Based on the time weight, "Transformer Insulation Aging Model" has a higher priority in the traversal queue. Then, the relationship strength between knowledge units is calculated. For example, the relationship strength between "Transformer Fault Diagnosis" and "Transformer Insulation Aging Model" is S = 0.4 * 0.905 + 0.3 * 1 + 0.3 * 0.8 = 0.902, while the relationship strength with "Oil Chromatography Analysis Method" is 0.4 * 0.741 + 0.3 * 0.9 + 0.3 * 0.7 = 0.7764. Based on these calculation results, the traversal queue is reordered, and "Transformer Insulation Aging Model" is placed at the top.The algorithm prioritizes traversing 2,000 highly relevant concept-level nodes among 1 million nodes, forming a knowledge tracing network that is both extensive and focused, covering a multi-level knowledge structure from high-rise transformer theory to specific fault cases.

[0041] Step S108: Based on the knowledge tracing results, analyze and organize the tracing data, statistically analyze the type distribution, correlation strength, and time distribution characteristics of the tracing knowledge units, organize them according to the predefined power professional report template, fill in the report chapter content, and generate a complete power professional report.

[0042] A set of knowledge units is obtained. Based on a preset neighborhood radius and minimum sample size, the DBSCAN algorithm is used to perform clustering analysis on the knowledge unit set. Noise points that cannot be clustered are identified and classified into other types. The correlation strength between the knowledge units is calculated, and a relationship graph of the knowledge units is constructed, where the nodes of the relationship graph are the knowledge units and the edges are the correlations between the knowledge units. Based on the relationship graph, the PageRank algorithm is used to evaluate the importance of the knowledge units, where the damping factor of the PageRank algorithm is a preset value and the number of iterations is a preset number. Combining the weights of the correlation edges in the relationship graph, a comprehensive correlation strength score between the knowledge units is obtained. Time series analysis is performed on the knowledge units to identify... The update cycle and trend of the knowledge units are analyzed; a heatmap of the time distribution of the knowledge units is plotted, with the horizontal axis representing the year and the vertical axis representing the month, and the color intensity representing the number of knowledge units; the update trend of the knowledge units is predicted within a future preset time period, and a time distribution feature report of the knowledge units is generated; a predefined power professional report template is obtained, and the structural features and keywords of the report template are extracted; the knowledge units are preprocessed, including word segmentation and stop word removal; the text feature vector of the knowledge unit is calculated; based on the text feature vector, the cosine similarity between the knowledge unit and the chapter of the report template is calculated; if the cosine similarity is higher than a preset threshold, the knowledge unit is filled into the corresponding chapter of the report template to generate a complete power professional report.

[0043] Specifically, the knowledge units obtained through source tracing are classified and statistically analyzed. The DBSCAN algorithm is used for cluster analysis of knowledge units, with a neighborhood radius of 0.5 and a minimum sample size of 5. Noise points that cannot be clustered are classified separately as "Other" to determine the main knowledge types and their distribution. The correlation strength between knowledge units is calculated, and a knowledge unit relationship graph is constructed, where nodes represent knowledge units and edges represent the correlations between units. The PageRank algorithm is used to evaluate the importance of knowledge units, with a damping factor of 0.85 and 100 iterations. The comprehensive correlation strength score between knowledge units is obtained by combining the weights of the associated edges. The temporal distribution characteristics of knowledge units are analyzed, and time series analysis is used to identify the periodicity and trend of knowledge updates. A time distribution heatmap is drawn using data visualization tools, with the x-axis representing the year, the y-axis representing the month, and the color intensity representing the number of knowledge units. At the same time, time series prediction is performed to predict the knowledge update trend for the next 12 months, generating a time distribution characteristic report. Based on a predefined power industry report template, the structural features and keywords of the template are extracted. Natural language processing tools are used for text preprocessing, including word segmentation and stop word removal. TF-IDF was used to calculate text feature vectors, and then cosine similarity was used to calculate the matching degree between knowledge units and report chapters. A matching threshold of 0.6 was set, and content exceeding the threshold was automatically filled into the corresponding chapters to generate a complete power industry report. In the process of generating the power industry knowledge graph report, 100,000 source-acquired knowledge units were first classified, statistically analyzed, and clustered. Using the DBSCAN algorithm with a neighborhood radius of 0.5 and a minimum sample size of 5, 95% of the knowledge units were successfully divided into 8 main categories, including "equipment failure," "operation and maintenance," and "energy management," with the remaining 5% classified as "other." Subsequently, a knowledge unit relationship graph containing 100,000 nodes and 500,000 edges was constructed. The PageRank algorithm was used to evaluate unit importance, with a damping factor of 0.85, and convergence after 100 iterations. Combined with edge weights, the comprehensive association strength score of the "transformer fault diagnosis" unit was calculated to be 0.92, ranking in the top 5%. In the time distribution analysis, a heatmap was used to visualize the knowledge update situation over 5 years, revealing that March and September of each year are peak periods for knowledge updates. Through time series forecasting, knowledge related to "smart grids" is projected to grow by 20% over the next 12 months. Finally, based on a predefined power industry report template, 500 structural features and 1000 keywords were extracted. TF-IDF was used to calculate text feature vectors, and 100,000 knowledge units were matched with 50 chapters of the report, with a matching threshold set to 0.6. Ultimately, 90% of the chapter content was automatically populated, generating a complete 150,000-word power industry report covering multiple aspects from equipment operation and maintenance to energy policy, providing systematic knowledge support for power industry decision-making.

[0044] Step S109: Utilize the power industry knowledge graph and power knowledge base to construct an intelligent question-answering system for the power industry. Through natural language understanding methods, user questions are mapped to nodes and relationships in the knowledge graph. Semantic search and reasoning methods are then combined to obtain the optimal answer to the question, forming an intelligent human-computer interaction.

[0045] The system performs natural language processing on the user-input question to extract key entities and determines the relationships between these entities using dependency parsing. Based on the key entities and relationships, the question type is identified; if the question's confidence level is below a preset threshold, it is labeled as another type. The extracted entities and relationships are mapped to nodes and edges in a pre-built knowledge graph, and a pre-trained BERT model is used to generate semantic vectors for the question and knowledge graph elements. Based on the semantic vectors, the most relevant subgraph structure is determined using cosine similarity calculation, and the coverage and relevance scores of the subgraph structure are assessed. A depth-first search is performed on the subgraph structure, and the PageRank algorithm is used to rank the search paths, selecting the optimal inference path to generate a set of candidate answers. The candidate answers are then ranked and rewritten using a pre-trained language model, and terms are selected from a pre-built power industry terminology dictionary to generate a final answer that conforms to natural language expression.

[0046] Specifically, natural language processing is performed on the user-input questions. Named entity recognition technology is used to extract key entities from the questions, and dependency parsing is used to identify the relationships between entities. Simultaneously, rule-based methods and support vector machine classifiers are combined to identify question types, such as factual, causal, and comparative, with a confidence threshold of 0.8. Questions below this threshold are marked as "other types." The extracted entities and relationships are mapped to nodes and edges in the knowledge graph. A pre-trained BERT model is used to generate semantic vectors for the questions and knowledge graph elements, with a vector dimension of 768. Cosine similarity is used to calculate semantic similarity, with a threshold of 0.7, to determine the most relevant subgraph structure. The obtained subgraph structures are evaluated, calculating their coverage and relevance scores. Based on the determined subgraph structures, a depth-first search algorithm is used to traverse the graph, with a maximum depth of 5. The PageRank algorithm is used to sort the paths, with 100 iterations and a damping coefficient of 0.85, selecting the optimal reasoning path and generating a set of candidate answers. A pre-trained language model is used to sort and rewrite the candidate answers, generating the final answer that conforms to natural language expression. In this process, a power industry terminology dictionary containing 5,000 professional terms was constructed, prioritizing terms from the dictionary to ensure the professionalism of the answers. Rule templates were used to handle specific power problem types, such as load forecasting and fault diagnosis. The source and reasoning process of the answers were recorded to achieve interpretable intelligent question answering. In the implementation of the intelligent question answering system for the power industry, the user-input question "What are the causes of excessive transformer oil temperature?" was processed first. A BiLSTM-CRF-based named entity recognition model was used to identify the key entity "transformer oil temperature," achieving an F1 score of 0.95. Dependency parsing was used to identify the relationship "cause." Simultaneously, a support vector machine classifier was used to identify the question type, with a confidence score of 0.92, classifying it as a causal question. Subsequently, a BERT model was used to generate 768-dimensional semantic vectors, and the cosine similarity between "transformer oil temperature" and "cause" and nodes in the knowledge graph was calculated. The top 5 nodes with the highest similarity are "transformer", "oil temperature", "overheating", "fault", and "load", with similarities of 0.89, 0.87, 0.85, 0.82, and 0.79, respectively. A subgraph containing these nodes is constructed with a coverage of 0.85. A depth-first search algorithm is used to traverse the subgraph, with a maximum depth of 5, to find 20 possible paths. The PageRank algorithm is applied to sort these paths, and after 100 iterations, the optimal path is found to be "transformer-oil temperature-overheating-overload". Based on this path, the candidate answer "the transformer oil temperature is likely due to overload" is generated. Finally, the answer is optimized using the GPT-3 model, and terms such as "insulation aging" and "cooling system failure" are selected from the power industry terminology dictionary to supplement the answer.The final generated answer was: "The main causes of excessive transformer oil temperature include overheating due to excessive load, increased heat loss due to insulation aging, and cooling system failure affecting heat dissipation efficiency. It is recommended to adjust the load and check the cooling device." This answer received a professionalism score of 0.88 and an interpretability score of 0.92, meeting the needs of intelligent question answering in the power industry.

[0047] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the protection scope of one or more embodiments of this specification.

Claims

1. A method for AI-powered intelligent question answering and professional report generation in the power industry, characterized in that, The method includes: A knowledge graph for the power industry is constructed. The knowledge graph contains nodes representing knowledge units and edges representing the relationships between knowledge units. A multi-level node representation method is adopted in the knowledge graph. High-level nodes represent concept-level knowledge units, and low-level nodes represent data-level knowledge units. Power industry knowledge concepts are abstracted into high-level nodes, and various types of power data are organized into low-level nodes. A mapping relationship between the high-level nodes and the low-level nodes is established. Key information is extracted from the constructed knowledge graph to form a preliminary power knowledge base structure. The extraction process includes identifying high-frequency nodes, analyzing the connectivity between nodes, and evaluating the importance weight of nodes. Based on the preliminary knowledge base structure, structured and unstructured data in the power sector are integrated to form a complete power knowledge base, which is kept updated synchronously with the knowledge graph. Based on the power knowledge base, a natural language processing model for the power field is trained, including sequentially using a deep learning model for named entity recognition and relation extraction. Using a trained natural language processing model, semantic annotation is performed on the nodes and edges in the knowledge graph. Through entity alignment and relation mapping, new semantic information is integrated into the knowledge graph. Adaptive knowledge tracing is performed based on the knowledge graph. Starting from a given target knowledge unit, the knowledge graph is traversed using a breadth-first search strategy. The depth and breadth parameters of traversing target knowledge units of different granularities are dynamically adjusted to obtain other knowledge units that are directly or indirectly related to the target knowledge unit. The adjustment is based on the hierarchy, relevance, and time attributes of the knowledge units. In the process of knowledge tracing, when the target knowledge unit is a concept-level node, higher-level nodes are traversed first; when the target knowledge unit is a data-level node, the search is focused on the lower-level nodes. At the same time, the relationship between knowledge units at different time points is dynamically weighted. Based on the knowledge tracing results, the tracing data is analyzed and organized, and the type distribution, correlation strength, and time distribution characteristics of the tracing knowledge units are statistically analyzed. The data is then organized according to a predefined power professional report template, the report chapters are filled in, and a complete power professional report is generated. By leveraging the power industry knowledge graph and power knowledge base, an intelligent question-answering system for the power industry is constructed. Through natural language understanding methods, user questions are mapped to nodes and relationships in the knowledge graph. Combined with semantic search and reasoning methods, the optimal answer to the question is obtained, forming an intelligent human-computer interaction.

2. The method according to claim 1, characterized in that, The construction of a power industry knowledge graph includes nodes representing knowledge units and edges representing the relationships between knowledge units. The knowledge graph employs a multi-level node representation, with higher-level nodes representing concept-level knowledge units and lower-level nodes representing data-level knowledge units. Power industry knowledge concepts are abstracted into higher-level nodes, and various types of power data are organized into lower-level nodes. A mapping relationship is established between the higher-level nodes and the lower-level nodes, including: Acquire power industry concepts and power data, abstract the power industry concepts into high-level nodes, and organize the power data into low-level nodes; Semantic analysis is performed based on the high-level nodes and low-level nodes to determine the mapping relationship between nodes; Calculate the strength of the association between nodes to obtain the weight attributes of the association edges; By combining timestamps and version numbers to record changes in associated edges, the temporal attributes of the associated edges can be obtained; Acquire new data streams and identify new knowledge units from the new data streams; Calculate the semantic similarity between the newly added knowledge unit and the existing nodes. If the semantic similarity is higher than a preset threshold, an association relationship is established. Reasoning is performed on the aforementioned relationships to obtain the association analysis results of explicit knowledge.

3. The method according to claim 1, characterized in that, The process of extracting key information from the constructed knowledge graph to form a preliminary power knowledge base structure includes identifying high-frequency nodes, analyzing the connectivity between nodes, and evaluating the importance weight of nodes, including: The knowledge graph is traversed using a depth-first search algorithm to obtain the frequency of each node, and a hash table is used to record the mapping relationship between the node ID and its frequency of occurrence. Based on the mapping relationship between node ID and occurrence frequency, the connectivity degree of each node in the knowledge graph is calculated, and the connection relationship between nodes is represented by a weighted adjacency list structure. After the weighted adjacency list structure is represented, the multidimensional feature vector of the node is obtained based on the node frequency, connectivity, node type and number of attributes. The importance weight of a node is determined using the multidimensional feature vector. Based on the importance weights of the nodes, key nodes and their related edges are selected to construct a subgraph structure; After the subgraph structure is constructed, the Louvain algorithm is used to perform community detection on the subgraph to determine the preliminary hierarchical structure of the power knowledge base.

4. The method according to claim 1, characterized in that, Based on the preliminary knowledge base structure, a complete power knowledge base is formed by integrating structured and unstructured data from the power sector, and this knowledge base is kept synchronized with the knowledge graph. This includes: Obtain structured data from power equipment databases, SCADA systems, and energy management systems; The structured data is processed to unify units and standardize naming to obtain standardized power equipment data; The BERT model is used to process power industry documents, operation manuals, and fault reports to identify the names, parameters, and operating actions of power equipment. Based on the identification results, a relationship extraction model is constructed to extract the topological relationships and operational process information between devices; Obtain a predefined ontology model for the power domain; Based on the aforementioned power domain ontology model, semantic alignment is performed on standardized power equipment data and extracted equipment relationship information; The Word2Vec algorithm was used to train word vectors for the power industry. The similarity between concepts is calculated based on the word vectors. If the similarity exceeds a preset threshold, the corresponding concepts will be merged. Establish a two-way synchronous update mechanism between the knowledge graph and the knowledge base; If a knowledge graph update message is detected, the incremental update procedure of the knowledge base is triggered. Based on the results of real-time data stream processing, perform corresponding knowledge base update operations.

5. The method according to claim 1, characterized in that, The training of a natural language processing model for the power field based on the power knowledge base includes sequentially employing a deep learning model for named entity recognition and relation extraction, including: The received power industry standard documents, equipment manuals, and operation reports are preprocessed to obtain a power industry dictionary; Based on the power industry dictionary and the preprocessed data, the preprocessed data is used to train a named entity recognition task to obtain a named entity recognition model; Using the entity information identified by the named entity recognition model and the known relationships in the knowledge base, the unannotated text is annotated to obtain annotated text; For the annotated text, a recurrent neural network is used to construct a relation extraction model, which determines the relationship between entity pairs by inputting the contextual representation of the entity pairs.

6. The method according to claim 1, characterized in that, The step involves using a trained natural language processing model to semantically annotate nodes and edges in the knowledge graph, and integrating new semantic information into the knowledge graph through entity alignment and relation mapping, including: The trained natural language processing model is used to obtain node information in the knowledge graph, and the node information is semantically labeled to obtain the semantic type and attributes of the node. Based on the semantic type and attributes of the nodes, the Word2Vec algorithm is used to generate semantic vector representations of the nodes; Based on the semantic vector representation of the nodes, entity alignment is performed using a cosine similarity calculation method, which includes setting a similarity threshold. If the cosine similarity is higher than the preset threshold, the entity attributes are compared. If the overlap of the entity attributes is lower than a preset value, a hierarchical clustering algorithm is used to cluster the entity attributes and determine the splitting boundary. Relations are extracted from the edge information in the knowledge graph, and semantic annotation of the edge information is performed using a convolutional neural network. Based on the semantic annotation results, the semantic types of edges are mapped to predefined relation ontology; The mapped semantic information is integrated with the original knowledge graph to obtain the integrated knowledge graph; Using the batch update function of the graph database, the attributes of nodes and edges are updated according to the integrated knowledge graph to obtain the updated knowledge graph.

7. The method according to claim 1, characterized in that, The adaptive knowledge tracing based on the knowledge graph starts from a given target knowledge unit, traverses the knowledge graph using a breadth-first search strategy, dynamically adjusts the depth and breadth parameters of traversing target knowledge units of different granularities, and obtains other knowledge units directly or indirectly related to the target knowledge unit. The adjustment is based on the hierarchy, relevance, and time attributes of the knowledge units, including: Receive a search request carrying a target knowledge unit identifier, initialize a breadth-first search queue according to the search request, and set the initial search depth and breadth parameters; Calculate the search weight, which is determined by the node level, wherein the weight of the level where the target knowledge unit is located is 1; Assign search priorities to nodes at different levels based on the search weights; Obtain the semantic vectors of the target knowledge unit and the currently traversed node, and calculate the similarity score of the semantic vectors; The embedded representation of the learning nodes is used to calculate the association score between nodes; By combining the similarity score, the relevance score, and the edge weight attribute, a comprehensive relevance score is obtained. Receive the time attribute information of the node; Calculate the time decay coefficient based on the time attribute information; Multiplying the time decay coefficient by the node importance yields a node priority that incorporates the time factor. If the number of traversed nodes reaches a preset threshold, then calculate the average correlation and standard deviation of the traversed nodes. Update the preset threshold based on the average correlation degree and standard deviation; The search depth and breadth parameters are dynamically adjusted using the updated thresholds.

8. The method according to claim 1, characterized in that, In the knowledge tracing process, when the target knowledge unit is a concept-level node, higher-level nodes are traversed first; when the target knowledge unit is a data-level node, the search focuses on lower-level nodes. Simultaneously, the relationships between knowledge units at different time points are dynamically weighted, including: The target knowledge unit is obtained, and the type of the target knowledge unit is identified by a support vector machine classifier, and the target knowledge unit is divided into concept-level nodes or data-level nodes. The search parameters are dynamically adjusted based on the type of the target knowledge unit. If the target knowledge unit is a concept-level node, the search breadth parameter will be increased from the default value to the preset value, and the depth limit will be reduced from the preset value to the specified value. If the target knowledge unit is a data-level node, the search depth limit will be increased from the preset value to the specified value, and the breadth parameter will be decreased from the default value to the preset value. Analyze the timestamp information of the target knowledge unit and calculate the time weight of the target knowledge unit; By combining the type and hierarchical information of the target knowledge unit with the time weight, the relationship strength between the target knowledge unit and other knowledge units is calculated; Based on the strength of the relationship, the traversal queue is prioritized and the traversal order and depth are dynamically adjusted to optimize the knowledge tracing path.

9. The method according to claim 1, characterized in that, Based on the knowledge tracing results, the tracing data is analyzed and organized, and the type distribution, correlation strength, and time distribution characteristics of the tracing knowledge units are statistically analyzed. The data is then organized according to a predefined power industry report template, and the report chapters are filled in to generate a complete power industry report, including: Obtain a set of knowledge units, and perform cluster analysis on the set of knowledge units using the DBSCAN algorithm based on the preset neighborhood radius and minimum number of samples; Identify the noise points in the knowledge unit set that cannot be clustered, and classify the noise points into other types; Calculate the association strength between the knowledge units, construct a relationship graph of the knowledge units, where the nodes of the relationship graph are the knowledge units and the edges are the association relationships between the knowledge units; Based on the relationship diagram, the importance of the knowledge unit is evaluated using the PageRank algorithm, where the damping factor of the PageRank algorithm is a preset value and the number of iterations is a preset number. By combining the weights of the associated edges in the relationship graph, a comprehensive association strength score between the knowledge units is obtained; Time series analysis is performed on the knowledge units to identify their update cycle and trend. Draw a heatmap of the time distribution of the knowledge units, where the horizontal axis represents the year, the vertical axis represents the month, and the color intensity represents the number of knowledge units. Predict the update trend of the knowledge unit within a preset time period in the future, and generate a time distribution characteristic report of the knowledge unit; Obtain a predefined power industry report template and extract its structural features and keywords. The knowledge units are preprocessed with text, including word segmentation and stop word removal; Calculate the text feature vector of the knowledge unit; Based on the text feature vector, calculate the cosine similarity between the knowledge unit and the chapter of the report template; If the cosine similarity is higher than a preset threshold, the knowledge unit will be filled into the corresponding chapter of the report template to generate a complete power industry report.

10. The method according to claim 1, characterized in that, The system utilizes a power industry knowledge graph and knowledge base to construct an intelligent question-answering system for the power industry. It maps user questions to nodes and relationships in the knowledge graph using natural language understanding methods, and combines semantic search and reasoning methods to obtain the optimal answer, forming an intelligent human-computer interaction. This includes: Natural language processing is performed on the user-input question to extract key entities from the question, and the relationships between the entities are determined through dependency parsing. Based on the key entities and relationships, identify the problem type; if the confidence level of the problem is lower than a preset threshold, mark the problem as another type. The extracted entities and relationships are mapped to nodes and edges in a pre-built knowledge graph, and a pre-trained BERT model is used to generate semantic vectors for questions and knowledge graph elements. Based on the semantic vector, the most relevant subgraph structure is determined by cosine similarity calculation, and the coverage and relevance score of the subgraph structure are judged. A depth-first search is performed on the subgraph structure, and the PageRank algorithm is used to sort the search paths. The optimal reasoning path is selected to generate a set of candidate answers. The candidate answers are sorted and rewritten using a pre-trained language model, and terms are selected from a pre-established dictionary of power industry terms to generate a final answer that conforms to natural language expression.

Citation Information

Cited By

  • Power multi-modal knowledge base question-answer pair construction method and device

    CN121745252A

  • Method and device for constructing power multi-modal knowledge base question pair

    CN121745252B

  • Defense equipment intelligence automatic report generation method based on large language model

    CN122173629B