Power knowledge processing method and device based on layered architecture, computer device and storage medium
By processing multi-source heterogeneous power knowledge data through a layered architecture and constructing a dynamic knowledge graph, the challenges of knowledge management and utilization in the power industry have been solved. This has enabled efficient management of power system knowledge and accurate report generation, thereby enhancing the scientific decision-making capabilities of power businesses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG POWER GRID CO LTD
- Filing Date
- 2025-09-04
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional power knowledge processing technologies cannot effectively manage and utilize multi-source heterogeneous power knowledge data, making it difficult to accurately classify and label, construct knowledge graphs, and generate high-quality processing reports. They cannot meet the power industry's needs for efficient knowledge management and utilization.
A hierarchical architecture-based approach is adopted to acquire multi-source heterogeneous power knowledge data. Through standardized processing and a multi-dimensional classification system, structured, unstructured, and real-time streaming data are processed to construct a dynamic knowledge graph and generate processing reports. This approach breaks through the limitations of single-dimensional classification, improves data consistency and reliability, enhances the traceability and relevance of knowledge, and supports intelligent reasoning and the practicality of reports.
It achieves comprehensive coverage and precise management of power system knowledge, improves data utilization efficiency and reporting accuracy, meets the power industry's needs for efficient knowledge management and utilization, and supports scientific decision-making in power business.
Smart Images

Figure CN121093277B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically to a method, apparatus, medium, and device for processing power knowledge based on a hierarchical architecture. Background Technology
[0002] Against the backdrop of rapid development in today's power industry, the scale of power systems continues to expand, and their complexity is constantly increasing. From various energy generating units at the generation end to the intricate lines and substations in the transmission network, and then to the diverse user needs on the consumption side, the entire power ecosystem involves massive amounts of data and knowledge. Regarding structured data, power companies have accumulated a large amount of ledger information at the equipment asset level, such as detailed parameters, purchase dates, and maintenance records for various generating equipment, transmission lines, and substation equipment. This data is stored in structured formats such as tables. Simultaneously, business-level process forms cover forms from all stages of power planning, construction, operation and maintenance to marketing, such as project progress reports, equipment maintenance work orders, and customer electricity application forms. Unstructured data also occupies an important position in the power sector. Policy documents in the knowledge type dimension, including national and local power industry policies and environmental protection policies that constrain the power industry, as well as technical documents in the professional dimension, such as academic journal articles, internal enterprise technical reports, and equipment operation manuals, contain a large amount of valuable industry knowledge. Regarding real-time streaming data, with the construction of smart grids, a large amount of equipment operation status monitoring data has been generated in the power system, such as real-time parameters of generator speed, temperature, voltage, and current, and status data of transmission lines such as sag, icing, and galloping. This real-time streaming data is characterized by large data volume, high flow rate, and high timeliness.
[0003] Traditional knowledge classification often focuses on only one or a few dimensions, such as simple classification based on equipment type or business process, which is insufficient to adapt to the complex knowledge structure of the power industry. This makes it difficult to accurately locate knowledge during storage and retrieval, reducing the efficiency of knowledge utilization. When generating processing reports, traditional technologies lack a deep understanding and targeted analysis of power business scenarios. When facing equipment operation and maintenance scenarios, they cannot comprehensively integrate equipment asset information, historical failure cases, professional technical standards, and other knowledge to generate accurate and instructive fault diagnosis reports and effective historical case matching results; in policy compliance scenarios, it is difficult to quickly and accurately analyze the compatibility between solutions and policies, and generate high-quality solution and policy compatibility reports.
[0004] In summary, the relevant technologies have significant shortcomings in processing multi-source heterogeneous power knowledge data, classifying and labeling it, constructing knowledge graphs, and generating processing reports, and cannot meet the power industry's needs for efficient management and utilization of knowledge. Summary of the Invention
[0005] In view of this, the present invention provides a power knowledge processing method, apparatus, medium and equipment based on a hierarchical architecture to solve the problem that related technologies have obvious defects in processing multi-source heterogeneous power knowledge data, classification and labeling, knowledge graph construction and processing report generation, and cannot meet the power industry's needs for efficient management and utilization of knowledge.
[0006] In a first aspect, the present invention provides a power knowledge processing method based on a hierarchical architecture. The method includes: acquiring multi-source heterogeneous power knowledge data, including structured data, unstructured data, and real-time streaming data; performing standardization processing on the structured data, unstructured data, and real-time streaming data respectively; classifying and labeling the standardized structured knowledge data based on a multi-dimensional classification system, including a primary dimension and a secondary dimension; converting the standardized unstructured text into triplet data; constructing a dynamic knowledge graph based on the triplet data, labeled data, and standardized real-time streaming data, with equipment, faults, and procedures as core entities; and generating a processing report based on the power business scenario type and the dynamic knowledge graph.
[0007] The power knowledge processing method based on a layered architecture provided in this invention firstly acquires multi-source heterogeneous power knowledge data, including structured equipment ledgers and business forms, unstructured policy documents and technical literature, and real-time equipment operation monitoring data, achieving comprehensive coverage of knowledge data and providing a complete data source for subsequent in-depth processing. By integrating multiple types of knowledge, it avoids analytical bias caused by data dispersion, ensuring mastery of knowledge across all scenarios of power system operation and management. Furthermore, by establishing a foundation for the correlation of multi-source data, it creates conditions for mining potential connections between different types of data. Secondly, by unifying and verifying the fields of structured data, it solves the problems of conflicting formats and inconsistent parameters in structured data from different sources, improving data consistency and reliability. Through parsing, extraction, and semantic cleaning of unstructured data, it transforms text and documents that are difficult to use directly into standardized, processable formats, releasing the value of unstructured knowledge. Through edge preprocessing and standardized encapsulation of real-time streaming data, it filters noise and retains key features, ensuring the effectiveness and timeliness of real-time data. Finally, through targeted processing of the three types of data, it lays a unified data foundation for subsequent classification, labeling, and knowledge fusion, reducing processing obstacles caused by data format differences. Then, by combining primary and secondary dimensions in its classification approach, the limitations of single-dimensional classification are overcome, enabling precise characterization of knowledge attributes from multiple perspectives, including knowledge type, business domain, professional scope, equipment association, and organizational affiliation. Intelligent semantic parsing and feature extraction improve the accuracy and efficiency of classification and labeling, reducing the subjectivity of manual intervention. Hierarchical category matching and encoding generation give knowledge a clear organizational structure, facilitating rapid retrieval and location. Dynamic filtering and priority sorting ensure that classification and labeling can adapt to the needs of different business scenarios, enhancing the flexibility of knowledge management. By forming standardized labeled data, a structured classification basis is provided for knowledge graph construction, promoting the systematic organization of knowledge. Subsequently, by accurately identifying core entities such as equipment, faults, and procedures, the problem of extracting key information from unstructured text was solved, and the core elements of knowledge were clarified. By parsing the semantic relationships between entities, the logical connections within knowledge were revealed, transforming scattered textual information into structured "entity-relationship" units. By associating multi-dimensional classification information, triplet data and classification systems were linked, enhancing the traceability and relevance of knowledge. Through the structured representation of triples, basic data was provided for establishing entity relationships in knowledge graphs, improving the reusability and reasoning ability of knowledge.Furthermore, by building a framework around equipment, faults, and procedures as core entities, the key knowledge elements of the power system are focused on, ensuring the domain relevance and practicality of the graph. By integrating triples, labeled data, and real-time streaming data, the organic combination of static knowledge and dynamic information is achieved, enriching the content dimensions of the graph. By mining implicit relationships between entities and completing paths, the correlation between knowledge is enhanced, supporting deeper intelligent reasoning. Through cross-dimensional consistency verification and dynamic update mechanisms, knowledge conflicts are promptly discovered and corrected, ensuring the accuracy and timeliness of the graph. Through the dynamic characteristics of the graph, it can adapt to changes in the power system, continuously providing reliable knowledge support for business operations. Finally, by customizing report frameworks for different scenarios such as equipment operation and maintenance and policy compliance, the report content is ensured to accurately match business needs, thus improving the report's practicality. Through in-depth analysis of entity-related data and combined with generative technologies, complex knowledge is transformed into easily understandable structured conclusions, lowering the barrier to entry for professional knowledge. By integrating multi-dimensional knowledge and annotating the sources of evidence, the credibility and interpretability of the report conclusions are enhanced. By outputting customized reports, the decision-making needs of different business scenarios are met, improving work efficiency. Through standardized report presentation, the effective transfer and application of knowledge are promoted, providing strong support for scientific decision-making in the power industry. In summary, by implementing this invention, the significant shortcomings of related technologies in processing multi-source heterogeneous power knowledge data, classification and annotation, knowledge graph construction, and report generation are addressed, failing to meet the power industry's needs for efficient knowledge management and utilization.
[0008] In one optional implementation, the structured data includes ledger information at the equipment asset level and process form data at the business level; the unstructured data includes policy documents at the knowledge type level and technical documents at the professional level; and the real-time streaming data includes equipment operation status monitoring data. The above-mentioned structured data, unstructured data, and real-time streaming data are subjected to standardization processing, including: performing unified field mapping on the structured data based on preset metadata standards and verifying compliance of the structured data by associating it with corresponding professional technical standards; parsing and extracting the unstructured data using multimodal documents, converting the unstructured data into a parsable format, and performing semantic cleaning; performing noise filtering and feature parameter extraction on the real-time streaming data through edge nodes, and standardizing and encapsulating the processed real-time streaming data based on time series and equipment identifiers.
[0009] In one optional implementation, the main dimensions include knowledge type LX, business YW, and professional ZY, and the auxiliary dimensions include equipment asset SB and organizational structure JG. The standardized structured knowledge data is classified and labeled based on a multi-dimensional classification system to obtain labeled data. This includes: using a pre-trained large-scale model in the power industry to perform semantic parsing on the structured knowledge data, extracting business theme features, technical field features, and core object features; matching knowledge type LX, business YW, and professional ZY based on the business theme features, technical field features, and core object features, and forcibly associating and sorting them according to the order from knowledge type LX to business YW to professional ZY; and dynamically filtering equipment assets based on knowledge attributes. The equipment assets (SB) and organizational structures (JG) are ranked based on their relevance. These knowledge attributes are determined by business scenario characteristics, business theme characteristics, technical field characteristics, and core object characteristics. For the ranked primary dimension, the hierarchical category knowledge graph under the primary dimension is invoked. Similarity calculations between feature words and category attributes are performed, matching down to the finest-grained category for annotation, resulting in primary dimension annotation data. For the ranked secondary dimension, the hierarchical category knowledge graph under the secondary dimension is invoked. Similarity calculations between feature words and category attributes are performed, matching down to the finest-grained category for annotation, resulting in secondary dimension annotation data. Annotated data is then formed based on the primary dimension annotation data and the secondary dimension annotation data.
[0010] In one optional implementation, the above-mentioned conversion of standardized unstructured text into triple data includes: using natural language processing technology to segment and identify entities in the unstructured text to obtain core entities, the core entities including equipment, faults, procedures, and business activities, the entity identification being implemented based on a pre-trained model in the power field; using a semantic relation extraction algorithm to analyze the semantic correlation between the core entities to determine the relationship type between the subject entity and the object entity; and constructing triples based on the format of subject entity, relation type to object entity, wherein the subject entity, relation type, and object entity are all associated with dimensional information in a multi-dimensional classification system.
[0011] In one optional implementation, the above-mentioned construction of a dynamic knowledge graph based on triplet data, labeled data, and standardized real-time streaming data includes: constructing a basic framework of the graph with equipment entities, fault entities, and procedure entities as core entities; establishing initial associations between core entities through triplet data, wherein equipment entities are associated with labeled data of equipment assets SB, and procedure entities are associated with labeled data of knowledge type LX; extracting temporal features from the standardized real-time streaming data, and associating the real-time streaming data with the corresponding core entities based on timestamps and equipment identifiers; mining the relationships between core entities using graph neural network algorithms, and using entity vector similarity calculations to complete the implicit association paths between equipment, faults, and procedures; employing a cross-dimensional consistency verification mechanism, using graph neural networks to detect whether there are logical conflicts between different dimensional classification systems; if logical conflicts exist, marking the difference nodes and triggering a source tracing verification process based on labeled data, and updating the dynamic knowledge graph.
[0012] In one optional implementation, the above-mentioned generation of a processing report based on power business scenario types and dynamic knowledge graphs includes: determining the power business scenario type, which includes equipment operation and maintenance scenarios and policy compliance scenarios; calling the category system of business YW corresponding to the above-mentioned power business scenario type in the dynamic knowledge graph to determine the content framework and data extraction scope of the processing report; retrieving entity association data related to the above-mentioned power business scenario type from the dynamic knowledge graph, which includes equipment assets SB, historical fault cases, procedural documents in knowledge type LX, and technical standards in professional ZY, to form the basic dataset for the report; using graph neural networks to perform deep analysis on the above-mentioned entity association data, combined with generative large model retrieval enhancement generation technology, to generate structured analysis content including scenario adaptation conclusions, basis tracing, and optimization suggestions; and integrating the above-mentioned structured analysis content into a processing report according to the output specifications of the above-mentioned power business scenario type. The processing report output for the equipment operation and maintenance scenario includes a fault diagnosis report and historical case matching results, and the processing report output for the policy compliance scenario includes a solution and policy adaptability report.
[0013] In one optional implementation, the above-mentioned classification and labeling of standardized structured knowledge data based on a multi-dimensional classification system to obtain labeled data further includes: adding dimensional attribute labels to core entities based on the labeling information of knowledge type LX, business YW, professional ZY, equipment asset SB, and organizational structure JG in the labeled data; and adjusting the weight of the relationship between core entities using a time-series decay algorithm based on the weight adjustment rule, wherein the real-time data weight of equipment asset SB increases dynamically over time, and the weight adjustment rule is related to the process node requirements of business YW.
[0014] Secondly, the present invention provides a power knowledge processing device based on a hierarchical architecture. The device includes: an acquisition module for acquiring multi-source heterogeneous power knowledge data, including structured data, unstructured data, and real-time streaming data; a standardization module for performing standardization processing on the structured data, unstructured data, and real-time streaming data respectively; an annotation module for classifying and annotating the standardized structured knowledge data based on a multi-dimensional classification system, including a primary dimension and a secondary dimension; a transformation module for converting the standardized unstructured text into triple data; a construction module for constructing a dynamic knowledge graph based on the triple data, annotated data, and standardized real-time streaming data, with equipment, faults, and procedures as core entities; and a generation module for generating a processing report based on the power business scenario type and the dynamic knowledge graph.
[0015] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the power knowledge processing method based on a hierarchical architecture as described in the first aspect or any corresponding embodiment thereof.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the power knowledge processing method based on a hierarchical architecture as described in the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a power knowledge processing method based on a hierarchical architecture according to an embodiment of the present invention.
[0019] Figure 2 This is a structural block diagram of a power knowledge processing device based on a hierarchical architecture according to an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Against the backdrop of rapid development in today's power industry, the scale of power systems continues to expand, and their complexity is constantly increasing. From various energy generating units at the generation end to the intricate lines and substations in the transmission network, and then to the diverse user needs on the consumption side, the entire power ecosystem involves massive amounts of data and knowledge. Regarding structured data, power companies have accumulated a large amount of ledger information at the equipment asset level, such as detailed parameters, purchase dates, and maintenance records for various generating equipment, transmission lines, and substation equipment. This data is stored in structured formats such as tables. Simultaneously, business-level process forms cover forms from all stages of power planning, construction, operation and maintenance to marketing, such as project progress reports, equipment maintenance work orders, and customer electricity application forms. Unstructured data also occupies an important position in the power sector. Policy documents in the knowledge type dimension, including national and local power industry policies and environmental protection policies that constrain the power industry, as well as technical documents in the professional dimension, such as academic journal articles, internal enterprise technical reports, and equipment operation manuals, contain a large amount of valuable industry knowledge. Regarding real-time streaming data, with the construction of smart grids, a large amount of equipment operation status monitoring data has been generated in the power system, such as real-time parameters of generator speed, temperature, voltage, and current, and status data of transmission lines such as sag, icing, and galloping. This real-time streaming data is characterized by large data volume, high flow rate, and high timeliness.
[0023] Traditional knowledge classification often focuses on only one or a few dimensions, such as simple classification based on equipment type or business process, which is insufficient to adapt to the complex knowledge structure of the power industry. This makes it difficult to accurately locate knowledge during storage and retrieval, reducing the efficiency of knowledge utilization. When generating processing reports, traditional technologies lack a deep understanding and targeted analysis of power business scenarios. When facing equipment operation and maintenance scenarios, they cannot comprehensively integrate equipment asset information, historical failure cases, professional technical standards, and other knowledge to generate accurate and instructive fault diagnosis reports and effective historical case matching results; in policy compliance scenarios, it is difficult to quickly and accurately analyze the compatibility between solutions and policies, and generate high-quality solution and policy compatibility reports.
[0024] In summary, the relevant technologies have significant shortcomings in processing multi-source heterogeneous power knowledge data, classifying and labeling it, constructing knowledge graphs, and generating processing reports, and cannot meet the power industry's needs for efficient management and utilization of knowledge.
[0025] The power knowledge processing method based on a layered architecture provided in this invention firstly acquires multi-source heterogeneous power knowledge data, including structured equipment ledgers and business forms, unstructured policy documents and technical literature, and real-time equipment operation monitoring data, achieving comprehensive coverage of knowledge data and providing a complete data source for subsequent in-depth processing. By integrating multiple types of knowledge, it avoids analytical bias caused by data dispersion, ensuring mastery of knowledge across all scenarios of power system operation and management. Furthermore, by establishing a foundation for the correlation of multi-source data, it creates conditions for mining potential connections between different types of data. Secondly, by unifying and verifying the fields of structured data, it solves the problems of conflicting formats and inconsistent parameters in structured data from different sources, improving data consistency and reliability. Through parsing, extraction, and semantic cleaning of unstructured data, it transforms text and documents that are difficult to use directly into standardized, processable formats, releasing the value of unstructured knowledge. Through edge preprocessing and standardized encapsulation of real-time streaming data, it filters noise and retains key features, ensuring the effectiveness and timeliness of real-time data. Finally, through targeted processing of the three types of data, it lays a unified data foundation for subsequent classification, labeling, and knowledge fusion, reducing processing obstacles caused by data format differences. Then, by combining primary and secondary dimensions in its classification approach, the limitations of single-dimensional classification are overcome, enabling precise characterization of knowledge attributes from multiple perspectives, including knowledge type, business domain, professional scope, equipment association, and organizational affiliation. Intelligent semantic parsing and feature extraction improve the accuracy and efficiency of classification and labeling, reducing the subjectivity of manual intervention. Hierarchical category matching and encoding generation give knowledge a clear organizational structure, facilitating rapid retrieval and location. Dynamic filtering and priority sorting ensure that classification and labeling can adapt to the needs of different business scenarios, enhancing the flexibility of knowledge management. By forming standardized labeled data, a structured classification basis is provided for knowledge graph construction, promoting the systematic organization of knowledge. Subsequently, by accurately identifying core entities such as equipment, faults, and procedures, the problem of extracting key information from unstructured text was solved, and the core elements of knowledge were clarified. By parsing the semantic relationships between entities, the logical connections within knowledge were revealed, transforming scattered textual information into structured "entity-relationship" units. By associating multi-dimensional classification information, triplet data and classification systems were linked, enhancing the traceability and relevance of knowledge. Through the structured representation of triples, basic data was provided for establishing entity relationships in knowledge graphs, improving the reusability and reasoning ability of knowledge.Furthermore, by building a framework around equipment, faults, and procedures as core entities, the key knowledge elements of the power system are focused on, ensuring the domain relevance and practicality of the graph. By integrating triples, labeled data, and real-time streaming data, the organic combination of static knowledge and dynamic information is achieved, enriching the content dimensions of the graph. By mining implicit relationships between entities and completing paths, the correlation between knowledge is enhanced, supporting deeper intelligent reasoning. Through cross-dimensional consistency verification and dynamic update mechanisms, knowledge conflicts are promptly discovered and corrected, ensuring the accuracy and timeliness of the graph. Through the dynamic characteristics of the graph, it can adapt to changes in the power system, continuously providing reliable knowledge support for business operations. Finally, by customizing report frameworks for different scenarios such as equipment operation and maintenance and policy compliance, the report content is ensured to accurately match business needs, thus improving the report's practicality. Through in-depth analysis of entity-related data and combined with generative technologies, complex knowledge is transformed into easily understandable structured conclusions, lowering the barrier to entry for professional knowledge. By integrating multi-dimensional knowledge and annotating the sources of evidence, the credibility and interpretability of the report conclusions are enhanced. By outputting customized reports, the decision-making needs of different business scenarios are met, improving work efficiency. Through standardized report presentation, the effective transfer and application of knowledge are promoted, providing strong support for scientific decision-making in the power industry. In summary, by implementing this invention, the significant shortcomings of related technologies in processing multi-source heterogeneous power knowledge data, classification and annotation, knowledge graph construction, and report generation are addressed, failing to meet the power industry's needs for efficient knowledge management and utilization.
[0026] According to an embodiment of the present invention, an embodiment of a power knowledge processing method based on a hierarchical architecture is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0027] This embodiment provides a power knowledge processing method based on a hierarchical architecture. Figure 1 This is a flowchart of a power knowledge processing method based on a hierarchical architecture according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:
[0028] Step S101: Obtain multi-source heterogeneous power knowledge data, which includes structured data, unstructured data, and real-time streaming data.
[0029] Specifically, for structured data, by connecting to the power company's internal business systems and databases, batch extraction of equipment asset-level ledger information, such as equipment model, purchase date, and installation location, as well as business-level process form data, such as maintenance work orders and approval records, is performed. Data interface adaptation technology is used during the extraction process to ensure data format compatibility across different systems. For unstructured data, a document management system is used to collect policy documents of the knowledge type dimension, including industry regulations and corporate rules, as well as technical documents of the professional dimension, such as equipment manuals and academic papers. Simultaneously, web crawling tools are used to selectively crawl publicly available power technology materials, with filtering rules set to select valid content. Real-time streaming data is collected through sensors and monitoring equipment deployed at substations, transmission lines, and other sites, covering status parameters such as voltage, current, and temperature during equipment operation. During the collection process, edge gateways are used to achieve real-time data transmission and initial aggregation, ensuring data continuity and timeliness. During data acquisition, a data source identification mechanism needs to be established, adding source tags to each type of data. For example, structured data should be labeled "ERP System" or "Asset Management System," unstructured data should be labeled "Internal Document Library" or "External Knowledge Base," and real-time streaming data should be labeled "Sensor Number" or "Monitoring Point Location," facilitating subsequent data source tracing. Simultaneously, a data collection verification process should be implemented. This includes checking the integrity of extracted structured data to ensure no key fields are missing; validating the format of collected unstructured data and removing damaged or unopenable files; and monitoring the transmission quality of real-time streaming data, triggering a retransmission mechanism in case of data interruption or packet loss. Furthermore, a data acquisition log should be established to record the time, quantity, and status of data acquisition, creating a traceable record of the acquisition process.
[0030] Step S102: Perform standardization processing on the structured data, unstructured data and real-time streaming data respectively.
[0031] Specifically, the structured data includes ledger information at the equipment asset level and process form data at the business level; the unstructured data includes policy documents at the knowledge type level and technical documents at the professional level; and the real-time streaming data includes equipment operation status monitoring data. Step S102 includes:
[0032] Step a1: Based on the preset metadata standard, perform unified field mapping on the above structured data, and associate the corresponding professional technical standards to perform compliance verification on the above structured data.
[0033] Furthermore, for the ledger information in the dimension of equipment assets and the process form data in the business dimension, first sort out the field differences of structured data from different sources, and establish a comparison table including original fields, standard fields, and data type mapping relationships. For example, map the "equipment model" and "specification model" in different systems to the standard field of "equipment specification model" uniformly; for the missing standard fields such as "manufacturer information" in the ledger, supplement them by associating with procurement records, and only retain the latest record for the repeated "approver" field in the process form. During compliance verification, call the professional technical standard rule library of the power industry, and compare the equipment parameters in the ledger with the corresponding technical specifications, such as whether the equipment withstand voltage value meets the industry safety standards; for the process form data, check whether the required items are complete and whether the approval process complies with the business specifications. If it is found that the data does not meet the standards, mark the abnormal fields and generate a verification report including deviation content and corresponding standard clauses, and re-execute the mapping and verification after correction until all fields meet the preset standards.
[0034] Step a2: Use multimodal documents to parse and extract the above unstructured data, convert the above unstructured data into a parsable format, and perform semantic cleaning.
[0035] Furthermore, for the policy documents in the knowledge type dimension and the technical literature in the professional dimension, use multimodal document parsing tools to process them by type: directly decompose pure text files into semantic units through a word segmentation tool; for documents containing charts, extract the text information in the charts through image recognition technology, and use a table structuring tool to convert the data table into a two-dimensional array; for technical literature containing formulas, convert them into editable formula expressions through a formula recognition engine. During format conversion, uniformly encapsulate the parsed content into structured data including "title - chapter - keywords - body content", where policy documents are subdivided into fields according to "issuing department - effective date - clause content", and technical literature is divided into modules according to "technical field - core conclusion - reference literature". In the semantic cleaning stage, first filter out meaningless words such as "of" and "and" through the power industry stop word list, and then unify synonyms such as "knife switch" and "disconnector" into the standard expression of "disconnector" based on the domain term library; for ambiguous expressions in policy documents such as "relevant regulations", replace them with specific regulation names through context association; finally, delete duplicate paragraphs and redundant information irrelevant to power business to ensure the consistency of text semantics.
[0036] Step a3: Perform noise filtering processing and feature parameter extraction on the above real-time stream data through edge nodes, and perform standardized encapsulation on the processed real-time stream data based on time series and device identification.
[0037] Furthermore, when edge nodes process equipment operation status monitoring data, they first select filtering algorithms according to data type: for continuously collected analog data such as voltage and current, a sliding window averaging method is used to smooth high-frequency noise, and instantaneous fluctuations are filtered by setting a reasonable window size; for discrete data such as equipment start-up and shutdown status, abnormal states that do not conform to the time sequence are filtered based on state transition logic rules. For example, when the same equipment simultaneously exhibits "running" and "shutdown" states, records that are continuous with the preceding and following states are retained. In the feature parameter extraction stage, key features such as "average operating value - peak value - fluctuation frequency" are extracted from real-time data, combined with equipment operating characteristics. For example, the hourly average temperature, maximum temperature, and number of temperature surges are extracted from transformer temperature data. During standardized encapsulation, the processed data is arranged in chronological order, and a unique device identifier, data acquisition timestamp, and edge node number are added to each record. Feature parameters and metadata are integrated into a standardized data unit containing "device identifier - time series - feature parameter set - data quality label". The data quality label is marked as "excellent", "average", or "to be reviewed" based on the noise filtering results. Finally, the data is uploaded to the data center in a fixed format through a preset data interface.
[0038] Step S103: Classify and label the standardized structured knowledge data based on a multi-dimensional classification system to obtain labeled data. The multi-dimensional classification system includes a main dimension and a secondary dimension.
[0039] Specifically, the main dimensions mentioned above include knowledge type LX, business YW, and professional ZY, while the secondary dimensions include equipment assets SB and organizational structure JG. Step S103 includes:
[0040] Step b1 involves using a pre-trained large model in the power sector to perform semantic parsing on the structured knowledge data, extracting business theme features, technical field features, and core object features from the structured knowledge data.
[0041] Furthermore, the standardized structured knowledge data undergoes preprocessing to remove redundant fields and format errors, ensuring data format consistency. Subsequently, the processed structured data is input into a pre-trained large-scale model in the power sector. The model performs semantic parsing of the data through word segmentation, part-of-speech tagging, and dependency parsing. During parsing, the model identifies core keywords and semantic units in the data. For example, it extracts business theme keywords such as "maintenance" and "troubleshooting" from equipment maintenance records, identifies technical terminology such as "transmission lines" and "transformer equipment" from technical parameter descriptions, and locates core objects such as "transformer" and "operation procedures" from process descriptions. Through clustering and summarizing these keywords and semantic units, the business theme features, technical technology features, and core object features of the structured knowledge data are finally extracted and output as a list of feature words, providing a basis for subsequent matching of the main dimension.
[0042] Step b2: Based on business theme features, technical field features, and core object features, match knowledge type LX, business YW, and professional ZY, and force association sorting based on the order of knowledge type LX to business YW to professional ZY.
[0043] Furthermore, a mapping relationship library between feature words and main dimension categories is established, where business theme feature words correspond to categories of knowledge type LX, and technical field feature words correspond to categories of business YW and profession ZY. Next, the feature words extracted in step b1 are compared with the categories in the mapping relationship library to initially determine the matching knowledge type, business, and profession. For example, if the business theme feature words contain "procedure" or "standard," then it matches "standard category" in knowledge type LX; if the technical field feature words contain "transmission" or "line," then it matches "transmission business" in business YW and "line profession" in profession ZY. Then, a forced association sorting is performed according to the order of knowledge type LX → business YW → profession ZY, that is, first fix the knowledge type, then associate the corresponding business under that knowledge type, and finally associate the corresponding profession under that business, forming a hierarchical association chain of "knowledge type - business - profession," ensuring the orderliness and uniqueness of the main dimension classification.
[0044] Step b3: Dynamically filter equipment assets SB and organizational structures JG based on knowledge attributes, and sort the equipment assets SB and organizational structures JG based on the priority of relevance. The above knowledge attributes are determined by business scenario characteristics, business theme characteristics, technical field characteristics and core object characteristics.
[0045] Furthermore, a knowledge attribute feature set is constructed based on business scenario characteristics, business theme characteristics, technical field characteristics, and core object characteristics. For example, the feature set of a certain knowledge might include features such as "operation and maintenance scenario," "equipment maintenance," "substation technology," and "transformer." Then, based on preset association rules, the relevance between knowledge attributes and auxiliary dimensions is determined. If a knowledge attribute contains equipment-related features, then equipment assets (SB) are selected; if it contains department or unit-related features, then organizational structures (JG) are selected. Next, the semantic similarity between the knowledge attribute features and the auxiliary dimension categories is calculated to determine the degree of association. For example, the higher the similarity between "transformer" and the "substation equipment" category in equipment assets (SB), the higher the degree of association. Finally, the selected auxiliary dimensions are sorted according to the priority of association from high to low. If a piece of knowledge is associated with both equipment assets and organizational structures, and the association with equipment assets is higher, then equipment assets (SB) are ranked before organizational structures (JG).
[0046] Step b4: For the sorted main dimension, call the hierarchical category knowledge graph under the main dimension, and match it level by level to the finest category by calculating the similarity between feature words and category attributes, and then label it to obtain the main dimension labeled data.
[0047] Furthermore, the hierarchical category knowledge graph under the main dimension is invoked. This graph contains the hierarchical structure of knowledge type LX, business YW, and professional ZY, progressively subdivided from the top-level category to the finest-grained category. Then, for each sorted main dimension, its corresponding feature words are compared with the attributes of the top-level category in the knowledge graph for that dimension. For example, the feature words of the knowledge type are compared with the attributes of top-level categories such as "standards" and "case studies," and the category with the highest similarity is selected as the matching result. Next, based on this matching result, the process of calculating the similarity between feature words and category attributes is repeated at the next level until the finest-grained category is matched. For example, the knowledge type is matched step-by-step from "standards" to "operation and maintenance standards" and then to "transmission operation and maintenance standards," ultimately completing the annotation and forming the main dimension annotation data, such as "knowledge type: transmission operation and maintenance standards."
[0048] Step b5: For the sorted secondary dimension, call the hierarchical category knowledge graph under the secondary dimension, and match it step by step to the finest granular category by calculating the similarity between feature words and category attributes, and then label it to obtain the secondary dimension labeled data.
[0049] Furthermore, the hierarchical category knowledge graph under the auxiliary dimension is invoked. This graph contains the hierarchical structure of equipment assets (SB) and organizational structures (JG). For each sorted auxiliary dimension, the similarity between its corresponding feature words and the attributes of the top-level category in that dimension of the graph is calculated. For example, the feature words of equipment assets are compared with the attributes of top-level categories such as "transmission equipment" and "transformer equipment," and the most matching category is selected. Then, the matching proceeds level by level to the next level category until the finest granularity is reached. For example, equipment assets are matched level by level from "transformer equipment" to "transformer" and "oil-immersed transformer." Finally, the matching results are labeled to form auxiliary dimension labeled data, such as "equipment assets: oil-immersed transformer."
[0050] Step b6: Based on the main dimension annotation data and the secondary dimension annotation data, form the annotation data.
[0051] Furthermore, a unique code is assigned to each primary dimension, secondary dimension, and its subcategories. The dimension code distinguishes between knowledge type (LX), business (YW), specialty (ZY), equipment asset (SB), and organization (JG), while the category code identifies the specific category within each dimension. Next, the primary dimension annotation data obtained in step b4 is converted to a "dimension code + category code" format. For example, the knowledge type "transmission operation and maintenance specifications" corresponds to "LX+001," and the business "transmission business" corresponds to "YW+002." Similarly, the secondary dimension annotation data obtained in step b5 is encoded, such as the equipment asset "oil-immersed transformer" corresponding to "SB+003." Finally, the encoding results of the primary and secondary dimensions are integrated to form complete annotation data, such as "LX+001; YW+002; ZY+003; SB+003," achieving multi-dimensional encoding and annotation of structured knowledge data.
[0052] Step S104: Convert the standardized unstructured text into triplet data.
[0053] Specifically, step S104 includes:
[0054] Step c1 involves using natural language processing technology to segment and identify entities in the unstructured text to obtain core entities. These core entities include equipment, faults, procedures, and business activities. The entity recognition is based on a pre-trained model in the power industry.
[0055] Further, when performing word segmentation, first load the professional dictionary of the power industry, and use the bidirectional maximum matching method to segment the unstructured text, ensuring that compound professional terms such as "transformer" and "switchgear cabinet" are not split. At the same time, filter out虚词 such as "de" (的) and "zai" (在) that have no practical meaning. In the entity recognition stage, input the segmented text into the pre-trained model in the power field. This model encodes the text through a multi-layer Transformer structure and outputs the entity label probabilities of each word or phrase. Entities with labels of "equipment", "fault", "regulation", and "business activity" are selected according to the probability values. For example, the model will label "cable line" as an equipment entity, "short circuit fault" as a fault entity, "live working regulation" as a regulation entity, and "line inspection" as a business activity entity. Then, deduplicate and standardize the identified entities to unify the entity representation form. For example, merge "transformer" and "main transformer" into the same equipment entity, and finally form a core entity list.
[0056] Step c2, use the semantic relation extraction algorithm to analyze the semantic association degree between the above core entities, and determine the relationship type between the subject entity and the object entity.
[0057] Further, first combine the identified core entities pairwise to construct entity pairs. For each entity pair, intercept the context fragment in the unstructured text, and extract the sentence or paragraph containing the entity pair as the context for relationship analysis. Use the semantic relation extraction algorithm based on dependency syntactic analysis to parse the syntactic structure between entities in the context, identify syntactic relations such as subject-predicate, verb-object, and modifier-head, and initially judge the association direction between entities. At the same time, combined with the pre-set relationship type library in the power field, such as "occur", "cause", "basis", "execute", etc., determine the most likely relationship type by calculating the semantic similarity between the entity pair and each relationship type. For example, for the entity pair of "cable line" and "short circuit fault", through analyzing the context "The cable line has a short circuit fault", identify the relationship type of "occur"; for "short circuit fault" and "live working regulation", according to the context of "Handling short circuit faults requires referring to the live working regulation", determine the relationship type of "basis". Finally, perform artificial verification rule adaptation on the extracted relationship types to filter out incorrect relationships that do not conform to the power business logic.
[0058] Step c3, construct triples based on the format of subject entity, relationship type to object entity, where the subject entity, relationship type, and object entity are all associated with the dimension information in the multi-dimensional classification system.
[0059] Furthermore, following a fixed format from subject entity, relation type to object entity, the entity pairs and corresponding relation types determined in step c2 are combined into triples. For example, "cable line," "occurrence," and "short circuit fault" are combined into <cable line, occurrence, short circuit fault>, and "short circuit fault," "basis," and "live-line working procedure" are combined into <short circuit fault, basis, live-line working procedure>. When associating with a multi-dimensional classification system, subject entities and object entities are matched with corresponding dimension categories through preset mapping rules. For example, the equipment entity "cable line" is associated with the "line equipment" category of the equipment asset dimension, and the procedure entity "live-line working procedure" is associated with the "operation specifications" category of the knowledge type dimension. The relation type determines the corresponding classification information based on the associated entity dimension. For example, the "occurrence" relation is associated with the "equipment fault handling" category of the business dimension, and the "basis" relation is associated with the "safe operation" category of the professional dimension. The resulting triples not only contain entities and relationships, but also include corresponding dimensional classification labels for each part, thus enabling association with a multi-dimensional classification system.
[0060] Step S105: Construct a dynamic knowledge graph based on triplet data, labeled data, and standardized real-time streaming data. The knowledge graph uses equipment, faults, and procedures as core entities.
[0061] Specifically, step S105 includes:
[0062] Step d1: Construct a basic framework for the graph using equipment entities, fault entities, and procedure entities as core entities. Establish initial relationships between core entities through triplet data. Specifically, equipment entities are associated with labeled data of equipment assets SB, and procedure entities are associated with labeled data of knowledge types LX.
[0063] Furthermore, when constructing the basic framework of the graph, three core entities—equipment, faults, and procedures—are first extracted from the triplet data. Each entity is assigned a unique identifier and its basic attributes are defined, such as the equipment's model parameters, the fault's symptom description, and the procedure's release date. By traversing the subject entity, relation type, and object entity structure within the triplet, the core entities are connected according to their corresponding relationships. For example, "transformer" is associated with "insulation aging" through an "occurrence" relationship, and "insulation aging" is associated with "transformer maintenance procedure" through a "corresponding processing" relationship. Simultaneously, by matching the entity identifier with the equipment asset SB code and knowledge type LX code in the labeled data, equipment entities are automatically associated with their respective asset classification information, and procedure entities are bound to their corresponding knowledge type classification tags.
[0064] Step d2 involves extracting time-series features from the standardized real-time streaming data and associating the real-time streaming data with the corresponding core entities based on timestamps and device identifiers.
[0065] Furthermore, a sliding window method is used to extract time-series features, such as the average value of equipment parameters, fluctuation amplitude, and number of abrupt changes within a fixed time interval. A timestamp field and a unique equipment identifier field are added to each real-time data point. The equipment identifier is precisely matched with the identifier of the equipment entity in the graph, and the extracted time-series features are dynamically attached to the corresponding equipment entity as attributes. For example, if a real-time stream contains the timestamp "2023-10-01 08:00" and the identifier "Transformer T101", then its corresponding voltage fluctuation feature is associated with the "Real-time Operation Features" attribute of the "Transformer T101" entity in the graph, forming a dynamically updated entity attribute chain.
[0066] Step d3 involves using a graph neural network algorithm to mine the relationships between core entities and using entity vector similarity calculation to complete the implicit association paths of the equipment association fault association procedures.
[0067] Furthermore, the core entities and their relationships in the graph are first transformed into graph structure data, with each entity as a node and relationships as edges. A graph embedding algorithm is used to convert the nodes into low-dimensional vectors, and the cosine similarity between any two node vectors is calculated. Entity pairs with similarity higher than a set threshold are selected. For missing intermediate relationships in the "device-fault-procedure" path, such as knowing device A and procedure B but not directly related, if the similarity between device A and a fault C, and the similarity between fault C and procedure B both meet the threshold, the "device A - fault C - procedure B" relationship path is automatically completed and marked as "implicit relationship".
[0068] Step d4 employs a cross-dimensional consistency verification mechanism, utilizing graph neural networks to detect whether there are logical conflicts between different dimensional classification systems.
[0069] Furthermore, during cross-dimensional consistency verification, we first identify dimensions that may have overlapping relationships within the multi-dimensional classification system, such as equipment asset (SB) and professional (ZY), knowledge type (LX) and business (YW). We then define logical conflict rules, such as a mismatch between the professional (ZY) to which the equipment belongs and the professional requirements of the business (YW), or a conflict between the knowledge type (LX) of the procedure and the organizational structure (JG) permissions of the issuing department. We then use a graph neural network to model the entity relationships under different dimensional labels. The conflict probability value of the output layer determines whether there is a logical contradiction. For example, when the SB dimension of a certain equipment entity is labeled as "transmission equipment," but the YW dimension of its associated business process is labeled as "substation business," the model will identify it as a high-probability conflict.
[0070] In step d5, if a logical conflict exists, mark the differing nodes and trigger the source verification process based on the labeled data, and update the dynamic knowledge graph.
[0071] Furthermore, upon discovering a logical conflict, "difference markers" are added to the entity nodes and relationship edges involved in the conflict in the graph, recording the conflict type and associated dimensional information. When source tracing verification is triggered, the corresponding classification annotation process records are searched in reverse through the dimensional codes in the labeled data to check for errors in feature extraction, category matching, and code generation. For example, if the dimensional conflict between device and business stems from an error in category matching during annotation, the dimensional labels of the corresponding entities are corrected, the relevant relationships in the graph are updated synchronously, the erroneous paths are deleted, the entity vector similarity is recalculated, and the correct relationship path is completed.
[0072] Step S106: Generate a processing report based on the power business scenario type and dynamic knowledge graph.
[0073] Specifically, step S106 includes:
[0074] Step e1: Determine the type of power business scenario. The above includes equipment operation and maintenance scenarios and policy compliance scenarios.
[0075] Furthermore, the system receives user-inputted business requirements or automatically identified business trigger commands, and makes judgments based on keyword matching and scenario feature comparison. For example, when the input contains keywords such as "equipment failure," "maintenance plan," or "operational status," it is determined to be an equipment maintenance scenario; when the input involves content such as "policy document," "compliance inspection," or "solution approval," it is determined to be a policy compliance scenario. For ambiguous requirements, a scenario selection list automatically pops up, allowing the user to confirm the specific type. After determining the scenario type, a unique identifier is assigned to the business scenario, and a corresponding scenario attribute tag is associated with it, such as labeling an equipment maintenance scenario with the "maintenance" tag and a policy compliance scenario with the "compliance" tag, providing clear scenario guidance for subsequent steps.
[0076] Step e2 involves calling the category system of business YW corresponding to the above-mentioned power business scenario types in the dynamic knowledge graph to determine the content framework and data extraction scope of the processing report.
[0077] Furthermore, based on the identified power business scenario types, the system automatically invokes the category system of the business YW dimension in the dynamic knowledge graph. For example, equipment operation and maintenance scenarios are associated with business categories such as "equipment repair," "fault handling," and "status assessment," while policy compliance scenarios are associated with business categories such as "policy matching," "compliance verification," and "solution review." According to these categories, the system retrieves preset report content framework templates. The framework for equipment operation and maintenance scenarios may include chapters such as "fault description," "diagnosis process," "handling solution," and "case reference," while the framework for policy compliance scenarios may include sections such as "solution overview," "policy clause matching," "compliance analysis," and "adjustment suggestions." Simultaneously, the data extraction scope is determined based on the business category. For example, equipment operation and maintenance scenarios are limited to extracting data related to equipment assets, fault cases, and operation and maintenance procedures, while policy compliance scenarios are limited to extracting data such as policy documents, solution texts, and compliance standards.
[0078] Step e3: Retrieve entity association data related to the above-mentioned power business scenario types from the dynamic knowledge graph. The above-mentioned entity association data includes equipment assets SB, historical fault cases, procedural documents in knowledge type LX, and technical standards in professional ZY, forming the basic dataset for the report.
[0079] Furthermore, when retrieving entity-related data from the dynamic knowledge graph, the search criteria are determined by the specific business scenario type, and the search is expanded in conjunction with the relationships between business categories. For equipment operation and maintenance scenarios, the unique identifier of the equipment entity is used to retrieve its corresponding equipment asset data (such as model parameters, maintenance records), historical fault cases (fault phenomena, handling processes, and results), operation and maintenance procedures (operation steps, safety specifications) in knowledge type LX, and technical standards (parameter thresholds, maintenance requirements) in professional category ZY. For policy compliance scenarios, policy documents related to the solution (issuing department, effective date, clause content), compliance specifications in knowledge type LX, industry standards in professional category ZY, and implementation requirements of relevant organizations JG are retrieved. All retrieved data is integrated according to entity relationships, and duplicate or irrelevant information is removed to form a structured basic dataset for reporting. Each data entry is labeled with the identifier of the source entity and its relationship.
[0080] Step e4: Graph neural networks are used to perform in-depth analysis on the above entity association data. Combined with generative large model retrieval enhancement generation technology, structured analysis content including scene adaptation conclusions, source tracing, and optimization suggestions is generated.
[0081] Furthermore, when using graph neural networks for deep analysis of entity-related data, the entities and relationships in the report's basic dataset are first transformed into a graph structure. Entities are then mapped to low-dimensional vectors using node embedding technology. Vector similarity is calculated to uncover hidden associations, such as the potential causal relationship between "abnormal equipment parameters" and "historical fault causes" in equipment operation and maintenance scenarios, and the matching degree between "solution clauses" and "policy requirements" in policy compliance scenarios. After analysis, the structured data is input into a generative large-scale model, and retrieval-enhanced generation technology is enabled to retrieve relevant procedural documents and technical standards from the knowledge graph in real time as the basis for generation. The generated structured analysis content must include scenario-adaptation conclusions (such as root causes of equipment failures and solution compliance determination), source tracing (specific procedural clauses and policy provisions cited), and optimization suggestions (such as adjustments to equipment maintenance cycles and directions for modifying solution clauses). Each conclusion must be supported by at least one entity data point.
[0082] Step e5: In accordance with the output specifications of the above power business scenario types, integrate the above structured analysis content into a processing report. The processing report output for the equipment operation and maintenance scenario includes a fault diagnosis report and historical case matching results, while the processing report output for the policy compliance scenario includes a solution and policy compatibility report.
[0083] Furthermore, when integrating reports according to the output specifications of power business scenario types, the corresponding report format template is first called. The template for equipment operation and maintenance scenarios must meet the standardized document requirements for operation and maintenance work, including fixed modules such as "Fault Phenomenon Description," "Diagnosis Basis," "Handling Steps," and "Effect Evaluation" for fault diagnosis reports, as well as sections such as "Similar Case List," "Case Comparison Analysis," and "Experience Learning" for historical case matching results. The template for policy compliance scenarios must follow the standardized format of compliance reports, including content such as "Policy Clause Comparison Table," "Explanation of Non-Compliance Items," "Rectification Measures," and "Compliance Rating" for solution and policy compatibility reports. The structured analysis content generated in step e4 is then filled into the template by module to ensure the correspondence between data and modules. Hyperlinks or reference numbers are added to referenced procedural documents and technical standards to facilitate traceability. Finally, the format is automatically validated, and the font, layout, and chapter numbering are adjusted to generate a processing report that conforms to the output specifications.
[0084] In some optional embodiments, the above-mentioned classification and annotation of standardized structured knowledge data based on a multi-dimensional classification system to obtain labeled data further includes:
[0085] Step f1: Based on the annotation information of knowledge type LX, business YW, professional ZY, equipment asset SB, and organization JG in the labeled data, add dimensional attribute labels to the core entities.
[0086] Furthermore, specific annotation information for knowledge type LX, business YW, professional ZY, equipment asset SB, and organizational structure JG is extracted from the labeled data. This information is presented in the format of "dimensional code + category code". Then, the correspondence between core entities and the annotation information of each dimension is determined. For example, the equipment entity corresponds to the annotation information of equipment asset SB, and the procedure entity corresponds to the annotation information of knowledge type LX. Next, dimension attribute tags are generated for each core entity according to the correspondence. The tag format is "dimensional name + category name", such as "equipment asset SB: transformer" and "knowledge type LX: maintenance procedure". For core entities that are associated with multiple dimension annotation information, all relevant dimension attribute tags need to be added to form a multi-dimensional attribute set for the entity. These dimension attribute tags are bound and stored with the core entity in the knowledge graph to ensure that the associated multi-dimensional attributes can be directly accessed when the entity is retrieved or analyzed.
[0087] Step f2 involves using a time-series decay algorithm to adjust the weights of the relationships between core entities based on the weight adjustment rules. The real-time data weight of equipment asset SB increases dynamically over time, and the weight adjustment rules are related to the process node requirements of business YW.
[0088] Furthermore, the weight adjustment rules set the initial weight benchmark and adjustment range for the relationships based on different process nodes of the business YW. A time-series decay algorithm is applied, using the establishment time or last update time of the relationship as a benchmark to dynamically adjust the weights of the relationships between core entities: for relationships involving non-real-time data, the weight is gradually reduced over time according to the decay coefficient; for relationships involving real-time data of equipment assets SB, the weight is increased according to a time-positive increment formula to ensure that the latest real-time data relationships have higher priority. Simultaneously, based on the current business YW process node, the corresponding adjustment rule parameters are invoked. For example, at the fault handling node, the weight increment of real-time data relationships will be greater than that at the daily inspection node. Finally, the weight calculation process is periodically triggered to update the weight values of all relationships according to the above rules and algorithms, and the time and basis of the weight adjustment are recorded in the knowledge graph for subsequent traceability.
[0089] The power knowledge processing method based on a layered architecture provided in this invention firstly acquires multi-source heterogeneous power knowledge data, including structured equipment ledgers and business forms, unstructured policy documents and technical literature, and real-time equipment operation monitoring data, achieving comprehensive coverage of knowledge data and providing a complete data source for subsequent in-depth processing. By integrating multiple types of knowledge, it avoids analytical bias caused by data dispersion, ensuring mastery of knowledge across all scenarios of power system operation and management. Furthermore, by establishing a foundation for the correlation of multi-source data, it creates conditions for mining potential connections between different types of data. Secondly, by unifying and verifying the fields of structured data, it solves the problems of conflicting formats and inconsistent parameters in structured data from different sources, improving data consistency and reliability. Through parsing, extraction, and semantic cleaning of unstructured data, it transforms text and documents that are difficult to use directly into standardized, processable formats, releasing the value of unstructured knowledge. Through edge preprocessing and standardized encapsulation of real-time streaming data, it filters noise and retains key features, ensuring the effectiveness and timeliness of real-time data. Finally, through targeted processing of the three types of data, it lays a unified data foundation for subsequent classification, labeling, and knowledge fusion, reducing processing obstacles caused by data format differences. Then, by combining primary and secondary dimensions in its classification approach, the limitations of single-dimensional classification are overcome, enabling precise characterization of knowledge attributes from multiple perspectives, including knowledge type, business domain, professional scope, equipment association, and organizational affiliation. Intelligent semantic parsing and feature extraction improve the accuracy and efficiency of classification and labeling, reducing the subjectivity of manual intervention. Hierarchical category matching and encoding generation give knowledge a clear organizational structure, facilitating rapid retrieval and location. Dynamic filtering and priority sorting ensure that classification and labeling can adapt to the needs of different business scenarios, enhancing the flexibility of knowledge management. By forming standardized labeled data, a structured classification basis is provided for knowledge graph construction, promoting the systematic organization of knowledge. Subsequently, by accurately identifying core entities such as equipment, faults, and procedures, the problem of extracting key information from unstructured text was solved, and the core elements of knowledge were clarified. By parsing the semantic relationships between entities, the logical connections within knowledge were revealed, transforming scattered textual information into structured "entity-relationship" units. By associating multi-dimensional classification information, triplet data and classification systems were linked, enhancing the traceability and relevance of knowledge. Through the structured representation of triples, basic data was provided for establishing entity relationships in knowledge graphs, improving the reusability and reasoning ability of knowledge.Furthermore, by building a framework around equipment, faults, and procedures as core entities, the key knowledge elements of the power system are focused on, ensuring the domain relevance and practicality of the graph. By integrating triples, labeled data, and real-time streaming data, the organic combination of static knowledge and dynamic information is achieved, enriching the content dimensions of the graph. By mining implicit relationships between entities and completing paths, the correlation between knowledge is enhanced, supporting deeper intelligent reasoning. Through cross-dimensional consistency verification and dynamic update mechanisms, knowledge conflicts are promptly discovered and corrected, ensuring the accuracy and timeliness of the graph. Through the dynamic characteristics of the graph, it can adapt to changes in the power system, continuously providing reliable knowledge support for business operations. Finally, by customizing report frameworks for different scenarios such as equipment operation and maintenance and policy compliance, the report content is ensured to accurately match business needs, thus improving the report's practicality. Through in-depth analysis of entity-related data and combined with generative technologies, complex knowledge is transformed into easily understandable structured conclusions, lowering the barrier to entry for professional knowledge. By integrating multi-dimensional knowledge and annotating the sources of evidence, the credibility and interpretability of the report conclusions are enhanced. By outputting customized reports, the decision-making needs of different business scenarios are met, improving work efficiency. Through standardized report presentation, the effective transfer and application of knowledge are promoted, providing strong support for scientific decision-making in the power industry. In summary, by implementing this invention, the significant shortcomings of related technologies in processing multi-source heterogeneous power knowledge data, classification and annotation, knowledge graph construction, and report generation are addressed, failing to meet the power industry's needs for efficient knowledge management and utilization.
[0090] This embodiment also provides a power knowledge processing device based on a hierarchical architecture, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0091] This embodiment provides a power knowledge processing device based on a hierarchical architecture, such as... Figure 2 As shown, it includes:
[0092] The acquisition module 201 is used to acquire multi-source heterogeneous power knowledge data, which includes structured data, unstructured data and real-time streaming data.
[0093] The standardization module 202 is used to perform standardization processing on the structured data, unstructured data and real-time streaming data mentioned above, respectively.
[0094] The annotation module 203 is used to classify and annotate the standardized structured knowledge data based on a multi-dimensional classification system to obtain annotated data. The multi-dimensional classification system includes a main dimension and an auxiliary dimension.
[0095] The transformation module 204 is used to transform the standardized unstructured text into triple data;
[0096] Module 205 is used to construct a dynamic knowledge graph based on triplet data, labeled data and standardized real-time streaming data. The knowledge graph has equipment, faults and procedures as core entities.
[0097] The generation module 206 is used to generate processing reports based on power business scenario types and dynamic knowledge graphs.
[0098] The power knowledge processing device based on a layered architecture provided in this invention firstly acquires multi-source heterogeneous power knowledge data, including structured equipment ledgers and business forms, unstructured policy documents and technical literature, and real-time equipment operation monitoring data, achieving comprehensive coverage of knowledge data and providing a complete data source for subsequent in-depth processing. By integrating multiple types of knowledge, it avoids analytical bias caused by data dispersion, ensuring mastery of knowledge across all scenarios of power system operation and management. Furthermore, by establishing a foundation for the correlation of multi-source data, it creates conditions for mining potential connections between different types of data. Secondly, by unifying and verifying the fields of structured data, it solves the problems of conflicting formats and inconsistent parameters in structured data from different sources, improving data consistency and reliability. Through parsing, extraction, and semantic cleaning of unstructured data, it transforms text and documents that are difficult to use directly into standardized, processable formats, releasing the value of unstructured knowledge. Through edge preprocessing and standardized encapsulation of real-time streaming data, it filters noise and retains key features, ensuring the effectiveness and timeliness of real-time data. Finally, through targeted processing of the three types of data, it lays a unified data foundation for subsequent classification, labeling, and knowledge fusion, reducing processing obstacles caused by differences in data formats. Then, by combining primary and secondary dimensions in its classification approach, the limitations of single-dimensional classification are overcome, enabling precise characterization of knowledge attributes from multiple perspectives, including knowledge type, business domain, professional scope, equipment association, and organizational affiliation. Intelligent semantic parsing and feature extraction improve the accuracy and efficiency of classification and labeling, reducing the subjectivity of manual intervention. Hierarchical category matching and encoding generation give knowledge a clear organizational structure, facilitating rapid retrieval and location. Dynamic filtering and priority sorting ensure that classification and labeling can adapt to the needs of different business scenarios, enhancing the flexibility of knowledge management. By forming standardized labeled data, a structured classification basis is provided for knowledge graph construction, promoting the systematic organization of knowledge. Subsequently, by accurately identifying core entities such as equipment, faults, and procedures, the problem of extracting key information from unstructured text was solved, and the core elements of knowledge were clarified. By parsing the semantic relationships between entities, the logical connections within knowledge were revealed, transforming scattered textual information into structured "entity-relationship" units. By associating multi-dimensional classification information, triplet data and classification systems were linked, enhancing the traceability and relevance of knowledge. Through the structured representation of triples, basic data was provided for establishing entity relationships in knowledge graphs, improving the reusability and reasoning ability of knowledge.Furthermore, by building a framework around equipment, faults, and procedures as core entities, the key knowledge elements of the power system are focused on, ensuring the domain relevance and practicality of the graph. By integrating triples, labeled data, and real-time streaming data, the organic combination of static knowledge and dynamic information is achieved, enriching the content dimensions of the graph. By mining implicit relationships between entities and completing paths, the correlation between knowledge is enhanced, supporting deeper intelligent reasoning. Through cross-dimensional consistency verification and dynamic update mechanisms, knowledge conflicts are promptly discovered and corrected, ensuring the accuracy and timeliness of the graph. Through the dynamic characteristics of the graph, it can adapt to changes in the power system, continuously providing reliable knowledge support for business operations. Finally, by customizing report frameworks for different scenarios such as equipment operation and maintenance and policy compliance, the report content is ensured to accurately match business needs, thus improving the report's practicality. Through in-depth analysis of entity-related data and combined with generative technologies, complex knowledge is transformed into easily understandable structured conclusions, lowering the barrier to entry for professional knowledge. By integrating multi-dimensional knowledge and annotating the sources of evidence, the credibility and interpretability of the report conclusions are enhanced. By outputting customized reports, the decision-making needs of different business scenarios are met, improving work efficiency. Through standardized report presentation, the effective transfer and application of knowledge are promoted, providing strong support for scientific decision-making in the power industry. In summary, by implementing this invention, the significant shortcomings of related technologies in processing multi-source heterogeneous power knowledge data, classification and annotation, knowledge graph construction, and report generation are addressed, failing to meet the power industry's needs for efficient knowledge management and utilization.
[0099] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0100] In this embodiment, the verification device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0101] This invention also provides a computer device having the above-described features. Figure 2 The apparatus shown.
[0102] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 3As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 3 Take a processor 10 as an example.
[0103] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.
[0104] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0105] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0106] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0107] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0108] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and stored on a remote storage medium or a non-transitory machine-readable storage medium and to be stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0109] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A power knowledge processing method based on a hierarchical architecture, characterized in that, The method includes: Acquire multi-source heterogeneous power knowledge data, which includes structured data, unstructured data, and real-time streaming data; Standardization processing is performed on the structured data, unstructured data, and real-time streaming data respectively; The standardized structured knowledge data is classified and labeled based on a multi-dimensional classification system to obtain labeled data. The multi-dimensional classification system includes a main dimension and an auxiliary dimension. Convert standardized unstructured text into triplet data; A dynamic knowledge graph is constructed based on triplet data, labeled data, and standardized real-time streaming data, with equipment, faults, and procedures as core entities. Generate processing reports based on power business scenario types and dynamic knowledge graphs; The structured data includes ledger information at the equipment asset level and process form data at the business level; the unstructured data includes policy documents at the knowledge type level and technical documents at the professional level; and the real-time streaming data includes equipment operation status monitoring data. The standardization processing performed on the structured data, unstructured data, and real-time streaming data includes: The structured data is mapped to fields in a unified manner based on a preset metadata standard, and the structured data is verified for compliance by associating it with the corresponding professional technical standards. The unstructured data is parsed and extracted using multimodal documents, and then converted into a parsable format and semantically cleaned. The real-time streaming data is subjected to noise filtering and feature parameter extraction through edge nodes, and the processed real-time streaming data is standardized and encapsulated based on time series and device identifier. The primary dimensions include knowledge type (LX), business (YW), and professional (ZY), while the secondary dimensions include equipment assets (SB) and organizational structure (JG). The standardized structured knowledge data is classified and labeled based on a multi-dimensional classification system to obtain labeled data, including: A pre-trained large model in the power sector is used to perform semantic parsing on structured knowledge data, extracting business theme features, technical field features, and core object features from the structured knowledge data; Based on business theme features, technical field features, and core object features, knowledge type LX, business YW, and professional ZY are matched, and the association is forced and sorted according to the order from knowledge type LX to business YW to professional ZY. The system dynamically filters equipment assets (SB) and organizational structures (JG) based on knowledge attributes, and sorts them based on the priority of relevance. The knowledge attributes are determined by business scenario characteristics, business theme characteristics, technical field characteristics, and core object characteristics. For the sorted main dimension, the hierarchical category knowledge graph under the main dimension is called. By calculating the similarity between feature words and category attributes, the matching is performed level by level down to the finest category, and the annotation is performed to obtain the main dimension annotation data. For the sorted secondary dimension, the hierarchical category knowledge graph under the secondary dimension is called. By calculating the similarity between feature words and category attributes, the data is matched step by step to the finest-grained category and labeled to obtain the secondary dimension labeled data. Annotated data is generated based on the primary dimension annotation data and the secondary dimension annotation data.
2. The method according to claim 1, characterized in that, The process of converting standardized unstructured text into triplet data includes: Natural language processing technology is used to segment and identify entities in unstructured text to obtain core entities, which include equipment, faults, procedures and business activities. The entity recognition is based on a pre-trained model in the power field. The semantic relation extraction algorithm is used to analyze the semantic correlation between the core entities and determine the relationship type between the subject entity and the object entity; Triples are constructed based on the format of subject entity, relation type to object entity, where the subject entity, relation type and object entity are all associated with dimensional information in a multi-dimensional classification system.
3. The method according to claim 2, characterized in that, The construction of a dynamic knowledge graph based on triplet data, labeled data, and standardized real-time streaming data includes: The basic framework of the graph is constructed with equipment entities, fault entities, and procedure entities as core entities. The initial association between core entities is established through triple data. Specifically, the equipment entity is associated with the labeled data of equipment asset SB, and the procedure entity is associated with the labeled data of knowledge type LX. Time-series features are extracted from the standardized real-time streaming data, and the real-time streaming data is associated with the corresponding core entities based on the timestamp and device identifier. A graph neural network algorithm is used to mine the relationships between core entities, and entity vector similarity is used to calculate and complete the implicit association paths of equipment association fault association procedures; A cross-dimensional consistency verification mechanism is adopted, and a graph neural network is used to detect whether there are logical conflicts between different dimensional classification systems. If a logical conflict exists, mark the differing nodes and trigger the source tracing and verification process based on the labeled data, and update the dynamic knowledge graph.
4. The method according to claim 3, characterized in that, The process report, generated based on power business scenario types and dynamic knowledge graphs, includes: Determine the types of power business scenarios, which include equipment operation and maintenance scenarios and policy compliance scenarios; The category system of business YW corresponding to the power business scenario type in the dynamic knowledge graph is invoked to determine the content framework and data extraction scope of the processing report; Retrieve entity association data related to the power business scenario type from the dynamic knowledge graph. The entity association data includes equipment assets (SB), historical fault cases, procedural documents in knowledge type LX, and technical standards in professional ZY to form the basic dataset for the report. The entity association data is analyzed in depth using graph neural networks, and combined with the retrieval enhancement generation technology of generative large models, to generate structured analysis content that includes scene adaptation conclusions, source tracing, and optimization suggestions. According to the output specifications of the power business scenario types, the structured analysis content is integrated into a processing report. The processing report output for the equipment operation and maintenance scenario includes a fault diagnosis report and historical case matching results, while the processing report output for the policy compliance scenario includes a solution and policy compatibility report.
5. The method according to claim 4, characterized in that, The process of classifying and labeling standardized structured knowledge data based on a multi-dimensional classification system to obtain labeled data also includes: Based on the annotation information of knowledge type LX, business YW, professional ZY, equipment asset SB and organization JG in the annotated data, add dimensional attribute tags to the core entities; The weight adjustment rules are based on a time-series decay algorithm to adjust the weights of the relationships between core entities. The real-time data weight of equipment assets (SB) increases dynamically over time, and the weight adjustment rules are related to the process node requirements of business (YW).
6. A power knowledge processing device based on a hierarchical architecture, characterized in that, The device includes: The acquisition module is used to acquire multi-source heterogeneous power knowledge data, which includes structured data, unstructured data, and real-time streaming data. The standardization module is used to perform standardization processing on the structured data, unstructured data, and real-time streaming data respectively; The annotation module is used to classify and annotate the standardized structured knowledge data based on a multi-dimensional classification system to obtain annotated data. The multi-dimensional classification system includes a main dimension and an auxiliary dimension. The transformation module is used to convert standardized unstructured text into triplet data; The construction module is used to build a dynamic knowledge graph based on triple data, labeled data and standardized real-time streaming data. The knowledge graph has equipment, faults and procedures as core entities. The generation module is used to generate processing reports based on power business scenario types and dynamic knowledge graphs; The structured data includes ledger information at the equipment asset level and process form data at the business level; the unstructured data includes policy documents at the knowledge type level and technical documents at the professional level; the real-time streaming data includes equipment operation status monitoring data; and the standardization module is also used to perform the following steps: The structured data is mapped uniformly based on a preset metadata standard, and the structured data is verified for compliance by associating it with corresponding professional technical standards. Multimodal documents are used to parse and extract the unstructured data, converting it into a parsable format and performing semantic cleaning. Noise filtering and feature parameter extraction are performed on the real-time streaming data through edge nodes, and the processed real-time streaming data is standardized and encapsulated based on time series and device identifiers. The main dimensions include knowledge type (LX), business (YW), and professional (ZY), and the secondary dimensions include equipment asset (SB) and organizational structure (JG). The standardized structured knowledge data is classified and labeled based on a multi-dimensional classification system to obtain labeled data, including: using a pre-trained large model in the power industry to perform semantic parsing on the structured knowledge data, extracting business theme features, technical field features, and core object features of the structured knowledge data; based on business theme features... The system matches knowledge types LX, business YW, and professional ZY based on technical field characteristics and core object characteristics, and forces association sorting based on the order of knowledge type LX to business YW to professional ZY. It dynamically filters equipment assets SB and organizational structures JG based on knowledge attributes, and sorts them based on association priority. These knowledge attributes are determined by business scenario characteristics, business theme characteristics, technical field characteristics, and core object characteristics. For the sorted main dimension, the system calls the hierarchical category knowledge graph under the main dimension, calculates the similarity between feature words and category attributes, matches step-by-step to the finest-grained category, and annotates it to obtain the main dimension annotated data. For the sorted secondary dimension, the system calls the hierarchical category knowledge graph under the secondary dimension, calculates the similarity between feature words and category attributes, matches step-by-step to the finest-grained category, and annotates it to obtain the secondary dimension annotated data. Annotated data is formed based on the main dimension annotated data and the secondary dimension annotated data.
7. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the power knowledge processing method based on a hierarchical architecture as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the power knowledge processing method based on a hierarchical architecture as described in any one of claims 1 to 5.