Unified modeling and archiving method for multi-source data in power generation process and related equipment

By constructing a unified knowledge graph based on a three-layer semantic model and employing a differential update strategy, the problem of dispersed storage of multi-source heterogeneous data in power generation enterprises was solved, enabling efficient data management and intelligent diagnosis, and improving the level of intelligent operation and maintenance of power generation equipment.

CN121919280APending Publication Date: 2026-04-24HUANENG JINGMEN THERMAL POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUANENG JINGMEN THERMAL POWER CO LTD
Filing Date
2026-01-05
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, the decentralized storage of multi-source heterogeneous data in power generation enterprises leads to high management costs, and it is difficult to effectively link different types of data. Furthermore, the lack of a historical version tracing mechanism makes it impossible to meet the needs of in-depth data application.

Method used

By using semantic recognition and entity extraction from multi-source heterogeneous data, a unified knowledge graph with a three-layer semantic model is constructed. A differential update strategy and vector retrieval technology are adopted to achieve dynamic data fusion and intelligent diagnosis.

Benefits of technology

It enables efficient management and intelligent operation and maintenance of multi-source heterogeneous data, improves data fusion efficiency and knowledge retrieval accuracy, and supports equipment status monitoring and fault prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121919280A_ABST
    Figure CN121919280A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power generation data management, in particular to a power generation process multi-source data unified modeling and archiving method and related equipment, and through semantic recognition and entity extraction, structured data is mapped into parameter nodes, unstructured texts are mapped into knowledge nodes, and pictures and videos are mapped into multi-modal external entity nodes. A unified knowledge graph is constructed by adopting a three-layer semantic model of an equipment layer, a working condition layer and a knowledge layer, and a physical entity hierarchical structure, a dynamic operation attribute and an expert experience rule are respectively defined. And dynamic evolution management of the knowledge graph is realized through a differential updating strategy, and a total change history is recorded and is audited and graded by experts to take effect. Fault reason reasoning, operation state operation knowledge recommendation and data association retrieval functions are realized based on a vector retrieval technology, and intelligent decision and knowledge reuse in an industrial production scene are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power generation data management technology, specifically to a method and related equipment for unified modeling and archiving of multi-source data in the power generation process. Background Technology

[0002] In the power generation industry, especially in thermal power plants, existing technologies primarily employ a collaborative storage solution combining structured databases, time-series databases, and document libraries to address the management needs of multi-source heterogeneous data. Specifically, structured data such as real-time operating parameters are typically stored in time-series databases to support time-series analysis; semi-structured data such as inspection records and maintenance work orders are mostly stored in structured databases or document management systems; unstructured text data such as procedures, fault reports, and expert experience are stored as documents in document libraries; and multimedia data such as images and videos are managed through dedicated databases or document libraries. This distributed storage model has become the mainstream technology for data management in power generation companies, widely applied in various scenarios such as equipment status monitoring, operating parameter analysis, and maintenance work order management, forming an industry-recognized solution framework.

[0003] However, the current problems are: the dispersed storage of data sources in existing technical solutions leads to high management costs, and different types of data (such as structured operating parameters and unstructured expert experience) are difficult to effectively correlate, forming information silos; the frequently updated power generation process data lacks a historical version traceability mechanism, making it difficult to achieve complete tracking of the data change process; more importantly, existing technologies only focus on data storage and query functions, lacking reasoning capabilities and intelligent diagnostic support based on multi-source data fusion, and cannot fully explore the value of data to support in-depth application needs such as equipment status monitoring and fault prediction, becoming a key technical bottleneck restricting the intelligent upgrading of the power generation process. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a unified modeling and archiving method and related equipment for multi-source data in the power generation process, which addresses the shortcomings of the prior art and solves the technical problem of high management costs caused by the dispersed sources and heterogeneous formats of multi-source heterogeneous data generated in the current power generation process.

[0005] The objective of this invention is achieved through the following technical solutions: In a first aspect, the present invention provides a method for unified modeling and archiving of multi-source data in a power generation process, including: Acquire multi-source heterogeneous production data, perform semantic recognition and entity extraction on the multi-source heterogeneous production data to obtain structured data, unstructured text, images and videos; map the structured data in the multi-source heterogeneous production data to parameter nodes, map the unstructured text to knowledge nodes, and map the images and videos to multimodal external entity nodes. A unified knowledge graph is constructed based on a three-layer semantic model, which includes a device layer, an operating condition layer, and a knowledge layer. The device layer is used to define the hierarchical structure of physical entities, the operating condition layer describes the dynamic operating attributes of power generation equipment, and the knowledge layer contains expert experience and equipment operation rules. A differential update strategy is adopted to manage the evolution of the knowledge graph, record the full history of changes, and update the knowledge through a process of expert review and hierarchical implementation. It is also used to query the knowledge graph based on vector retrieval technology to realize fault cause reasoning, operation status knowledge recommendation, and data association retrieval.

[0006] As a further improvement of the present invention, semantic recognition and entity extraction are performed on multi-source heterogeneous production data, including: Multi-source heterogeneous production data is parsed and classified based on the power generation equipment's identification OID, process tag number, and business tag. The named entity recognition technology based on large models is used to extract key information such as equipment name, operating condition characteristics, fault phenomena and handling measures from unstructured text. Establish a unified indexing mechanism for multimodal external entity nodes to enable semantic association of cross-modal data.

[0007] As a further improvement of the present invention, the multi-source heterogeneous production data further includes semantic alignment processing, including: For structured data, statistical feature vectors of real-time device parameters are extracted, including mean, variance, spectral features, and time-frequency analysis results, to construct multi-dimensional feature representations of parameter nodes; For unstructured text, a domain-adapted pre-trained language model is used, combined with an electric power professional dictionary, to perform named entity recognition and relation extraction, extracting a quadruple of equipment name, fault mode, phenomenon description, and handling measures. For multimedia data, a multimodal feature extraction network is applied to generate semantic embedding vectors for images or videos and establish cross-modal associations with device entities; Based on spatiotemporal consistency constraints, the timestamps, geographical locations, and process locations of multi-source data are aligned and verified to eliminate data asynchrony and spatial offset.

[0008] As a further improvement of the present invention, the specific steps for constructing the three-layer semantic model include: In the equipment layer, the hierarchical topology of the generator set, system and components of the power generation equipment is constructed through the first set relationship, forming a physical entity hierarchical structure; In the operating condition layer, the operating parameters, status, and cycle of the power generation equipment are associated with the equipment layer through the second setting relationship; In the knowledge layer, a semantic network of fault states, handling procedures, and historical cases is constructed, and it interacts with the operating condition layer and the equipment layer through a third-party defined relationship. The equipment layer, operating condition layer, and knowledge layer are connected through bidirectional semantic relationships.

[0009] As a further improvement of the present invention, the differential update strategy specifically includes: Store the differences in the knowledge graph after the changes, establish a version identifier for each knowledge node, and record the change time, the person making the change, and the reason for the change. The knowledge validity level is set, which is used to classify knowledge into four states: draft, pending review, effective, and expired. Expired knowledge is logically archived rather than physically deleted, preserving a complete historical traceability chain.

[0010] As a further improvement to the present invention, the multi-expert review and tiered effectiveness process includes: The roles of audit experts in different professional fields are set up, including equipment experts, process experts and safety experts; Based on the type and scope of the knowledge update, the system will automatically assign appropriate expert roles for review. Establish a priority assessment mechanism for knowledge updates, and differentiate the review process between urgent and routine updates; conduct full-process traceability based on the electronic records of review opinions and the basis for knowledge updates in real time.

[0011] As a further improvement of the present invention, the steps of the vector retrieval technology for fault diagnosis include: The real-time running parameters are vectorized and similarity is calculated with historical similar working conditions; Based on graph neural networks, relational reasoning is performed on knowledge graphs to identify potential fault propagation paths; Based on the severity of the fault and the historical handling results, the recommended handling measures are dynamically ranked and recommended. Generates a visual diagnostic report containing a complete knowledge chain of fault causes, impact scope, and handling recommendations.

[0012] Secondly, the present invention provides a unified modeling and archiving system for multi-source data in the power generation process, comprising: The multi-source data access module is used to connect to real-time monitoring systems, equipment management systems, document management systems, and multimedia acquisition devices during the power generation process to acquire multi-source heterogeneous production data. A semantic recognition and entity extraction engine is used for automatic classification, entity recognition, and relation extraction of multi-source heterogeneous production data; A three-layer knowledge graph builder is used to construct a unified knowledge model that includes an equipment layer, a working condition layer, and a knowledge layer. The knowledge graph evolution manager is used to perform differential updates, version control, and review processes for the knowledge graph. The intelligent query and reasoning module is used to realize knowledge query, fault diagnosis and decision recommendation based on vector retrieval.

[0013] Thirdly, the present invention provides a computer-readable storage medium for storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the above-described method for unified modeling and archiving of multi-source data in the power generation process.

[0014] Fourthly, the present invention provides a computing device, comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include steps for performing the above-described method for unified modeling and archiving of multi-source data in the power generation process.

[0015] The beneficial effects of this invention are as follows: This invention provides a unified modeling and archiving method for multi-source data in the power generation process. By acquiring multi-source heterogeneous production data and performing semantic recognition and entity extraction, structured data is mapped to parameter nodes, unstructured text is mapped to knowledge nodes, and images and videos are mapped to multimodal external entity nodes. This achieves accurate conversion of multi-source heterogeneous data into structured knowledge nodes, improving data fusion efficiency and interpretability. Based on a three-layer semantic model—defining the physical entity hierarchy at the equipment layer, describing dynamic operating attributes at the operating condition layer, and integrating expert experience and operating rules at the knowledge layer—a unified knowledge graph is constructed, forming a hierarchical and attribute-based knowledge representation system, enhancing the understanding of power generation... The system provides a comprehensive depiction of equipment operating status; it employs a differential update strategy to manage the evolution of the knowledge graph, recording the entire change history and updating knowledge through expert review and a tiered effectiveness process to ensure the timeliness and accuracy of the knowledge graph, while also enabling traceability and controllability of the change process; it uses vector retrieval technology to query the knowledge graph, enabling fault cause reasoning, operational status knowledge recommendation, and data association retrieval, thus improving the efficiency and accuracy of knowledge retrieval. Ultimately, through the synergistic effect of the above technical features, a complete technical system of data fusion, knowledge modeling, dynamic updating, and intelligent retrieval is formed, significantly improving the intelligence level and decision support capabilities of power generation equipment operation and maintenance knowledge management. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1This is a schematic diagram of the process for unified modeling and archiving of multi-source data in the power generation process in an embodiment of the present invention; Figure 2 This is an internal structural diagram of the computer device in an embodiment of the present invention; Detailed Implementation To make the objectives and technical solutions of this invention clearer and easier to understand, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0018] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. The described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] Example 1 Currently, the multi-source heterogeneous data generated during the power generation process is scattered and has different formats, resulting in high management costs and difficulty in associating different types of equipment information. At the same time, the data is updated frequently but lacks a historical version traceability mechanism, making it impossible for the data to drive reasoning and intelligent diagnosis. Therefore, this embodiment provides a unified modeling and archiving method for multi-source data in the power generation process, realizing dynamic fusion of multiple data sources, correlation reasoning, and full life cycle management.

[0020] A unified modeling and archiving method for multi-source data in the power generation process includes the following steps: acquiring multi-source heterogeneous production data; performing semantic recognition and entity extraction on the multi-source heterogeneous production data; mapping structured data in the multi-source heterogeneous production data to parameter nodes, unstructured text to knowledge nodes, and images and videos to multimodal external entity nodes; constructing a unified knowledge graph based on a three-layer semantic model, which includes an equipment layer, an operating condition layer, and a knowledge layer. The equipment layer defines the hierarchical structure of physical entities, the operating condition layer describes the dynamic operating attributes of power generation equipment, and the knowledge layer contains expert experience and equipment operation rules; using a differential update strategy to manage the evolution of the knowledge graph, recording the full change history, and updating knowledge through expert review and a graded effectiveness process; and querying the knowledge graph based on vector retrieval technology to achieve fault cause reasoning, operating status operation knowledge recommendation, and data association retrieval.

[0021] The working principle of this embodiment is as follows: Multi-source heterogeneous production data from the power generation process is acquired through a standardized interface. After semantic recognition and entity extraction, different types of data are mapped to corresponding nodes. Then, a unified knowledge graph is constructed based on the equipment layer, operating condition layer, and knowledge layer, clarifying the semantic relationships between nodes at each level. When knowledge needs to be updated, a differential update strategy is used to record change information, and the update is completed after expert review and a tiered implementation process. Finally, based on vector retrieval technology, relevant nodes are quickly retrieved in the knowledge graph according to user query requests, enabling functions such as fault cause reasoning. This embodiment achieves dynamic fusion of multiple data sources, solving the problems of scattered data sources and heterogeneous formats, and reducing management costs. Through the association relationships of the knowledge graph, effective association of different types of equipment information is achieved. The differential update strategy and version management mechanism make the data update process traceable, meeting the needs of historical version tracing. The reasoning function based on vector retrieval realizes data-driven reasoning and intelligent diagnosis, improving the intelligent operation and maintenance level of power generation enterprises.

[0022] It should be noted that the multi-source heterogeneous production data in this embodiment comes from the real-time monitoring system, equipment management system, document management system and multimedia acquisition equipment in the power generation process. Specifically, it includes real-time operating parameters (structured), inspection records, maintenance work orders (semi-structured), procedures and regulations, fault reports, expert experience (unstructured), and multimedia data such as pictures and videos (unstructured).

[0023] Semantic recognition and entity extraction are performed on multi-source heterogeneous production data, including: parsing and classifying multi-source heterogeneous production data based on the identification ID (OID) of power generation equipment, process tag number, and business tag; using named entity recognition technology based on large models to extract key information such as equipment name, operating condition characteristics, fault phenomena, and handling measures from unstructured text; and establishing a unified indexing mechanism for multimodal external entity nodes to achieve semantic association of cross-modal data.

[0024] The method adopts named entity recognition technology based on a large model. The large model selected is the BERT model that has been fine-tuned with data from the power industry. The advantage of choosing this model is that it has a high accuracy rate in entity recognition in the professional field and can accurately extract key information such as equipment name, operating condition characteristics, fault phenomena and handling measures. In practical applications, other models such as GPT series can also be selected. This embodiment does not limit it.

[0025] In the three-layer semantic model of this embodiment, the equipment layer is used to define the hierarchical structure of physical entities, including units, systems, components, etc. For example, Unit #1 of a thermal power plant is a first-level node, the boiler system and turbine system under the unit are second-level nodes, and the furnace, economizer and other components under the boiler system are third-level nodes. Each node constructs a hierarchical topology structure through relationships such as "deployed in" and "belongs to".

[0026] The operating condition layer is used to describe the dynamic operating attributes of power generation equipment, including parameters (such as temperature, pressure, speed, etc.), status (normal, abnormal, alarm, etc.), and cycles (operation cycle, maintenance cycle, etc.). It is associated with the equipment layer through relationships such as "influence" and "historical correlation". For example, the speed parameters of the steam turbine (operating condition layer node) are associated with the steam turbine components (equipment layer node) through the "belong to" relationship.

[0027] The knowledge layer contains expert experience and equipment operation rules, including fault status, handling procedures, historical cases, etc. It interacts with the operating condition layer and equipment layer through relationships such as "cause" and "recommended handling measures". For example, a fault status (knowledge layer node) is associated with the corresponding equipment component (equipment layer node) and operating condition parameters (operating condition layer node) through the "cause" relationship.

[0028] This embodiment sets up review expert roles from different professional fields, such as equipment experts, process experts, and safety experts. Based on the type of knowledge update (e.g., parameter adjustment, addition of fault handling rules) and the scope of impact (e.g., single piece of equipment, the entire system, the whole plant), the corresponding review expert group is automatically assigned. For example, knowledge involving only the update of parameters for a specific component is assigned to an equipment expert for review alone; knowledge involving process adjustments and affecting the entire system is assigned to a joint review by equipment experts, process experts, and safety experts.

[0029] The data of each node in the knowledge graph is converted into vector form and stored in a vector database. Milvus is chosen as the vector database because it supports efficient similarity queries and large-scale data storage. In practical applications, other databases such as FAISS can also be selected, but this embodiment does not limit the choice. When a query is needed, the query request is converted into a vector and its similarity is calculated with the node vectors in the vector database. Nodes with a similarity higher than a set threshold (e.g., 0.8) are selected to enable fault cause reasoning, operational status knowledge recommendation, and data association retrieval.

[0030] Example 2 Based on the unified modeling and archiving method for multi-source data of the power generation process in Example 1, this example further illustrates the method.

[0031] Semantic recognition and entity extraction are performed on multi-source heterogeneous production data, including: parsing and classifying the multi-source heterogeneous production data based on the OID (Original Equipment ID), process tag number, and business tag of power generation equipment; extracting key information such as equipment name, operating condition characteristics, fault phenomena, and handling measures from unstructured text using large-scale model-based named entity recognition technology; and establishing a unified indexing mechanism for multi-modal external entity nodes to achieve semantic association between cross-modal data. Based on the OID, process tag number, and business tag of power generation equipment, preliminary parsing and classification of multi-source heterogeneous production data is performed to clarify the equipment and data type to which the data belongs. Then, an optimized large-scale model is used for deep processing of unstructured text to accurately extract key information. Finally, a unified index is established for multi-modal external entity nodes, and an association mapping table is constructed to achieve semantic association between different modalities of data, laying the foundation for subsequent knowledge graph construction and application.

[0032] The multi-source heterogeneous production data also includes semantic alignment processing, including: for structured data, extracting statistical feature vectors of real-time equipment parameters, including mean, variance, spectral features, and time-frequency analysis results, to construct multi-dimensional feature representations of parameter nodes; for unstructured text, using a domain-adapted pre-trained language model, combined with a power industry dictionary, to perform named entity recognition and relation extraction, extracting a quadruple of equipment name, fault mode, phenomenon description, and handling measures; for multimedia data, applying a multimodal feature extraction network to generate semantic embedding vectors for images or videos, and establishing cross-modal associations with equipment entities; and based on spatiotemporal consistency constraints, aligning and verifying the timestamps, geographical locations, and process locations of multi-source data to eliminate data asynchrony and spatial offset. This embodiment uses corresponding feature extraction and processing methods for different types of multi-source heterogeneous production data, converting them into feature representations or semantically related data in a unified format; then, through spatiotemporal consistency constraint verification, it eliminates differences in time, geographical location, and process location, achieving semantic alignment of multi-source data, enabling data from different sources and in different formats to be integrated within the same semantic framework.

[0033] The specific steps for constructing the three-layer semantic model include: In the equipment layer, a hierarchical topology of the generator set, system, and components is constructed through a first predefined relationship, forming a physical entity hierarchy; in the operating condition layer, the operating parameters, status, and cycle of the generator set are associated with the equipment layer through a second predefined relationship; in the knowledge layer, a semantic network of fault status, handling procedures, and historical cases is constructed, and interacts with the operating condition layer and equipment layer through a third predefined relationship; the equipment layer, operating condition layer, and knowledge layer are connected through bidirectional semantic relationships. First, the physical entity hierarchy of the equipment layer is constructed to clarify the organizational structure of the generator set; then, the operating parameters, status, and cycle of the operating condition layer are associated with the equipment layer to combine static and dynamic operating information; next, the semantic network of the knowledge layer is constructed, and its interaction relationship with the operating condition layer and equipment layer is established to achieve the fusion of knowledge information with equipment information and operating condition information; finally, the three-layer structure is connected through bidirectional semantic relationships to form a complete and interconnected semantic model, providing a foundation for data association queries and reasoning.

[0034] The first set of relationships in the device layer includes relationships such as "belongs to", "deployed in", and "contains".

[0035] The second set of relationships includes "belongs to," "affects," and "reflects." For example, boiler furnace temperature and pressure parameters are associated with furnace components through a "belongs to" relationship; turbine speed parameters are associated with the turbine rotor through a "belongs to" relationship. Normal, abnormal, and alarm states of equipment are associated with corresponding equipment components through a "reflects" relationship; for example, a furnace over-temperature alarm is associated with furnace components through a "reflects" relationship. Equipment operating cycles and maintenance cycles are associated with corresponding equipment or systems through a "belongs to" relationship; for example, the boiler system's maintenance cycle is associated with the boiler system through a "belongs to" relationship.

[0036] The third set of relationships includes relationships such as "cause", "recommended treatment measures", and "based on".

[0037] The equipment layer and the operating condition layer are connected through bidirectional relationships such as "influence" and "reflection." For example, changes in the state of equipment components affect operating condition parameters (from equipment layer to operating condition layer), and abnormalities in operating condition parameters reflect the state of equipment components (from operating condition layer to equipment layer). The equipment layer and the knowledge layer are connected through bidirectional relationships such as "cause" and "based on." For example, a failure of an equipment component will cause the corresponding knowledge node to be activated (from equipment layer to knowledge layer), and the processing measures in the knowledge node are used to solve the problem of the equipment component (from knowledge layer to equipment layer). The operating condition layer and the knowledge layer are connected through bidirectional relationships such as "cause" and "recommend processing measures." For example, abnormalities in operating condition parameters will cause the corresponding fault knowledge node to be activated (from operating condition layer to knowledge layer), and the processing measures in the knowledge node are used to improve the operating condition parameters (from knowledge layer to operating condition layer).

[0038] The differential update strategy specifically includes: storing the differences in the knowledge graph; establishing a version identifier for each knowledge node and recording the change time, the person making the change, and the reason for the change; setting knowledge effectiveness levels, which are used to classify knowledge into four states: draft, pending review, effective, and expired; and logically archiving expired knowledge instead of physically deleting it, thus preserving a complete historical traceability chain.

[0039] It should be noted that the draft status refers to the state of a knowledge node when it has just been created but has not yet been submitted for review. At this time, the knowledge node is only visible to the creator and does not participate in the query and reasoning of the knowledge graph.

[0040] Pending review status: After the creator submits the review request, the knowledge node enters the pending review status. At this time, review experts can view the knowledge node and perform the review operation.

[0041] Effective Status: Once a knowledge node has been approved by all designated review experts, it enters the effective status and can participate in the query and reasoning of the knowledge graph.

[0042] Expired status: When the equipment corresponding to a knowledge node has been phased out, the process has been updated, or the knowledge content has become invalid, it is set to the expired status. The knowledge node will no longer participate in the query and reasoning of the knowledge graph, but its data will be retained.

[0043] When the knowledge graph needs to be updated, only the changed parts are stored to reduce storage overhead. At the same time, a version identifier and change log are established for each knowledge node to record update-related information. By setting knowledge effectiveness levels, the knowledge update process is standardized to ensure that only knowledge that has been reviewed and approved can become effective and put into use. For expired knowledge, it is not physically deleted, but only logically marked and archived to retain complete historical data for subsequent traceability and query.

[0044] Furthermore, the multi-expert review and tiered approval process includes: setting up review expert roles in different professional fields, including equipment experts, process experts, and safety experts; automatically assigning appropriate review expert role combinations for review based on the type and scope of knowledge updates; establishing a priority assessment mechanism for knowledge updates to differentiate between the review processes for urgent and routine updates; and enabling real-time traceability of the entire process based on electronic records of review opinions and the basis for knowledge updates.

[0045] In this embodiment, the equipment expert has over 5 years of experience in the operation and maintenance of power generation equipment, is familiar with the equipment's structure, principles, fault diagnosis, and maintenance methods, and is responsible for reviewing equipment-related knowledge updates, such as equipment parameters, fault handling measures, and maintenance procedures. The process expert has over 5 years of experience in power generation process design and optimization, is familiar with power generation processes, process parameters, and process improvement methods, and is responsible for reviewing process-related knowledge updates, such as process adjustment plans and operational process optimization. The safety expert has over 3 years of experience in safety management within the power generation industry, is familiar with safety regulations, safe operating procedures, and risk assessment methods, and is responsible for reviewing safety-related knowledge updates, such as safe operating specifications and risk prevention measures.

[0046] The steps of using vector retrieval technology for fault diagnosis include: vectorizing real-time operating parameters and calculating their similarity with historical similar operating conditions; performing relational reasoning on the knowledge graph based on graph neural networks to identify potential fault propagation paths; dynamically ranking and recommending treatment measures based on fault severity and historical treatment effects; and generating a visual diagnostic report containing a complete knowledge chain of fault causes, impact scope, and treatment suggestions. The process involves converting real-time operating parameters into vector form and calculating their similarity with historical operating condition vectors to find similar historical operating conditions for reference; then using graph neural networks to perform relational reasoning on the knowledge graph to identify potential fault propagation paths; next, dynamically ranking treatment measures based on fault severity and historical treatment effects to improve the targeting of treatment measures; and finally, generating a visual diagnostic report that intuitively displays fault-related information and treatment suggestions to assist staff in fault handling.

[0047] In this embodiment, the graph neural network selected is GCN (Graph Convolutional Network). This model can effectively utilize the topological structure and node features of the knowledge graph for relational reasoning. The advantages of choosing this model are high reasoning accuracy and strong generalization ability. In practical applications, other models such as GAT can also be selected, but this embodiment does not limit this choice. The node features and adjacency matrix of the knowledge graph are input into the GCN model. Through multi-layer convolution operations, the model learns the potential relationships between nodes and identifies potential fault propagation paths. For example, it can infer other component nodes and operating condition nodes that may be affected by a fault node of a certain component.

[0048] Example 3 This embodiment provides a unified modeling and archiving system for multi-source data in the power generation process, including: a multi-source data access module for connecting to real-time monitoring systems, equipment management systems, document management systems, and multimedia acquisition devices in the power generation process; a semantic recognition and entity extraction engine for automatically classifying, recognizing entities, and extracting relationships from multi-source data; a three-layer knowledge graph builder for constructing a unified knowledge model including equipment, operating condition, and knowledge layers; a graph evolution manager for executing differential updates, version control, and review processes for the knowledge graph; and an intelligent query and reasoning module for realizing knowledge querying, fault diagnosis, and decision recommendation based on vector retrieval.

[0049] The multi-source data access module adopts a modular design, containing multiple standardized interface units that respectively interface with the real-time monitoring system (OPCUA interface unit), the device management system (JDBC interface unit), the document management system (RESTAPI interface unit), and the multimedia acquisition equipment (RTSP interface unit). Each interface unit is responsible for receiving, parsing, and format conversion of data, transforming data from different sources and in different formats into a unified JSON format before transmitting it to the semantic recognition and entity extraction engine.

[0050] The semantic recognition and entity extraction engine integrates named entity recognition, multimodal feature extraction, and semantic alignment algorithms based on a large model. Specifically, the named entity recognition algorithm uses a BERT model fine-tuned for the power industry, the multimodal feature extraction algorithm uses the CLIP model, and the semantic alignment algorithm employs an alignment algorithm based on spatiotemporal consistency constraints. After receiving uniformly formatted data from the multi-source data access module, the engine automatically identifies the data type (structured, unstructured text, multimedia) and calls the corresponding algorithms for processing, achieving automatic data classification, entity recognition, and relation extraction. The processing results are then transmitted to a three-layer knowledge graph builder.

[0051] The three-layer knowledge graph builder includes a device layer building unit, a working condition layer building unit, a knowledge layer building unit, and a hierarchical association unit.

[0052] Equipment layer construction unit: Based on equipment identification ID (OID), process tag number, and other information, construct the hierarchical topology of units, systems, and components.

[0053] Operating condition layer construction unit: Extracts information such as equipment operating parameters, status, and cycle, constructs operating condition layer nodes, and establishes the association relationship with the equipment layer.

[0054] Knowledge layer building unit: Integrate expert experience, fault handling procedures, historical cases and other knowledge to build a knowledge layer semantic network and establish the relationship with the equipment layer and operating condition layer.

[0055] Hierarchical association unit: Through preset semantic relationships (such as "cause", "impact", "recommended handling measures", etc.), bidirectional association is realized between the equipment layer, operating condition layer and knowledge layer.

[0056] The Knowledge Graph Evolution Manager comprises a Differential Update Unit, a Version Control Unit, and an Audit Process Unit. The Differential Update Unit uses incremental storage technology, storing only the changed portions of the knowledge graph, reducing storage overhead. The Version Control Unit assigns a version identifier to each knowledge node, records version change information, and supports historical version backtracking. The Audit Process Unit automates the management of multi-expert review and tiered activation processes, including expert allocation, review comment collection, and knowledge activation status updates.

[0057] The intelligent query and reasoning module integrates a vector retrieval unit, a relational reasoning unit, a treatment measure recommendation unit, and a visualization report generation unit. The vector retrieval unit uses the Milvus vector database to store knowledge graph node vectors, enabling efficient similarity queries. The relational reasoning unit, based on the GCN graph neural network model, performs relational reasoning on the knowledge graph, identifying potential fault propagation paths. The treatment measure recommendation unit dynamically ranks and recommends treatment measures based on fault severity and historical treatment effects. The visualization report generation unit uses the ECharts visualization tool to generate a visualized diagnostic report that includes the fault cause, impact scope, and treatment recommendations.

[0058] Example 4 In another embodiment of the present invention, a computer-readable storage medium is provided as a storage component within a terminal device, the function of which is to store programs and data. It should be noted that the computer-readable storage medium here encompasses not only the built-in storage components of the terminal device but also extended storage components supported by the device. Essentially, it is a tangible medium capable of containing or storing programs that can be invoked by or in conjunction with an instruction execution system, device, or apparatus. This storage medium provides storage areas for the terminal's operating system and stores one or more instructions suitable for processor loading and execution, which can constitute one or more computer programs containing program code.

[0059] Specifically, examples of computer-readable storage media (a non-exclusive list) include: electrical connections with one or more wires, portable disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable optical disc read-only memory, optical storage devices, magnetic storage devices, or any reasonable combination of the above types.

[0060] The storage medium may also include data signals propagated as part of a baseband portion or a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any reasonable combination of both. Furthermore, computer-readable storage medium may also refer to other readable media besides conventional readable storage media, capable of sending, propagating, or transmitting programs for use or operation by an instruction execution system, apparatus, or device. Program code on the storage medium can be transmitted via any suitable medium, including but not limited to wireless, wired, optical fiber, or any reasonable combination thereof.

[0061] The program code used to implement the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C. The execution modes of the program code include: running entirely on the user's computing device, running partially on the user's device as a standalone software package, running partially in a distributed manner on both the user's device and a remote computing device, or running entirely on a remote computing device or server. When a remote computing device is involved, the device can be connected to the user's computing device via any type of network such as a local area network (LAN) or a wide area network (WAN), or connected to an external computing device via the Internet through an Internet service provider.

[0062] The processor is capable of loading and executing one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the unified modeling and archiving method for multi-source data of the power generation process described in Example 1.

[0063] Example 5 Figure 2 This is a schematic diagram of a computer device provided according to an embodiment of the present invention.

[0064] Please see Figure 2 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the unified modeling and archiving method for multi-source data of the power generation process in this embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the computing system constituting the unified modeling and archiving method for multi-source data of the power generation process in this embodiment. To avoid repetition, these details are not elaborated here.

[0065] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 2This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.

[0066] The processor 61 may be a central processing unit (CPU), or other general-purpose processors, CPUs, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, quantum computing-based data processing logic units, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0067] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or RAM of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the computer device 60.

[0068] Furthermore, the memory 62 may include both internal storage units and external storage devices of the computer device 60. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.

[0069] Any references to memory, databases, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (Read-Only Memory). Memory includes ROM, magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0070] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

Claims

1. A method for unified modeling and archiving of multi-source data in a power generation process, characterized in that, include: Acquire multi-source heterogeneous production data, perform semantic recognition and entity extraction on the multi-source heterogeneous production data, and obtain structured data, unstructured text, images and videos; Structured data in multi-source heterogeneous production data is mapped to parameter nodes, unstructured text is mapped to knowledge nodes, and images and videos are mapped to multimodal external entity nodes. A unified knowledge graph is constructed based on a three-layer semantic model, which includes a device layer, an operating condition layer, and a knowledge layer. The device layer is used to define the hierarchical structure of physical entities, the operating condition layer describes the dynamic operating attributes of power generation equipment, and the knowledge layer contains expert experience and equipment operation rules. A differential update strategy is adopted to manage the evolution of the knowledge graph, record the full history of changes, and update the knowledge through a process of expert review and hierarchical implementation. It is used to query knowledge graphs based on vector retrieval technology, enabling fault cause reasoning, operational status knowledge recommendation, and data association retrieval.

2. The method for unified modeling and archiving of multi-source data in the power generation process according to claim 1, characterized in that, Semantic recognition and entity extraction of multi-source heterogeneous production data, including: Multi-source heterogeneous production data is parsed and classified based on the power generation equipment's identification OID, process tag number, and business tag. The named entity recognition technology based on large models is used to extract key information such as equipment name, operating condition characteristics, fault phenomena and handling measures from unstructured text. Establish a unified indexing mechanism for multimodal external entity nodes to enable semantic association of cross-modal data.

3. The method for unified modeling and archiving of multi-source data in the power generation process according to claim 1, characterized in that, The process of processing multi-source heterogeneous production data also includes semantic alignment processing, including: For structured data, statistical feature vectors of real-time device parameters are extracted, including mean, variance, spectral features, and time-frequency analysis results, to construct multi-dimensional feature representations of parameter nodes; For unstructured text, a domain-adapted pre-trained language model is used, combined with an electric power professional dictionary, to perform named entity recognition and relation extraction, extracting a quadruple of equipment name, fault mode, phenomenon description, and handling measures. For multimedia data, a multimodal feature extraction network is applied to generate semantic embedding vectors for images or videos and establish cross-modal associations with device entities; Based on spatiotemporal consistency constraints, the timestamps, geographical locations, and process locations of multi-source data are aligned and verified to eliminate data asynchrony and spatial offset.

4. The method for unified modeling and archiving of multi-source data in the power generation process according to claim 1, characterized in that, The specific steps for constructing the three-layer semantic model include: In the equipment layer, the hierarchical topology of the generator set, system and components of the power generation equipment is constructed through the first set relationship, forming a physical entity hierarchical structure; In the operating condition layer, the operating parameters, status, and cycle of the power generation equipment are associated with the equipment layer through the second setting relationship; In the knowledge layer, a semantic network of fault states, handling procedures, and historical cases is constructed, and it interacts with the operating condition layer and the equipment layer through a third-party defined relationship. The equipment layer, operating condition layer, and knowledge layer are connected through bidirectional semantic relationships.

5. The method for unified modeling and archiving of multi-source data in the power generation process according to claim 1, characterized in that, The differential update strategy specifically includes: Store the differences in the knowledge graph after the changes, establish a version identifier for each knowledge node, and record the change time, the person making the change, and the reason for the change. The knowledge validity level is set, which is used to classify knowledge into four states: draft, pending review, effective, and expired. Expired knowledge is logically archived rather than physically deleted, preserving a complete historical traceability chain.

6. The method for unified modeling and archiving of multi-source data in the power generation process according to any one of claims 1-5, characterized in that, The multi-expert review and tiered approval process includes: The roles of audit experts in different professional fields are set up, including equipment experts, process experts and safety experts; Based on the type and scope of the knowledge update, the system will automatically assign appropriate expert roles for review. Establish a priority assessment mechanism for knowledge updates, and differentiate the review process between urgent and routine updates; conduct full-process traceability based on the electronic records of review opinions and the basis for knowledge updates in real time.

7. The method for unified modeling and archiving of multi-source data in the power generation process according to claim 6, characterized in that, The steps for fault diagnosis using the vector retrieval technology include: The real-time running parameters are vectorized and similarity is calculated with historical similar working conditions; Based on graph neural networks, relational reasoning is performed on knowledge graphs to identify potential fault propagation paths; Based on the severity of the fault and the historical handling results, the recommended handling measures are dynamically ranked and recommended. Generates a visual diagnostic report containing a complete knowledge chain of fault causes, impact scope, and handling recommendations.

8. A unified modeling and archiving system for multi-source data in a power generation process, characterized in that, include: The multi-source data access module is used to connect to real-time monitoring systems, equipment management systems, document management systems, and multimedia acquisition devices during the power generation process to acquire multi-source heterogeneous production data. A semantic recognition and entity extraction engine is used for automatic classification, entity recognition, and relation extraction of multi-source heterogeneous production data; A three-layer knowledge graph builder is used to construct a unified knowledge model that includes an equipment layer, a working condition layer, and a knowledge layer. The knowledge graph evolution manager is used to perform differential updates, version control, and review processes for the knowledge graph. The intelligent query and reasoning module is used to realize knowledge query, fault diagnosis and decision recommendation based on vector retrieval.

9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform the unified modeling and archiving method for multi-source data of the power generation process as described in any one of claims 1 to 7.

10. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including steps for performing the steps in the unified modeling and archiving method for multi-source data of the power generation process according to any one of claims 1 to 7.