Method and device for constructing a knowledge graph of hydropower centralized control business based on a hybrid large model

Through the mixed large-scale model building of the hydropower integrated control business knowledge graph, the problem of insufficient semantic analysis and correlation capabilities of traditional models when processing unstructured data is solved, efficient and accurate data analysis and decision-making support are achieved, and the operation efficiency and data utilization of the hydropower integrated control center are improved.

CN119990282BActive Publication Date: 2025-07-29SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510452134.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-29
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

When the existing hydropower centralized control centers process unstructured data, it is difficult for traditional models to accurately analyze semantic information and complex associations in the data, resulting in low accuracy of information extraction and relationship modeling, and lack of deep semantic association capabilities for multi-source heterogeneous data, which affects the flexibility and efficiency of decision support.

Method used

The hybrid large model is used to build the knowledge graph of water and electricity integrated control business, and the model layer is constructed through the network ontology language OWL, combined with the BERT layer and the Transformer fusion layer, so as to realize the deep semantic understanding of unstructured data and the refined feature extraction of multi-source heterogeneous data, establish the business logic graph, equipment entity graph and case graph, and use the graph database to construct and infer the knowledge graph.

Benefits of technology

It improves the efficiency and accuracy of the construction of the knowledge graph, can better adapt to the needs of the water and power integrated control business, provide fast and accurate decision-making support, reduce manual intervention, and improves the comprehensive data utilization rate and the operation efficiency of the centralized control center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990282B_ABST
    Figure CN119990282B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for constructing a knowledge graph of hydropower centralized control services based on a hybrid large model. The method includes: constructing a schema layer of the knowledge graph; establishing a mapping relationship between unstructured data and triple data; constructing a hybrid large model; outputting triples of the knowledge graph of hydropower centralized control services through the hybrid large model, and further obtaining a data layer of the knowledge graph of hydropower centralized control services. The present invention can effectively improve the efficiency and accuracy of knowledge graph construction, which is of great significance for the intelligent operation and equipment management of the centralized control center, and provides a new solution for business decision-making and fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of constructing a knowledge graph for hydropower centralized control services, and particularly relates to a method and device for constructing a knowledge graph for hydropower centralized control services based on a hybrid large model. Background Art

[0002] With the continuous development of the hydropower industry, hydropower centralized control centers have gradually become the core of power system dispatching and operation. However, the challenges faced by centralized control centers mainly come from the form of data they process - a large amount of unstructured data. These data come from various information sources such as equipment operation logs, fault records, dispatching regulations, and maintenance manuals, with complex and diverse forms, and are difficult to be effectively integrated and utilized through traditional linear analysis models. When dealing with the diversity and uncertainty of this kind of data, existing data processing methods often lack sufficient flexibility and accuracy, resulting in slow information analysis speed and lagging decision support, thus affecting the overall operation efficiency and reliability of centralized control centers.

[0003] At the present stage, hydropower centralized control centers rely on traditional rule-based models and manual analysis methods in the fields of equipment status monitoring, fault diagnosis, emergency response, etc. These methods often require a large amount of manual intervention and the participation of domain experts, and it is difficult to achieve rapid integration and real-time analysis of multi-source data. Moreover, due to the lack of automated information extraction means, traditional methods are difficult to provide rapid and accurate decision support in the face of emergencies such as equipment failures and dispatching anomalies.

[0004] Traditional knowledge graph construction models (such as rule bases, bag-of-words models, etc.) still have drawbacks. When dealing with unstructured data, traditional models are difficult to accurately parse the semantic information and complex associations in the data, resulting in low accuracy of information extraction and relationship modeling. At the same time, these models usually rely on rule-driven or shallow machine learning, and it is difficult to mine deep semantic associations in large-scale, multi-source heterogeneous data, with poor scalability and adaptability. Therefore, traditional models lack the ability to model the global dependency relationships between data, restricting the ability of logical reasoning and inference.

[0005] The introduction of hybrid large models has become a promising technological innovation. By integrating specialized modules for identifying unstructured data, hybrid large models can not only effectively reduce computational complexity but also enhance the ability to capture global information in long-sequence data. Through a multi-level and multi-module design, such models can perform refined feature extraction on multi-source heterogeneous data, revealing the complex correlation structures between data to the greatest extent. For example, a method for automatically constructing a domain knowledge graph based on a large language model and prompt engineering disclosed in Chinese Patent Document CN117035076A constructs a knowledge graph through a large language model. However, the current technology for constructing a knowledge graph using large models still has the following deficiencies: (1) being broad but not refined, performing poorly in tasks related to specific vertical domains, especially in the field of hydropower centralized control business; (2) the model is proficient in processing text data while ignoring the large amount of sequence data that appears in the hydropower centralized control business field, such as flow processes and rainfall processes. Summary of the Invention

[0006] To at least solve one of the problems existing in the prior art, the present invention proposes a method and device for constructing a knowledge graph of hydropower centralized control business based on a hybrid large model, which uses the hybrid large model to analyze the massive heterogeneous data generated in the hydropower centralized control business and constructs a business logic graph of the centralized control center, an equipment entity graph, a concept graph, and a case graph. Compared with traditional rule-based models and manual analysis methods, it has higher model efficiency, can effectively improve the comprehensive utilization rate of data in the hydropower centralized control center, provide technical references for the daily operation and emergency response of the hydropower centralized control center, and is of great significance for the development of dispatching work in the hydropower centralized control center, reducing water waste, and minimizing property losses of personnel in sudden accidents.

[0007] To achieve the objectives of the present invention, the method for constructing a knowledge graph of hydropower centralized control business based on a hybrid large model provided by the present invention includes the following steps:

[0008] Relying on the data in the research area, use the Web Ontology Language (OWL) to construct the schema layer of the knowledge graph of the centralized control center, and the schema layer includes the schema layers of the business logic graph of the centralized control center, the equipment entity graph, the concept graph, and the case graph;

[0009] Establish a mapping relationship between unstructured data and triple data through open-source data and labeled heterogeneous data of the hydropower centralized control center, and establish a mapping training set and a test set for the mapping of unstructured text of sequence data to triples, a mapping training set and a test set for the mapping of text data triples, and a mapping training set and a test set for the mapping of sequence data and triple splicing data to triple data after information fusion, where a triple represents the relationship between one entity and another entity;

[0010] Build a hybrid large model, which includes an embedding layer, BERT1 layer, BERT2 layer, and Transformer fusion layer for capturing semantic relationships between data. The BERT1 layer takes sequence data as the independent variable and water condition form description as the dependent variable; the BERT2 layer takes text data as the independent variable and knowledge graph triples as the dependent variable; the Transformer fusion layer takes the TOKEN concatenation value of the expert analysis text and knowledge graph triples after converting the sequence data as the independent variable, and the output is the hydropower centralized control graph triples; adjust and optimize the model hyperparameters;

[0011] Taking the schema layer as the basis, use the hybrid large model to extract triple data from the unstructured data of centralized control, and instantiate the business logic graph, device entity graph, concept graph, and case graph of the centralized control center to form the data layer of the hydropower centralized control business knowledge graph.

[0012] Furthermore, by collecting the business processes, operation specifications, historical water condition data, work tickets, and operation tickets in the study area, build a schema layer of the centralized control center knowledge graph based on the Web Ontology Language OWL.

[0013] Furthermore, in the constructed schema layer, use the dispatching regulations, business processes, operation specifications, historical water condition data, work tickets, and operation tickets to build the ontology relationships between business personnel, devices, operations, concepts, and cases; use OWL to define the hierarchical relationships, attributes, and constraint conditions between business personnel, devices, operations, concepts, and cases, and define the schema layers of the business logic graph, device entity graph, concept graph, and case graph.

[0014] The built-in reasoning mechanism of the Web Ontology Language OWL can be used to perform logical consistency checks and reasoning analysis on the defined schema layer to automatically identify and correct potential inconsistencies.

[0015] Furthermore, the open-source data can adopt the data in the open-source databases of DBpedia, Wikidata, and Freebase.

[0016] The open-source databases contain a large amount of unstructured data, entities, attributes, and relationship information. Use the unstructured data and the corresponding entity, attribute, and relationship information to establish a mapping relationship from unstructured data to structured triples. Classify the unstructured data into sequence data and text data according to the unstructured data type, and establish a mapping training set and test set for sequence data unstructured text, and a mapping training set and test set for text data triples respectively.

[0017] Furthermore, by converting the sequence data into expert analysis text, knowledge fusion is performed using the triples extracted from the expert analysis text and the pure text data, and new triples incorporating the information of the sequence data are extracted. Then, the triples annotations that integrate the sequence description information and the text description information are analyzed, and a mapping training set and a test set from the triple splicing data to the triple data after information fusion are established.

[0018] Furthermore, the hybrid large model is an attention mechanism large model that trains the BERT1 and BERT2 large models separately in the sequence data recognition task and the text data recognition task, and performs semantic fusion of the extracted features in the Transformer fusion layer.

[0019] Furthermore, in the hybrid large model, the main task of the embedding layer is to convert discrete input data into a dense low-dimensional vector representation, capture the semantic relationships between data through the word embedding method. Its role is to reduce the feature dimension, optimize the computational performance of the model, reduce the consumption of memory and computing resources during the operation, and at the same time improve the expression ability of the data. During the mapping process, the embedding layer can better capture the semantic information of the input data, enhance the model's understanding ability of complex features, and effectively reduce the influence of noise and irrelevant information;

[0020] The Transformer fusion layer includes a multi-head attention layer and a feed-forward neural network, which are used to capture the context relationships between words. Its multi-head attention mechanism is as follows:

[0021] The input of the Transformer fusion layer is a one-dimensional vector , ,..., , the one-dimensional vector is obtained by concatenating the output of the BERT1 layer , ,..., and the output of the BERT2 layer , ,..., . Among them, represents the feature vector at the -th position in the one-dimensional vector, ∈ , represents the length of the entire one-dimensional vector of the input. For each position , its corresponding query vector , key vector and value vector are obtained through the following linear transformation:

[0022] ;

[0023] ;

[0024] ;

[0025] Among them, 、 、 are linear transformation matrices for query, key, and value respectively;

[0026] By calculating the similarity score between the query vector and the key vectors at other positions, the attention weight is obtained;

[0027] ;

[0028] Among them, represents the dimension of the query vector and the key vector.

[0029] Using the attention weight to perform weighted summation on the value vector , the output attention vector

[0030] at the corresponding position is obtained;

[0031] BERT Layer 1 and BERT Layer 2: Both BERT Layer 1 and BERT Layer 2 are BERT large models. Through the bidirectional encoder architecture in the Transformer model, multiple stacked Transformers are used to achieve in-depth semantic understanding of unstructured data. The BERT large model captures the context information between sentences and words by introducing the masked language model MLM and the next sentence prediction NSP task during the training process.

[0032] The principle of the masked language model MLM is as follows:

[0033] Given the input vectors , ,..., , represents the th one-dimensional vector. The masked language model MLM randomly masks some words in the input vector, denoted as the masked sequence , ,..., , is the th masked word; for each masked word , the masked language model MLM will generate a probability distribution , and this probability distribution characterizes the masked word in the given context The distribution in

[0034] The principle of predicting NSP in the latter sentence is as follows:

[0035] Given an input sentence and a sentence , denotes the TOKEN converted from different types of one-dimensional vectors, and for the sentence , Add the [CLS] token as the start symbol of the sentence pair, add the [SEP] token to separate the two sentences, and combine the input sentences into an embedding vector ;

[0036] Use a linear classifier to process the embedding vector to obtain the sentence coherence score .

[0037] ;

[0038] Among them denotes the weight matrix of the linear classifier, denotes the embedding vector, denotes the bias value of the linear classifier.

[0039] Furthermore, use the cross-entropy loss function to optimize the weight matrix of the linear classifier and the bias value , and the formula of the cross-entropy loss function is as follows:

[0040] ;

[0041] Among them, denotes the cross-entropy loss of the linear classifier; denotes the total number of samples, that is, the number of sentence pairs in the training set; denotes the true label of the th sample; denotes the prediction probability of the masked language model MLM or the latter sentence prediction NSP for the th sample.

[0042] Use the unstructured text mapping training set and test set of sequence data to train the BERT1 layer, use the text data triple mapping training set and test set to train the BERT2 layer, and fuse the collected vector features in the Transformer fusion layer.

[0043] Further, input the dispatching regulations, business processes, operation specifications, historical water regime data, work tickets, and operation ticket data into the constructed hybrid large model, and output triple sets of business logic graphs, equipment entity graphs, concept graphs, and case graphs, and instantiate the business logic graph, equipment entity graph, concept graph, and case graph of the centralized control center to form the knowledge graph data layer.

[0044] Further, the steps to form the hydropower centralized control business knowledge graph data layer include:

[0045] Extract triples based on the hydropower centralized control knowledge graph triples, with the format of "subject - predicate - object", where the subject and object are defined as nodes in the knowledge graph, and the predicate represents the edge connecting the nodes;

[0046] Use a graph database to import the nodes and relationships to form a graph structure, create nodes and edges, and form the hydropower centralized control business knowledge graph data layer.

[0047] The hydropower centralized control business knowledge graph construction device provided by the present invention includes the following modules:

[0048] The schema layer construction module is used to construct the schema layer of the centralized control center knowledge graph using the Web Ontology Language (OWL). The schema layer includes the schema layers of the business logic graph, equipment entity graph, concept graph, and case graph of the centralized control center;

[0049] example graph.

[0050] The dataset construction module is used to establish the mapping relationship between unstructured data and triple data through open - source data and labeled heterogeneous data of the hydropower centralized control center, and establish the mapping training sets and test sets of sequence data - unstructured text mapping, text data - triple group mapping, and sequence data and triple splicing data - information - fused triple data mapping respectively. Here, a triple represents the relationship between one entity and another entity;

[0051] The model construction module is used to construct a hybrid large model. The hybrid large model includes an embedding layer for capturing the semantic relationship between data, a BERT1 layer, a BERT2 layer, and a Transformer fusion layer. The BERT1 layer takes sequence data as the independent variable and the water regime form description as the dependent variable; the BERT2 layer takes text data as the independent variable and the knowledge graph triples as the dependent variable; the Transformer fusion layer takes the TOKEN splicing value of the expert analysis text after sequence data conversion and the knowledge graph triples as the independent variable, and the output is the hydropower centralized control graph triples;

[0052] The knowledge graph construction module is used to take the schema layer as the basis, utilize the hybrid large model to extract triple data from the unstructured data of the centralized control, and instantiate the business logic graph of the centralized control center, the equipment entity graph, the concept graph, and the case graph to form the data layer of the hydropower centralized control business knowledge graph.

[0053] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0054] (1) The present invention applies the hybrid large model to the construction of the knowledge graph of the centralized control center, which can realize the analysis and processing of heterogeneous data driven by actual semantics. Compared with the disadvantages existing in the application of traditional knowledge graph construction models (such as rule bases, bag-of-words models, etc.), such as heavy expert work and inability to understand context language, etc., the hybrid large model only needs to construct one model to complete the simultaneous analysis of a large amount of heterogeneous data, and utilizes the characteristic that the BERT model can understand semantics. After training, it can consider the semantic space correlation of all data instances, and the accuracy and construction efficiency of the knowledge graph can be greatly improved.

[0055] (2) The present invention separates the sequence data task and the text data task, and adopts the BERT1 layer and the BERT2 layer for training respectively, reducing the mutual interference between semantic information and improving the model training efficiency and inference efficiency. At the same time, specialized optimization training is carried out for a large amount of potential sequence data (such as flow time series, etc.) in the hydropower centralized control scheduling business, so as to realize the specialization of the hydropower centralized control business model and be more adaptable to the working scenario of the hydropower centralized control business.

[0056] (3) The present invention constructs the knowledge graph of the hydropower centralized control vertical field by fusing text data and sequence data through the hybrid large model, making the constructed knowledge graph more in line with the actual business requirements. Description of the Drawings

[0057] Figure 1 It is a flowchart of the semantic-driven hydropower centralized control business knowledge graph construction method based on the hybrid large model according to the embodiment of the present invention.

[0058] Figure 2 It is a schematic diagram of the schema layer of the business logic knowledge graph according to the embodiment of the present invention.

[0059] Figure 3 It is a schematic diagram of the hybrid large model structure adopted in the implementation of the present invention.

[0060] Figure 4 It is a schematic diagram of the business logic knowledge graph after instantiating the unstructured data according to the embodiment of the present invention. Detailed Embodiments

[0061] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not limited to the present invention.

[0062] A method for constructing a semantic-driven hydropower centralized control service knowledge graph of a hybrid large model provided by the present invention, as Figure 1 shown, includes the following steps:

[0063] S1. Relying on the basic data of the research area, use the Web Ontology Language (OWL) to construct the schema layer of the centralized control center knowledge graph. The schema layer of the centralized control center knowledge graph includes a business logic graph schema layer, a device entity graph schema layer, a concept graph schema layer, and a case graph schema layer.

[0064] In each constructed schema layer, use dispatching regulations, business processes, operation specifications, historical water condition data, work tickets, and operation ticket data to construct the ontological relationships between business personnel, devices, operations, concepts (basic water conservancy concepts), and cases (historical handling cases); use the Web Ontology Language (OWL) to define the hierarchical relationships, attributes, and constraint conditions between business personnel, devices, operations, concepts, and cases, and define the business logic graph schema layer, the device entity graph schema layer, the concept graph schema layer, and the case graph schema layer.

[0065] In this step, collect the business processes, operation specifications, historical water condition data, work tickets, and operation ticket data of the research area, and construct the schema layer of the centralized control center knowledge graph based on the Web Ontology Language (OWL). In some embodiments of the present invention, taking the local business logic graph as an example, construct a small-scale business logic graph. In the constructed small-scale business logic graph, the small-scale business logic graph can be generated in the hybrid large model by using the remote centralized control mains power outage dispatching regulation data and the schema layer. The remote centralized control mains power outage logic graph schema layer is one of the contents of the centralized control center business logic graph schema layer. Now, take the data of this part of the content as an example, as Figure 2 shown. The centralized control center business logic graph, device entity graph, concept graph, and case graph are all constructed using the same method under different data, that is, first construct the schema layer based on the Web Ontology Language (OWL), and then construct the instance graph using the hybrid large model under the constraints of the schema layer. As Figure 2The figure shows a business logic graph model layer for remote centralized control of utility power outages, built using the OWL (ontology language). Starting from the node "Remote centralized control of utility power outages," this business logic graph model layer uses the "phenomenon" edge to connect the corresponding fault phenomenon nodes when the fault occurs. The "next step" edge indicates that the next node will be advanced when the execution conditions are met. The "operation start point" node is a marker node that indicates the next action to be taken. Following the "operation start point" node are corresponding measures to be taken from three different perspectives: the "dispatcher," the "dispatcher," and the "power station side." From the "dispatcher" perspective, there is an "edge" consisting of conditional branches: "Line forced transmission successful" and "Hydropower station busbar differential protection activated or line forced transmission unsuccessful." Different branch conditions correspond to different measures. The device entity graph model layer, the concept graph model layer, and the case graph model layer are all built using the OWL (ontology language). Figure 2 The specific logic of the pattern layer graph in is as follows:

[0066] The business logic knowledge graph model layer created a total of 15 nodes and 20 relationships. Among them, the 15 nodes include 1 plan node (remote emergency response plan for depressurization of the entire hydropower and new energy plant in the centralized control center), 1 situation node (remote centralized control, shutdown operation, depressurization of the entire plant), 4 phenomenon nodes and 8 operation nodes.

[0067] The four phenomenon nodes include:

[0068] Phenomenon node 1: Phenomenon (non-monitoring - manual confirmation of protection action, accident tripping);

[0069] Phenomenon Node 2: Phenomenon (Equipment monitoring - two outgoing lines Ⅰ and Ⅱ, 201 and 202, tripped simultaneously, causing the 220kV line and busbar to lose voltage simultaneously);

[0070] Phenomenon node 3: Phenomenon (equipment monitoring - unit outlet switches 021, 022, and 023 tripped);

[0071] Phenomenon Node 4: Phenomenon (Equipment Monitoring - 10kV and 400V Factory Power Supplies All Disappeared);

[0072] The 8 operation nodes include:

[0073] Operation Node 1: Operation (the dispatcher reports to the higher-level dispatching agency, disconnects all switches on the busbar without waiting for orders, and notifies the power station to restore power);

[0074] Operation Node 2: Operation (dispatcher - report the situation to the superior organization and notify the power station to check the status of primary and secondary equipment);

[0075] Operation Node 3: Operation (Dispatcher - Report to the superior dispatching agency the factors affecting the forced power transmission of the line, and report the detailed information such as fault recording and traveling wave ranging to the superior dispatching agency);

[0076] Operation Node 4: Operation (Dispatcher-in-Chief - The dispatcher-in-chief reports the situation to the person in charge of the dispatching operation department);

[0077] Operation Node 5: Operation (Power Station Side - The controlled power station reports the equipment inspection situation);

[0078] Operation Node 6: Operation (Dispatcher - The affected station service power should be restored as soon as possible);

[0079] Operation Node 7: Operation (Dispatcher - Apply to the superior agency for remote startup or on-site disposal, and perform remote startup or order on-site personnel to conduct disposal);

[0080] Operation Node 8: Operation (Dispatcher-in-Chief - The dispatcher-in-chief reports the disposal situation to the department head and makes a record of the situation);

[0081] Among them, the relationships between the nodes are as follows:

[0082] Plan Node → Situation Node: The relationship is "Situation", indicating the specific situation targeted by the plan;

[0083] Situation Node → Phenomenon Node 1: The relationship is "Phenomenon", indicating a phenomenon in this situation;

[0084] Situation Node → Phenomenon Node 2: The relationship is "Phenomenon", indicating a phenomenon in this situation;

[0085] Situation Node → Phenomenon Node 3: The relationship is "Phenomenon", indicating a phenomenon in this situation;

[0086] Situation Node → Phenomenon Node 4: The relationship is "Phenomenon", indicating a phenomenon in this situation;

[0087] Situation Node → Operation Starting Point: The relationship is "Next Step", indicating the starting point of entering the operation process from the situation;

[0088] Operation Starting Point → Operation Node 1: The relationship is "Dispatcher", indicating the first operation of the dispatcher;

[0089] Operation Starting Point → Operation Node 4: The relationship is "Dispatcher-in-Chief", indicating the first operation of the dispatcher-in-chief;

[0090] Operation Starting Point → Operation Node 5: The relationship is "Power Station Side", indicating the first operation of the power station side;

[0091] Operation Node 1 → Operation Node 2: The relationship is "Next Step", indicating the order of operations;

[0092] Operation Node 2 → Operation Node 3: The relationship is "Next Step", indicating the order of operations;

[0093] Operation Node 3 → Operation Node 6: The relationship is "Successful Line Strong Transmission", indicating the next operation in the case of success;

[0094] Operation Node 3 → Operation Node 7: The relationship is "Hydroelectric Power Station Bus Differential Protection Action or Unsuccessful Line Strong Transmission", indicating the next operation in the case of failure;

[0095] Operation Node 4 → Operation Node 8: The relationship is "Next Step", indicating the operation process of the dispatching chief.

[0096] Utilize the built-in reasoning mechanism of the Web Ontology Language OWL to conduct logical consistency checks and reasoning analysis on each defined schema layer to automatically identify and correct potential inconsistencies.

[0097] S2. Establish a mapping relationship between unstructured data and triple data through an open-source database and the labeled heterogeneous data of the hydropower centralized control center, and establish a training set and a test set for the mapping of unstructured text of sequence data, a training set and a test set for the mapping of triple data of text data, and a training set for the mapping of the concatenated data of sequence data and triples to triple data after information fusion. Among them, a triple represents the relationship between one entity and another entity, and the format is (head entity, relationship, tail entity).

[0098] In some embodiments of the present invention, the required open-source data can be extracted from open-source databases such as DBpedia, Wikidata, and Freebase.

[0099] The open-source database contains a large amount of unstructured data, entities, attribute, and relationship information. Utilize the unstructured data and the corresponding entity, attribute, and relationship information to establish a mapping relationship from unstructured data to structured triples. According to the type of unstructured data, classify the unstructured data into sequence data and text data, and establish a training set and a test set for the mapping of unstructured text of sequence data and a training set and a test set for the mapping of triple data of text data respectively.

[0100] In some embodiments of the present invention, the construction steps of the labeled heterogeneous data of the hydropower centralized control center include:

[0101] Experts analyze the logical structure in daily operation documents such as regulations and work tickets, select the key entities in the text according to operation experience, and establish the relationships between the key entities to form triple annotations of (head entity, relationship, tail entity);

[0102] For sequence data, such as rainfall, flow rate, etc., experts analyze the key concepts in the sequence data and give relevant description standards;

[0103] Concatenate the annotation of text description data (entity triples extracted from text description data) and the annotation of sequence data (expert analysis text made based on sequence data) as the input of the Transformer fusion layer. The output of the Transformer fusion layer is the triple of the hydropower centralized control atlas.

[0104] By converting sequence data into expert analysis text, use this expert analysis text to perform knowledge fusion with the triples extracted from pure text data such as dispatching regulations and work tickets, and extract new triples that combine the information of sequence data. Then experts analyze and annotate the triples that integrate sequence description information and text description information, and establish a mapping training set and test set from triple splicing data to post-triple data of information fusion.

[0105] The annotated heterogeneous data of the hydropower centralized control center is used to optimize the performance of the hybrid large model.

[0106] S3. Build a hybrid large model. The hybrid large model includes an embedding layer, a BERT1 layer, a BERT2 layer, and a Transformer fusion layer. For the BERT1 layer, the independent variable is sequence data and the dependent variable is the form description of the water regime; for the BERT2 layer, the independent variable is text data and the dependent variable is the knowledge graph triple; for the Transformer fusion layer, the independent variable is the TOKEN concatenation value of the expert analysis text converted from sequence data and the knowledge graph triple extracted by BERT2, and the dependent variable is the knowledge graph triple, and adjust and optimize the model hyperparameters. The specific input and output data structures adopted in the hybrid large model are as follows Figure 3 .

[0107] In some embodiments of the present invention, the independent variables of the BERT1 layer include historical rainfall data and flow rate data. The rainfall duration is 180 minutes, and the rainfall amount every 10 minutes is used as a feature, with a total of 18 features. The flow rate data includes the flow rate changes in the past 30 minutes, with a total of 18 flow rate features. The dependent variable is the form description text of the flood control and power generation tasks under the current rainfall and flow rate conditions. For the sequence data generated under 60 simulated rainstorm and waterlogging scenarios, it is randomly divided into a training set and a test set at a ratio of 70:30. The training set contains 42 samples and the test set contains 18 samples, which are used to evaluate the performance of the BERT1 model.

[0108] The main content of the triples of the text data of the BERT2 layer includes the theme, related terms, and background information. The structure of the triples is (head entity, relation, tail entity), where the theme is used as the input feature, and the related terms and background information are the output features. The theme refers to the name, approximate terms, common sayings, etc. of the current main water conservancy tasks; the related terms refer to the relative pronouns that appear in a series of water conservancy tasks generated around this theme; the background information is various concepts and important knowledge points related to this theme. In some embodiments of the present invention, a total of 100 themes are extracted, 300 related terms are associated, and the extracted themes include "flood risk", "drainage system", "reservoir management", etc. The related terms include "including", "drainage", "flood storage", etc., and the background information includes "basic geographic information", "topography and landforms", etc. In text data processing, it is divided according to the integrity and information richness of the text, and randomly allocated in a ratio of 60:40 to ensure the consistency of the theme distribution of the text data triple mapping training set and the test set. The training set contains 60 themes, and the test set contains 40 themes.

[0109] In some embodiments of the present invention, the Transformer fusion layer is trained using a spliced TOKEN mixed data set with a ratio of 1:1 of knowledge graph triples and sequence data. The BERT1 layer is trained using a sequence data unstructured text mapping training set, and the BERT2 layer is trained using a text data triple mapping training set. The specific input and output data structures used in the hybrid large model are shown in Table 1.

[0110] Table 1 Examples of input and output of the hybrid large model

[0111]

[0112] The hybrid large model is an attention mechanism large model that trains the BERT1 and BERT2 large models separately in the sequence data recognition task and the text data recognition task, and performs semantic fusion of the extracted features in the Transformer fusion layer.

[0113] The main work of the embedding layer is to convert the discrete input data, namely sequence data and text data, into a dense low-dimensional vector representation, capture the semantic relationship between data through the word embedding method, its role is to reduce the feature dimension, optimize the computational performance of the hybrid large model, reduce the consumption of memory and computing resources during the operation process, and at the same time improve the expression ability of the data. During the mapping process, the embedding layer can better capture the semantic information of the input data, enhance the understanding ability of the hybrid large model for complex features, and effectively reduce the influence of noise and irrelevant information, such as Figure 3 shown, , ,..., represents a single TOKEN of the sequence data, , ,..., represents a single TOKEN of the text data.

[0114] The Transformer fusion layer includes a multi-head attention layer and a feed-forward neural network, which are used to capture the context relationship between words. The feed-forward neural network includes a feed-forward layer and an embedding layer; the input of the Transformer fusion layer is a one-dimensional sequence, and the working mechanism of its multi-head attention layer is as follows:

[0115] Input one-dimensional vector , ,..., , one-dimensional vector is obtained by concatenating the outputs of BERT1 layer , ,..., and the outputs of BERT2 layer , ,..., , where represents the feature vector at the -th position in the one-dimensional vector, ∈ , represents the length of the entire one-dimensional input vector. For each position , its corresponding query vector , key vector and value vector can be obtained through the following linear transformation:

[0116] ;

[0117] ;

[0118] ;

[0119] where , , are the linear transformation matrices for query, key, and value respectively. By calculating the similarity score between the query vector and the key vectors of other positions, the attention weights can be obtained:

[0120] ;

[0121] where represents the dimension of the query vector and the key vector.

[0122] After obtaining the attention weights, the value vectors are weighted and summed using the attention weights to obtain the output attention vectors at the corresponding positions : :

[0123] ;

[0124] where represents the data at the -th position;

[0125] The feed-forward layer is a simple linear connection layer, and its calculation formula is as follows:

[0126] ;

[0127] where represents the output of the -th position component, and represents the feed-forward layer transformation matrix.

[0128] Both BERT1 layer and BERT2 layer are BERT large models. Through the bidirectional encoder architecture in the Transformer model, multiple stacked Transformer encoders are used to achieve in-depth semantic understanding of unstructured data.

[0129] The BERT large model captures the context information between sentences and words by introducing the masked language model MLM and the next sentence prediction NSP tasks during the training process. The principles of the masked language model MLM and the next sentence prediction NSP are as follows:

[0130] Masked Language Model MLM: Given the input vectors , ,..., , represents the -th one-dimensional vector with a certain number of dimensions. TOKEN or intermediate variables during model operation are in the form of one-dimensional vectors. The masked language model MLM randomly masks some words in the input vectors, denoted as the masked sequence , ,..., , is the -th masked one-dimensional vector, which can be regarded as a word here. For each masked word , the masked language model MLM will generate a probability distribution , which characterizes the distribution of the masked word in the given context .

[0131] Next Sentence Prediction NSP: Given the input sentence and sentences , represents TOKENs converted from different types of one-dimensional vectors, for sentences 、 Add the [CLS] token as the start token of the sentence pair, add the [SEP] token to separate the two sentences, and merge the input sentences into an embedding vector ; Use a linear classifier to process the embedding vector to obtain the sentence coherence score :

[0132] ;

[0133] Among them, represents the weight matrix of the linear classifier, represents the embedding vector, represents the bias value of the linear classifier.

[0134] Use the cross-entropy loss function to optimize the weight matrix and bias value of the linear classifier. The formula of the cross-entropy loss function is as follows:

[0135] ;

[0136] Among them, represents the cross-entropy loss of the linear classifier; represents the total number of samples, that is, the number of sentence pairs in the training set of the text data triple mapping; represents the true label of the th sample; represents the predicted probability of the th sample in the masked language model MLM or the next sentence prediction NSP. Here, there is a one-to-one correspondence between the th position and the th sample. Essentially, it is to mask the th position of a complete sentence and regard the th position as a sample.

[0137] Use the training set and test set of the unstructured text mapping of sequence data to train the BERT1 layer, use the training set and test set of the text data triple group mapping to train the BERT2 layer, and the collected vector features are fused in the Transformer fusion layer to output the triple of the hydropower centralized control knowledge graph. Subsequently, based on the triple of the hydropower centralized control knowledge graph, the hydropower centralized control knowledge graph can be created using a graph database.

[0138] When constructing the hybrid large model, parameters need to be set and optimized. In some embodiments of the present invention, the parameters of the hybrid large model are optimized by the grid search method and the 5-fold cross-validation method. The optimized parameter settings are shown in Table 2.

[0139] Table 2 Parameters and values of the BERT1 layer, BERT2 layer, and Transformer fusion layer used

[0140]

[0141] S4. Input the unstructured data into the hybrid large model, and output the business logic graph triple, device entity graph triple, concept graph triple, and case graph triple. Instantiate the business logic graph, device entity graph, concept graph, and case graph of the centralized control center to form the data layer of the hydropower centralized control business knowledge graph.

[0142] Based on the schema layer, use the hybrid large model to convert unstructured data (centralized control center dispatching regulations, business processes, operation specifications, historical water condition data, work tickets, and operation tickets data) into triple data. This triple data extracts the key information from the above unstructured data and establishes a directed graph relationship between the key information. Instantiate the business logic graph, device entity graph, concept graph, and case graph of the centralized control center to form the data layer of the hydropower centralized control business knowledge graph.

[0143] The process of constructing the business logic graph of the centralized control center according to the hydropower centralized control knowledge graph triples output by the hybrid large model (the processes of the device entity graph, concept graph, and case graph are the same as the process of constructing the business logic, only the data and schema layers are different) includes extracting triples in the format of "subject - predicate - object", covering the devices, processes, and responsible persons in the centralized control center. The subject and object are defined as nodes in the knowledge graph, while the predicate represents the edge connecting these nodes. For example, in the triple (Emergency Working Group, responsible for, Equipment Maintenance), "Emergency Working Group" and "Equipment Maintenance" are nodes, and "responsible for" is the connecting edge.

[0144] Use the graph database to import the nodes and relationships to form a graph structure, and create nodes and edges through the corresponding query language of the graph database syntax (in some embodiments of the present invention, the graph database can use Neo4j, and the syntax can use Cypher. It can be understood that in other embodiments, other graph databases and syntax can also be used). At the same time, supplement additional node attributes (such as device status, responsible person, etc.) and complex relationships (such as the sequence of events). Finally, use the inference engine to check the rationality of the graph to ensure logical consistency. In some embodiments of the present invention, the example results are as Figure 4 , Figure 4It includes key objects such as "battery", "USP power supply", and "environmental monitoring system", business personnel such as "centralized control dispatcher" and "emergency work group", and measures to be taken such as "prepare tools and materials", "rush to the scene for handling", "turn off unnecessary equipment", "notify the emergency work group", and "check the operation status of equipment". These node items are connected through "trigger", "facility", "execution", "leadership", "non-compliance", and "monitoring" relationships. This knowledge graph can meet basic business logic relationships and can be applied to logical reasoning and judgment.

[0145] In some embodiments of the present invention, there is provided a device for constructing a hydropower centralized control business knowledge graph based on a hybrid large model, and the device includes the following modules:

[0146] A schema layer construction module, configured to construct the schema layer of the centralized control center knowledge graph by using the Web Ontology Language OWL, and the schema layer includes the schema layers of the centralized control center business logic graph, the device entity graph, the concept graph, and the case graph;

[0147] A dataset construction module, configured to establish a mapping relationship between unstructured data and triple data through open-source data and labeled heterogeneous data of the hydropower centralized control center, and respectively establish a mapping training set and a test set for the unstructured text of sequence data, a mapping training set and a test set for the triple group of text data, and a mapping training set and a test set for the data obtained by splicing sequence data and triples to the triple data after information fusion, where a triple represents the relationship between one entity and another entity;

[0148] A model construction module, configured to construct a hybrid large model, and the hybrid large model includes an embedding layer, a BERT1 layer, a BERT2 layer, and a Transformer fusion layer for capturing semantic relationships between data. The BERT1 layer takes sequence data as an independent variable and the water situation form description as a dependent variable; the BERT2 layer takes text data as an independent variable and the knowledge graph triples as a dependent variable; the Transformer fusion layer takes the TOKEN splicing value of the expert analysis text after the sequence data is transformed and the knowledge graph triples as an independent variable, and the output is the hydropower centralized control graph triples;

[0149] A knowledge graph construction module, configured to use the hybrid large model to extract triple data from the centralized control unstructured data with the schema layer as the base, and instantiate the centralized control center business logic graph, the device entity graph, the concept graph, and the case graph to form the data layer of the hydropower centralized control business knowledge graph.

[0150] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a knowledge graph of hydropower centralized control business based on a hybrid large model, characterized in that It includes the following steps: Construct the schema layer of the centralized control center knowledge graph using the Web Ontology Language (OWL); the schema layer of the centralized control center knowledge graph includes the schema layers of the centralized control center business logic graph, device entity graph, concept graph, and case graph; Establish the mapping relationship between unstructured data and triple data through open-source data and labeled heterogeneous data of the hydropower centralized control center; Taking the schema layer of the centralized control center knowledge graph as the basis, use the pre-constructed hybrid large model to extract triple data from the unstructured data of the centralized control, instantiate the centralized control center knowledge graph, and form the data layer of the hydropower centralized control business knowledge graph; The hybrid large model includes an embedding layer, BERT1 layer, BERT2 layer, and Transformer fusion layer for capturing semantic relationships between data. The BERT1 layer takes sequence data as the independent variable and water condition form description as the dependent variable; the BERT2 layer takes text data as the independent variable and knowledge graph triples as the dependent variable; the Transformer fusion layer takes the concatenated value of the expert analysis text and knowledge graph triples after the transformation of sequence data as the independent variable, and the output is the hydropower centralized control graph triples; The establishment of the mapping relationship between unstructured data and triple data includes: respectively establishing the mapping training set and test set of sequence data unstructured text, the mapping training set and test set of text data triple groups, and the mapping training set and test set of sequence data and triple concatenated data to the triple data after information fusion, where triples represent the relationship between one entity and another entity; In the constructed schema layer, use dispatching regulations, business processes, operation specifications, historical water condition data, work ticket, and operation ticket data to construct the ontology relationship between business personnel, equipment, operations, concepts, and cases; use the Web Ontology Language (OWL) to define the hierarchical relationship, attributes, and constraint conditions between business personnel, equipment, operations, concepts, and cases, and define the schema layers of the business logic graph, device entity graph, concept graph, and case graph.

2. The method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to claim 1, wherein The open-source data contains a large amount of unstructured data, entities, attributes, and relationship information. Use the unstructured data and the corresponding entity, attribute, and relationship information to establish the mapping relationship from unstructured data to structured triples; according to the type of unstructured data, respectively establish the mapping training set and test set of sequence data unstructured text, and the mapping training set and test set of text data triple groups.

3. The method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to claim 1, wherein, By converting sequence data into expert analysis text, use the expert analysis text and the triples extracted from pure text data for knowledge fusion, and extract new triples that combine the information of sequence data; Then analyze the triple annotation that fuses sequence description information and text description information, and establish the mapping training set and test set of triple concatenated data to the triple data after information fusion.

4. The method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to claim 1, wherein The Transformer fusion layer includes a multi-head attention layer and a feed-forward neural network for capturing the context relationship between words; Both the BERT1 layer and the BERT2 are BERT large models. Through the bidirectional encoder architecture in the Transformer model, multiple stacked Transformers are used to achieve in-depth semantic understanding of unstructured data. The BERT large model captures the context information between sentences and words by introducing the masked language model (MLM) and next sentence prediction (NSP) during the training process.

5. The method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to claim 4, wherein The subsequent sentence prediction NSP includes: a given input sentence and the sentence , representing TOKENs converted from different types of one-dimensional vectors, for the sentence , adding the [CLS] token as the start symbol of the sentence pair, adding the [SEP] token to separate the two sentences, and merging the input sentences into an embedding vector ; processing the embedding vector using a linear classifier to obtain the sentence coherence score .

6. The method for constructing a knowledge graph of hydropower centralized control services based on a hybrid large model according to claim 5, wherein The weight matrix and bias value of the linear classifier are optimized using the cross-entropy loss function. The formula for the cross-entropy loss function is as follows: Among them, represents the cross-entropy loss of the linear classifier; represents the total number of samples; represents the true label of the th sample; represents the predicted probability of the th sample in the masked language model MLM or next sentence prediction NSP.

7. The method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to any one of claims 1-6, characterized in that The steps to form the data layer of the hydropower centralized control business knowledge graph include: Extract triples based on the hydropower centralized control knowledge graph triples, in the format of "subject - predicate - object". The subject and object are defined as nodes in the knowledge graph, and the predicate represents the edge connecting the nodes; Use the graph database to import the nodes and relationships to form a graph structure, create nodes and edges, and form the data layer of the hydropower centralized control business knowledge graph.

8. An apparatus for implementing the method for constructing a hydropower centralized control service knowledge graph based on a hybrid large model according to claim 1, characterized in that, The device includes the following modules: A schema layer construction module, which is used to construct the schema layer of the centralized control center knowledge graph using the Web Ontology Language (OWL); A dataset construction module, which is used to establish the mapping relationship between unstructured data and triple data through open-source data and the labeled heterogeneous data of the hydropower centralized control center; A model construction module, which is used to construct a hybrid large model. The hybrid large model includes an embedding layer, a BERT1 layer, a BERT2 layer, and a Transformer fusion layer for capturing the semantic relationships between data. The BERT1 layer takes sequence data as the independent variable and the water regime form description as the dependent variable; the BERT2 layer takes text data as the independent variable and the knowledge graph triples as the dependent variable; the Transformer fusion layer takes the concatenated value of the expert analysis text and the knowledge graph triples after the sequence data conversion as the independent variable, and the output is the hydropower centralized control graph triples; A knowledge graph construction module, which is used to use the hybrid large model to extract triple data from the unstructured data of the centralized control, instantiate the business logic graph of the centralized control center, the device entity graph, the concept graph, and the case graph based on the schema layer, and form the data layer of the hydropower centralized control business knowledge graph.

Citation Information

Patent Citations

  • Domain knowledge graph automatic construction method based on large language model and prompt project

    CN117035076A

  • Knowledge reasoning method based on industrial mechinery fault diagnosis knowledge graph

    CN113961718A

  • Semantic understanding method, computer equipment and storage medium

    CN114117008A