Method and device for constructing a large model of water conservancy engineering knowledge graph
By constructing a large water conservancy knowledge graph model that integrates multiple business knowledge and combining it with the self-iteration technology of the large language model, the problem of weak knowledge reasoning ability of existing water conservancy large models in typical business areas of water conservancy projects is solved, and effective support for multi-objective collaborative reasoning and intelligent decision-making is achieved.
Patent Information
- Application Number
- CN202510847148.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing large water conservancy models cannot effectively support typical business (L3) applications of water conservancy projects, such as dam safety monitoring and unit failure prediction, and have weak knowledge reasoning capabilities.
A method for constructing the water conservancy knowledge graph data layer, concept layer and instance layer by integrating multiple business knowledge is adopted, combined with self-iteration technology based on a large language model, to improve the knowledge graph of typical business fields of water conservancy projects. Through the deep integration of the knowledge graph and the large language model, a large model of the water conservancy project knowledge graph is created.
It has broken through the technical bottleneck of multi-objective collaborative reasoning and intelligent decision-making in typical business areas of water conservancy projects, and can support multi-objective collaborative solution reasoning and intelligent decision-making in flood control, water supply, power generation and other aspects of water conservancy projects, and promote the intelligent transformation of the water conservancy and hydropower industry.
Smart Images

Figure CN120372019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital twin engineering, and in particular to a method and device for constructing a large model of a water conservancy engineering knowledge graph. Background Art
[0002] The Water Conservancy Big Model is a standardized, systematic big model designed to meet the needs of the water conservancy and hydropower industry. It utilizes a "1+2+N" paradigm, leveraging a knowledge base comprised of structured and unstructured data as its engine, and centered around mechanistic and artificial intelligence models. This model targets water conservancy and hydropower business scenarios, primarily enabling intelligent empowerment of water conservancy projects and supporting high-quality development of the water conservancy sector. The "1" here refers to the water conservancy professional knowledge base, the "2" refers to the mechanistic and artificial intelligence models, and the "N" refers to multiple business scenarios. Following the development of artificial intelligence technology and big language models, two Chinese universities and research institutes have pioneered big model development projects for the water conservancy and hydropower industry. The Department of Water Conservancy at Tsinghua University established an interdisciplinary research team and, in April 2024, launched a big model project for the water conservancy industry based on the independently developed GLM big model. The China Institute of Water Resources and Hydropower Research officially released its independently developed water science big model, "SkyLIM," in May 2024. These big models will significantly advance the intelligent transformation of the water conservancy and hydropower industry and develop new high-quality water conservancy productivity.
[0003] Current large-scale water conservancy models are developed to address the goal of building a "digital twin river basin." They primarily address the needs of the water conservancy and hydropower industry at the river basin level. While the model data base effectively supports applications in broad business areas (Level 2) such as water disasters, water resources, water environment, and water ecology, it lacks support for typical business areas (Level 3) such as dam safety monitoring and unit failure prediction for water conservancy projects. The knowledge engine of current large-scale water conservancy models, driven by large language models (LLMs), excels in knowledge representation, extraction, and fusion. However, it suffers from weak knowledge reasoning capabilities, making it incapable of supporting reasoning for multi-objective collaborative solutions and intelligent decision-making for water conservancy projects. Knowledge graphs, by representing knowledge base information as a network of entities (nodes) and relationships (edges), mimic the structure of human knowledge. They capture not only the basic information of the knowledge base but also the higher-order relationships among multimodal information, including Euclidean and non-Euclidean data, and possess powerful reasoning capabilities.
[0004] In view of this, under the urgent need for intelligent transformation and upgrading of water conservancy, how to use advanced technical means such as knowledge graphs and large language models to propose a deep fusion model of knowledge graphs and large language models to break through the technical bottleneck of intelligent decision-making in water conservancy projects is a direction that urgently needs technical research. Summary of the Invention
[0005] In order to overcome the shortcomings of the above-mentioned existing technologies, the present invention provides a method and device for constructing a large model of water conservancy engineering knowledge graph, which can break through the technical bottlenecks of "multi-objective collaborative reasoning" and "intelligent decision-making" of large models in typical business fields of water conservancy projects, promote the intelligent transformation of the water conservancy and hydropower industry, and develop new quality productivity of water conservancy.
[0006] According to one aspect of the present invention, a method for constructing a large model of a water conservancy project knowledge graph is provided, comprising:
[0007] Based on the water conservancy project management business knowledge base, a water conservancy knowledge graph data layer, concept layer and instance layer construction method that integrates multiple business knowledge is adopted to build a knowledge graph of typical water conservancy project business areas, and refine the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction business;
[0008] Improve the knowledge graph of typical water conservancy engineering business areas using self-iterative knowledge graph completion technology based on a large language model;
[0009] Based on the knowledge graph retrieval enhancement generation technology, entities, relationships and their attributes are automatically extracted from structured and unstructured knowledge bases to create a large model of water conservancy project knowledge graph to support the reasoning application of multi-objective collaborative solutions for water conservancy projects.
[0010] The above technical solution adopts retrieval enhancement generation technology based on knowledge graph to automatically extract entities, relationships and their attributes from structured and unstructured knowledge bases, and create a large model of water conservancy project knowledge graph for flood control, water supply, power generation scheduling, etc. The constructed large model of water conservancy project knowledge graph combines the advantages of knowledge graph and large model in knowledge representation and knowledge analysis to realize the reasoning application of multi-objective collaborative solutions such as flood control, water supply, and power generation of water conservancy projects.
[0011] The above technical solution targets the complex knowledge structure and rich potential associations of water conservancy business, and improves the completeness and accuracy of the knowledge graph of typical water conservancy business areas through self-iterative water conservancy business knowledge graph completion technology based on a large language model.
[0012] As a further technical solution, a knowledge graph of typical water conservancy engineering business areas is constructed, including:
[0013] The BERT-BiLSTM-CRF model based on the Transformer architecture is used to construct the initial water conservancy knowledge graph based on the acquired knowledge base source data. The pre-trained language model BERT is used to pre-encode the text and capture the deep semantic features of the text; the bidirectional long short-term memory network BiLSTM is used to process the output of the pre-trained language model BERT, using its bidirectional characteristics to better capture contextual information; the conditional random field model CRF is used to determine the optimal label sequence in the sequence labeling task.
[0014] The above technical solution adopts the water conservancy knowledge graph data layer, concept layer and instance layer construction method of multi-business knowledge integration to construct a knowledge graph of typical business fields of water conservancy projects, and extract the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction businesses, so as to combine the advantages of knowledge graphs and large models in knowledge representation and knowledge analysis to build a large model of water conservancy project knowledge graphs.
[0015] As a further technical solution, we use the self-iterative knowledge graph completion technology of typical water conservancy engineering business areas based on a large language model to improve the knowledge graph of typical water conservancy engineering business areas, including:
[0016] Use a generative language model to perform semantic mining on the text description of the knowledge graph, and generate semantic analysis results as newly mined knowledge through autoregression.
[0017] Use the knowledge graph embedding model to embed the newly mined knowledge and the existing knowledge in the knowledge graph, and fuse them through semantic similarity and structural matching between the new and old knowledge;
[0018] Through continuous iteration, knowledge is extracted from new data and the knowledge graph is updated.
[0019] As a further technical solution, based on the retrieval enhancement generation technology of knowledge graph, entities, relationships and their attributes are automatically extracted from structured and unstructured knowledge bases to create a large model of water conservancy engineering knowledge graph, including:
[0020] Use graph neural networks to learn embedded representations of knowledge graphs and retrieve subgraphs or knowledge fragments related to the input question from the knowledge graph;
[0021] The retrieved subgraph or knowledge fragment is fed into the Transformer-based reasoning model together with the input question to generate text output.
[0022] As a further technical solution, the method further includes:
[0023] The knowledge graph is represented as a graph structure and input into the GraphRAG-GCN fusion model. After the graph structure is preprocessed by GraphRAG technology, GCN is used to perform multi-layer graph convolution operations to gradually aggregate neighborhood information to form a high-quality feature representation;
[0024] Based on the formed high-quality feature representation, we use graph neural networks to learn embedded representations of knowledge graphs, and introduce an attention mechanism in the learning process.
[0025] Based on the features learned after embedding representation, a contrastive learning method is introduced in the pre-training stage of the Transformer-based inference model to screen out knowledge graph data related to the pre-training task.
[0026] As a further technical solution, we use the filtered knowledge graph data and the Transformer-based reasoning model to perform reasoning, including:
[0027] Perform semantic analysis on user questions and convert them into query vectors;
[0028] Based on SPARQL and graph topology information, GraphRAG optimizes subgraph retrieval of knowledge graphs to find subgraphs relevant to the query;
[0029] The subgraph information is encoded into a feature vector and input into the Transformer-based reasoning model together with the query vector. The Transformer's self-attention mechanism is used to process the input information and generate the reasoning result.
[0030] According to one aspect of the present invention, a multi-objective collaborative scheduling method based on the water conservancy project knowledge graph large model is provided, comprising:
[0031] Obtain scheduling requirements;
[0032] The scheduling requirements are input into the constructed water conservancy project knowledge graph model, and with the help of Langchain's multi-objective optimization tool, a multi-objective collaborative scheduling plan for water conservancy projects is generated.
[0033] According to one aspect of the present invention, a device for constructing a large model of a water conservancy project knowledge graph is provided, comprising:
[0034] The first main module is used to construct an initial water conservancy knowledge graph based on the acquired knowledge base source data;
[0035] The second main module is used to introduce the generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph and obtain the optimized water conservancy knowledge graph;
[0036] The third main module is used to perform knowledge graph embedded learning based on the optimized water conservancy knowledge graph using graph neural networks, and introduces attention mechanisms and contrastive learning methods to filter out knowledge graph data related to pre-training tasks;
[0037] The fourth main module is used to perform reasoning based on the filtered knowledge graph data and the Transformer-based reasoning model to obtain a large model of the water conservancy engineering knowledge graph that deeply integrates the knowledge graph and the large language model.
[0038] Furthermore, in the large model of the water conservancy project knowledge graph, based on the water conservancy project management business knowledge base, a water conservancy knowledge graph data layer, concept layer and instance layer construction method integrating multiple business knowledge is adopted to construct a knowledge graph of typical business fields of water conservancy projects, and extract the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction businesses. Among them, the knowledge graph of typical business fields of water conservancy projects is designed to improve the completeness and accuracy of the knowledge graph of typical business fields of water conservancy projects through the self-iterative water conservancy business knowledge graph completion technology based on a large language model, targeting the characteristics of the complex knowledge structure and rich potential associations of water conservancy business.
[0039] According to one aspect of the present invention, a multi-objective collaborative scheduling device based on the water conservancy project knowledge graph large model is provided, comprising:
[0040] Data acquisition module, used to obtain scheduling requirements;
[0041] The collaborative scheduling module is used to input the scheduling requirements into the constructed water conservancy project knowledge graph model, and generate a multi-objective collaborative scheduling plan for water conservancy projects with the help of Langchain's multi-objective optimization tool.
[0042] According to one aspect of the present invention, an execution device is provided, including a processor and a memory, wherein the processor is coupled to the memory; the memory is used to store programs; the processor is used to execute the programs in the memory, so that the execution device executes the method for constructing a large model of water conservancy project knowledge graph or the multi-objective collaborative scheduling method based on the large model of water conservancy project knowledge graph.
[0043] According to one aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method for constructing a large model of a water conservancy project knowledge graph or the multi-objective collaborative scheduling method based on a large model of a water conservancy project knowledge graph.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention combines the advantages of knowledge graphs and large models in knowledge representation and knowledge analysis to construct a large model of water conservancy project knowledge graphs, breaking through the technical bottlenecks of "multi-target collaborative reasoning" and "intelligent decision-making" of large models in typical business fields of water conservancy projects; constructing a large model of water conservancy project knowledge graphs can not only capture the high-order relationships of multimodal information such as Euclidean data and non-Euclidean data in the water conservancy knowledge base, but also support the reasoning and intelligent decision-making of multi-target collaborative solutions such as flood control scheduling, power generation scheduling, water supply scheduling, dam safety monitoring, and unit fault prediction of water conservancy projects, promote the intelligent transformation of the water conservancy and hydropower industry, and develop new quality productivity of water conservancy. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction will be given below to the drawings used in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 A flow chart of a method for constructing a large model of a water conservancy engineering knowledge graph provided by an exemplary embodiment is shown.
[0048] Figure 2 A flowchart of a multi-objective collaborative scheduling method based on a large model of a water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0049] Figure 3 A logical principle diagram of a method for constructing a large model of a water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0050] Figure 4 A block diagram of a device for constructing a large model of a water conservancy engineering knowledge graph provided by an exemplary embodiment is shown.
[0051] Figure 5 A block diagram of a multi-objective collaborative scheduling device based on a large model of a water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0052] Figure 6 A block diagram of an execution device provided by an exemplary embodiment is shown. DETAILED DESCRIPTION
[0053] The terms "including" and "having" and any variations thereof in the description and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0054] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily need to be executed in the order described. For example, some operations / steps may be further decomposed, while others may be combined or partially combined, so the actual execution order may vary depending on the actual situation.
[0055] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention are arbitrarily combined with each other to form a new technical solution. This combination is not restricted by the sequence of steps and / or structural composition mode, but must be based on the ability of ordinary technicians in this field to implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that this combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0056] Please refer to Figure 1 ,in Figure 1 A flow chart of a method for constructing a large model of a water conservancy project knowledge graph provided by an exemplary embodiment is shown. The method includes:
[0057] Step 1: Based on the water conservancy project management business knowledge base, a water conservancy knowledge graph data layer, concept layer and instance layer construction method that integrates multiple business knowledge is adopted to construct a knowledge graph of typical business fields of water conservancy projects, and refine the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction businesses.
[0058] In step 1, the water conservancy project management business knowledge base source data obtained includes water conservancy professional data, scientific papers, journal documents, standards and specifications, laws and regulations, images, audio voiceprints, etc.
[0059] The construction of a knowledge graph for typical water conservancy engineering business areas involves employing the BERT-BiLSTM-CRF model based on the Transformer architecture to construct the data, concept, and instance layers of the water conservancy knowledge graph. BERT is used to pre-encode text and capture its deep semantic features; BiLSTM further processes BERT's output, leveraging its bidirectional nature to better capture contextual information; and CRF determines the optimal label sequence for sequence labeling tasks.
[0060] Text data from the water conservancy project management business knowledge base (such as scientific papers, journal articles, standards and specifications, and laws and regulations) undergoes preprocessing, such as word segmentation and annotation, to be converted into a format acceptable to the model. This data is then fed into the BERT model to obtain a pre-encoded feature representation. These features are then fed into a BiLSTM network, whose forward and backward LSTM layers simultaneously process the sequence, capturing contextual information. Finally, the BiLSTM output is fed into a CRF layer to calculate a score for each label sequence. In the CRF layer, given input features x and label sequence y, the score function S(x, y) is:
[0061]
[0062] Where A is the transfer matrix, Indicates that from the label Transfer to score; The label at position i is Finally, the optimal label sequence is determined by maximizing the score function, thereby extracting entities and relations and converting them into triples.
[0063] Step 2: Use the self-iterative knowledge graph completion technology of typical water conservancy engineering business fields based on the large language model to improve the knowledge graph of typical water conservancy engineering business fields;
[0064] Initial knowledge graph analysis (LLM technology) uses a fine-tuned Qwen model, leveraging its generative capabilities for semantic mining. The textual description of the initial water conservancy knowledge graph is fed into the fine-tuned Qwen model. During fine-tuning, the pre-trained Qwen model is further trained using text data from the water conservancy sector to better adapt it to semantic understanding in this field. The model generates semantic analysis of the knowledge graph using an autoregressive approach, mining potential semantic associations. During fine-tuning, the model parameters θ are updated by minimizing a loss function L. Common loss functions include the cross-entropy loss function:
[0065]
[0066] Where N is the number of samples, C is the number of categories, is the true label (0 or 1) of sample i belonging to category j, is the probability that the model predicts that sample i belongs to category j.
[0067] Using a method based on semantic similarity and structural matching, combined with knowledge graph embedding technology, the newly mined knowledge and the knowledge in the initial knowledge graph are embedded separately. The TransE model is used to map entities and relationships into a low-dimensional vector space. Semantic similarity, such as cosine similarity, is calculated between the new and old knowledge embedding vectors. At the same time, structural information of the knowledge graph, such as node degree and path length, is considered for structural matching. In the TransE model, the loss function for the triple (h, r, t) is:
[0068]
[0069] Where S is the set of true triples, is a set of incorrect triples generated by replacing the head entity h or the tail entity t in the true triples, is the interval parameter. By minimizing the loss function, the embedding vectors of entities and relations are obtained, which are used to calculate semantic similarity and structural matching.
[0070] Step 3: Based on the retrieval enhancement generation technology of knowledge graph, entities, relationships and their attributes are automatically extracted from structured and unstructured knowledge bases to create a large model of water conservancy project knowledge graph to support the reasoning application of multi-objective collaborative solutions for water conservancy projects.
[0071] Step 3 is a deep fusion of the knowledge graph and the large language model: based on the GraphRAG-GCN fusion model in the graph neural network (GNN), multi-layer graph convolution is used to aggregate neighborhood information. Specifically, the knowledge graph is represented as a graph structure and input into the GraphRAG-GCN fusion model. The model first pre-processes the graph structure through GraphRAG technology to enhance the representation of key nodes and relationships. Then, GCN is used to perform multi-layer graph convolution operations to gradually aggregate neighborhood information. Each layer of GCN updates the node features so that the node features contain more structural information. In the GCN layer, the feature update formula for node i in the l+1 layer is:
[0072]
[0073] Where, is the feature vector of node i in layer l, is the set of neighboring nodes of node i, and are the degrees of node i and its neighboring node l, is the weight matrix of layer l, is the bias term, It is a nonlinear activation function. Through multiple iterative updates, it provides high-quality feature representation for the subsequent fusion of knowledge graphs and large language models.
[0074] When GNN learns the embedding representation of the knowledge graph (KG), an attention-guided embedding method is used. In the LLM pre-training stage, a contrastive learning method is used to screen relevant knowledge graph data. In the GNN embedding representation learning, an attention mechanism is introduced to give higher weights to key water conservancy entity relationships (such as the relationship between dams and water flows), guiding GNN to pay more attention to the feature learning of these relationships. In the LLM pre-training stage, the knowledge graph data is divided into positive and negative examples. Through the contrastive learning method, LLM learns the difference between positive examples (relevant knowledge graph data) and negative examples (irrelevant data), thereby screening out knowledge graph data related to the pre-training task. In the attention mechanism, the attention weight of the relationship r is calculated. The formula is:
[0075]
[0076] Where, is the similarity score between the query vector q and the relation r. In contrastive learning, the loss function It can be defined as:
[0077]
[0078] Where, is the similarity between the query vector q and the key vector k, is the positive key vector, is the negative key vector, is the temperature parameter.
[0079] GraphRAG optimizes subgraph retrieval in knowledge graphs based on SPARQL, employing a Transformer-based reasoning model. Specifically, user questions are semantically parsed and converted into query vectors. Based on SPARQL and graph topology information, GraphRAG optimizes subgraph retrieval to find subgraphs relevant to the query. The subgraph information is encoded as a feature vector and fed into a Transformer-based reasoning model along with the query vector. The reasoning model utilizes the Transformer's self-attention mechanism to process the input information and generate inference results. In the Transformer-based reasoning model, the self-attention mechanism performs the following computational steps:
[0080]
[0081] In the formula, Q (Query), K (Key), and V (Value) are the input feature matrices. is the dimension of K.
[0082] Based on the aforementioned large model of water conservancy engineering knowledge graph, the embodiment of the present invention also provides a multi-objective collaborative scheduling method based on the large model of water conservancy engineering knowledge graph, see Figure 2 First, the scheduling requirements are obtained and converted into user questions; then, the user questions are input into the constructed water conservancy project knowledge graph model, and with the help of Langchain's multi-objective optimization tool, a multi-objective collaborative scheduling plan for water conservancy projects is generated.
[0083] The reasoning process of the multi-objective collaborative scheduling solution specifically includes:
[0084] Through the interaction between Langchain and the knowledge graph, constraints from the knowledge graph are incorporated into the reasoning process. In flood control scheduling, Langchain retrieves constraint information from the knowledge graph, such as the reservoir's flood control limit water level and the safe discharge volume of downstream rivers, and automatically verifies during reasoning whether the generated solution meets these constraints. In water supply scheduling, Langchain also relies on information such as residential and agricultural water demand from the knowledge graph to ensure that the reasoning results meet actual needs and safety standards, ensuring the rationality and reliability of decision-making.
[0085] Develop application services specifically for water conservancy projects, such as combining data and optimization algorithms to provide solutions for water project scheduling and using anomaly detection algorithms for early warning in dam safety monitoring. Leverage Langchain to quickly build applications. Use collaborative filtering algorithms to recommend knowledge services based on user behavior, and integrate GraphRAG to optimize recommendations through knowledge graph associations.
[0086] With the knowledge graph model inference engine at its core, Langchain's multi-objective optimization tools combine multiple objectives, such as flood control, water supply, and power generation, to generate a water project scheduling and operation model. Taking into account the weights of objectives such as power generation efficiency and downstream water supply security, and utilizing built-in multi-objective optimization algorithms such as the weighted sum method and the ε-constraint method, the inference engine generates an optimal scheduling plan that takes into account the needs of all parties.
[0087] For example, using the weighted summation method to combine multiple objective functions Linear combination into a comprehensive objective function , the formula is:
[0088]
[0089] Where, is the weight of the i-th objective function and satisfies , For example, let the power generation efficiency objective function be , the objective function of water supply security is , by reasonably allocating weights and , so that the comprehensive objective function It can achieve the optimal solution under certain constraints, thereby generating a scheduling plan that takes into account both power generation and water supply needs.
[0090] Preferably, the construction process of the large model of the water conservancy engineering knowledge graph provided by the embodiment of the present invention can be implemented through steps 2.1 to 2.2.
[0091] In step 2.1, first install the knowledge graph-based retrieval enhancement generation package graphrag based on the Python package.
[0092] pip install graphrag;
[0093] Then install the existing water knowledge base Waterknowledgebase and initialize the project
[0094] Pip install Waterknowledgebase
[0095] Waterknowledgebase install;
[0096] Then create a temporary folder graphrag to store running data.
[0097] mkdir . / graphrag / input;
[0098] curl https: / / www.xxx.com / xxx.txt >. / myTest / input / watersupply.txt / / Here is the sample code, which can be set according to the business to be tested, such as the txt text of water supply scheduling.
[0099] cd . / graphrag;
[0100] python -m graphrag.index –init;
[0101] At this point, the source data acquisition of the water conservancy knowledge base has been completed, including the import of water conservancy professional data, scientific papers, journal documents, standards and specifications, laws and regulations, images, audio voiceprints and other data.
[0102] In step 2.2, first modify the configuration file, modify the settings.yam file, and modify the large language model LLM to be used and the corresponding API_base, for example:
[0103] Encoding_model: cl100k_base;
[0104] Skip_workflows: [ ];
[0105] LLM:
[0106] API_key: ${GRAPHRAG_API_KEY};
[0107] Type: openai_chat # or azure_openai_chat;
[0108] Model: deepseek-chat #Customize your own large model;
[0109] Model_supports_json: false # recommend if this is available for your model;
[0110] API_base: http: / / api.agicto.cn / v1 #Domestic open source large model;
[0111] The modified configuration file is a global configuration, and subsequent calls to the model do not need to be modified again.
[0112] Next, we build an index for the knowledge graph-based retrieval enhancement software package:
[0113] Waterknowledgebase run wkb index –root . # Waterknowledgebase is the water conservancy knowledge base;
[0114] Finally, global query or local query is used to realize water conservancy project knowledge retrieval.
[0115] For example, global query focuses on retrieval of joint water supply scheduling information of river basin reservoir groups.
[0116] Waterknowledgebase run wkb query –root . --method global “Reservoir group joint water supply scheduling information”
[0117] For example, local query: focuses more on retrieving water supply scheduling information of a single reservoir
[0118] Waterknowledgebase run wkb query –root . --method local “Water supply scheduling information of each reservoir”.
[0119] Please refer to Figure 3 , Figure 3 This figure shows the logical principle of a method for constructing a large model of a water conservancy project knowledge graph, provided by one embodiment. This method supports intelligent water conservancy business operations through this inference logic, and supports the application of reasoning for multi-objective collaborative solutions such as flood control, water supply, and power generation in water conservancy projects.
[0120] Please refer to Figure 4 , which shows a block diagram of the structure of the water conservancy engineering knowledge graph large model construction device provided by this application, and the device includes: a first main module 102, which is used to construct an initial water conservancy knowledge graph based on the acquired knowledge base source data; a second main module 104, which is used to introduce a generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph to obtain an optimized water conservancy knowledge graph; a third main module 106, which is used to use a graph neural network to perform knowledge graph embedded learning based on the optimized water conservancy knowledge graph, and introduce an attention mechanism and a comparative learning method to screen out knowledge graph data related to the pre-training task; a fourth main module 108, which is used to perform reasoning based on the screened knowledge graph data in combination with a Transformer-based reasoning model to obtain a large model of the water conservancy engineering knowledge graph with deep integration of the knowledge graph and the large language model.
[0121] Please refer to Figure 5 , which shows a block diagram of the structure of the water conservancy engineering knowledge graph large model construction device provided by this application, and the model includes:
[0122] The data acquisition module 202 is used to obtain scheduling requirements, including water conservancy professional data, scientific papers, journals, standards and specifications, laws and regulations, images, audio and voiceprints, etc.
[0123] The collaborative scheduling module 204 is used to input the scheduling requirements into the constructed water conservancy project knowledge graph model, and generate a multi-objective collaborative scheduling plan for the water conservancy project with the help of Langchain's multi-objective optimization tool.
[0124] Current large-scale water conservancy models are developed to address the goal of building a "digital twin river basin." They primarily address the needs of the river basin and the broader water conservancy and hydropower sectors, and are unable to support typical (Level 3) applications such as dam safety monitoring and unit failure prediction for water conservancy projects. The knowledge engine of current large-scale water conservancy models, driven by large language models (LLMs), suffers from weak knowledge reasoning capabilities and is unable to support the current state of multi-objective collaborative solution reasoning and intelligent decision-making for water conservancy projects. By employing several modules in the figure, it is not only possible to capture high-order relationships across multimodal information, including Euclidean and non-Euclidean data, within the water conservancy knowledge base, but also to support multi-objective collaborative solution reasoning and intelligent decision-making for water conservancy projects, such as flood control scheduling, power generation scheduling, water supply scheduling, dam safety monitoring, and unit failure prediction. This will promote the intelligent transformation of the water conservancy and hydropower industries and develop new high-quality water conservancy productivity.
[0125] It should be noted that the device embodiments provided by the present invention are not only used to implement the methods in the above-mentioned method embodiments, but also used to implement the methods in other method embodiments provided by the present invention. The only difference is that the corresponding functional modules are set, and the principles are basically the same as the principles of the above-mentioned device embodiments provided by the present invention. As long as those skilled in the art refer to the specific technical solutions in other method embodiments on the basis of the above-mentioned device embodiments, obtain the corresponding technical means and the technical solutions composed of these technical means by combining technical features, and on the premise of ensuring the practicality of the technical solutions, improve the equipment in the above-mentioned device embodiments to obtain corresponding device embodiments for implementing the methods in other method embodiments.
[0126] Since the construction of a large model of water conservancy project knowledge graph and multi-objective collaborative scheduling methods have been discussed in detail in the previous method, the specific implementation of the above modules will not be repeated here.
[0127] Next, an execution device provided by an embodiment of the present invention is introduced. Figure 6 , Figure 6 This is a schematic diagram of a structure of an execution device provided by an embodiment of the present invention. The execution device 300 can be specifically manifested as an autonomous driving vehicle, a mobile phone, a tablet, a laptop computer, a desktop computer, a monitoring data processing device, etc., which is not limited here. Figure 1 The execution device 300 includes a receiver 301, a transmitter 302, a processor 303, and a memory 304 (the number of processors 303 in the execution device 300 can be one or more, Figure 6(taking one processor as an example), the processor 303 may include an application processor 3031 and a communication processor 3032. In some embodiments of the present application, the receiver 301, the transmitter 302, the processor 303 and the memory 304 may be connected via a bus or other means.
[0128] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 303. A portion of the memory 304 may also include non-volatile random access memory (NVRAM). The memory 304 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0129] Processor 303 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all of these buses are referred to as a bus system in the figure.
[0130] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 303. Processor 303 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 303. The above processor 303 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller. It can also include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 303 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 304, and processor 303 reads the information in memory 304 and performs the steps of the above method in conjunction with its hardware.
[0131] Receiver 301 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 302 can be used to output digital or character information through the first interface. Transmitter 302 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 302 can also include a display device such as a display screen.
[0132] In the embodiment of the present invention, the processor 303 is configured to execute Figure 1 The specific manner in which the application processor 3031 in the processor 303 performs the above steps is the same as that in the present application. Figure 1 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figure 1 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0133] Based on the same technical concept as the above-mentioned embodiment, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the water conservancy project knowledge graph large model construction method or the multi-objective collaborative scheduling method based on the water conservancy project knowledge graph large model.
[0134] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0136] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0138] In summary, the present invention is based on the water conservancy project management business knowledge base, and adopts a water conservancy knowledge graph data layer, concept layer and instance layer construction method that integrates multiple business knowledge to construct a knowledge graph of typical water conservancy project business fields, and extract the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction business; in view of the complex knowledge structure and rich potential associations of water conservancy business, the self-iterative water conservancy business knowledge graph completion technology based on the large language model is used to improve the completeness and accuracy of the knowledge graph of typical water conservancy business fields; in order to efficiently combine the advantages of knowledge graphs and large models in knowledge representation and knowledge analysis, the retrieval enhancement generation technology based on knowledge graphs is adopted to automatically extract entities, relationships and their attributes from structured and unstructured knowledge bases, and create a large model of water conservancy project knowledge graphs to support the reasoning application of multi-objective collaborative solutions such as flood control, water supply, and power generation of water conservancy projects.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a large model of a water conservancy engineering knowledge graph, characterized in that: include: Based on the water conservancy project management business knowledge base, a water conservancy knowledge graph data layer, concept layer and instance layer construction method that integrates multiple business knowledge is adopted to build a knowledge graph of typical water conservancy project business areas, and refine the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction business; Improve the knowledge graph of typical water conservancy engineering business areas using self-iterative knowledge graph completion technology based on a large language model; Retrieval-enhanced generation technology based on knowledge graphs automatically extracts entities, relationships, and their attributes from structured and unstructured knowledge bases to create a large water conservancy project knowledge graph model to support the application of multi-objective collaborative solution reasoning for water conservancy projects. This includes: using graph neural networks to learn embedded representations of knowledge graphs and retrieving subgraphs or knowledge fragments related to the input problem from the knowledge graph; The retrieved subgraph or knowledge fragment is input into the Transformer-based reasoning model together with the input question to generate text output; the knowledge graph is represented as a graph structure and input into the GraphRAG-GCN fusion model. After the graph structure is preprocessed by GraphRAG technology, multi-layer graph convolution operations are performed using GCN to gradually aggregate neighborhood information to form a high-quality feature representation; based on the formed high-quality feature representation, the knowledge graph is embedded in the representation using a graph neural network, and an attention mechanism is introduced in the learning process; based on the features after the embedded representation learning, the contrastive learning method is introduced in the pre-training stage of the Transformer-based reasoning model to screen out knowledge graph data related to the pre-training task.
2. The method for constructing a large model of a water conservancy project knowledge graph according to claim 1 is characterized in that: Construct a knowledge graph of typical water conservancy engineering business areas, including: The BERT-BiLSTM-CRF model based on the Transformer architecture is used to construct the initial water conservancy knowledge graph based on the acquired knowledge base source data. The pre-trained language model BERT is used to pre-encode the text and capture the deep semantic features of the text; the bidirectional long short-term memory network BiLSTM is used to process the output of the pre-trained language model BERT, using its bidirectional characteristics to better capture contextual information; the conditional random field model CRF is used to determine the optimal label sequence in the sequence labeling task.
3. The method for constructing a large model of a water conservancy project knowledge graph according to claim 1 is characterized in that: The self-iterative knowledge graph completion technology for typical water conservancy engineering business areas based on a large language model is used to improve the knowledge graph for typical water conservancy engineering business areas, including: Use a generative language model to perform semantic mining on the text description of the knowledge graph, and generate semantic analysis results as newly mined knowledge through autoregression. Use the knowledge graph embedding model to embed the newly mined knowledge and the existing knowledge in the knowledge graph, and fuse them through semantic similarity and structural matching between the new and old knowledge; Through continuous iteration, knowledge is extracted from new data and the knowledge graph is updated.
4. A multi-objective collaborative scheduling method based on the water conservancy project knowledge graph model according to any one of claims 1 to 3, characterized in that: include: Obtain scheduling requirements; The scheduling requirements are input into the constructed water conservancy project knowledge graph model, and with the help of Langchain's multi-objective optimization tool, a multi-objective collaborative scheduling plan for water conservancy projects is generated.
5. A large-scale model construction device for a water conservancy engineering knowledge graph, characterized in that: include: The first main module is used to construct an initial water conservancy knowledge graph based on the acquired knowledge base source data; The second main module is used to introduce the generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph and obtain the optimized water conservancy knowledge graph; The third main module is used to perform knowledge graph embedded learning based on the optimized water conservancy knowledge graph using graph neural networks, and introduces attention mechanisms and contrastive learning methods to filter out knowledge graph data related to pre-training tasks; The fourth main module is used to perform reasoning based on the filtered knowledge graph data and the Transformer-based reasoning model to obtain a large water conservancy project knowledge graph model that deeply integrates the knowledge graph and the large language model. This includes: using graph neural networks to learn embedded representations of the knowledge graph and retrieving subgraphs or knowledge fragments related to the input question from the knowledge graph; The retrieved subgraph or knowledge fragment is input into the Transformer-based reasoning model together with the input question to generate text output; the knowledge graph is represented as a graph structure and input into the GraphRAG-GCN fusion model. After the graph structure is preprocessed by GraphRAG technology, multi-layer graph convolution operations are performed using GCN to gradually aggregate neighborhood information to form a high-quality feature representation; based on the formed high-quality feature representation, the knowledge graph is embedded in the representation using a graph neural network, and an attention mechanism is introduced in the learning process; based on the features after the embedded representation learning, the contrastive learning method is introduced in the pre-training stage of the Transformer-based reasoning model to screen out knowledge graph data related to the pre-training task.
6. The multi-objective collaborative scheduling device based on the water conservancy project knowledge graph model according to claim 5 is characterized in that: include: Data acquisition module, used to obtain scheduling requirements; The collaborative scheduling module is used to input the scheduling requirements into the constructed water conservancy project knowledge graph model, and generate a multi-objective collaborative scheduling plan for water conservancy projects with the help of Langchain's multi-objective optimization tool.
7. An execution device, characterized in that: It includes a processor and a memory, the processor is coupled to the memory; the memory is used to store programs; the processor is used to execute the programs in the memory, so that the execution device executes the water conservancy project knowledge graph large model construction method as described in any one of claims 1 to 3 or the multi-objective collaborative scheduling method based on the water conservancy project knowledge graph large model as described in claim 4.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the water conservancy project knowledge graph large model construction method described in any one of claims 1 to 3 or the multi-objective collaborative scheduling method based on the water conservancy project knowledge graph large model described in claim 4.
Citation Information
Patent Citations
Water conservancy knowledge service system based on knowledge graph technology
CN118152525A