Water conservancy project knowledge graph large model construction method and device
By building a big model of water conservancy engineering knowledge graph, combining knowledge graph and large language model, the problem of insufficient knowledge reasoning capabilities of water conservancy models is solved, multi-objective collaborative solution reasoning and intelligent decision-making are achieved, and intelligent transformation of the water conservancy and hydropower industry is promoted.
Patent Information
- Application Number
- CN202510847148.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing water conservancy model has weak knowledge reasoning capabilities and cannot support the multi-objective collaborative solution reasoning and intelligent decision-making of water conservancy projects, especially inadequate applications in typical business areas such as dam safety monitoring and unit failure prediction.
Build a big model of knowledge graph of water conservancy engineering, and use the data layer, concept layer and example layer of multi-business knowledge fusion, combined with the self-iteration technology of the large language model, to automatically extract entities and relationships, and use graph neural network and Transformer model for deep fusion to achieve the combination of the advantages of knowledge graph and large model.
It has broken through the bottlenecks of multi-objective collaborative reasoning and intelligent decision-making in typical business areas of water conservancy engineering, and can support multi-objective collaborative solution reasoning and intelligent decision-making such as flood control scheduling, power generation scheduling, water supply scheduling, dam safety monitoring, unit failure prediction, and promote the intelligent transformation of the water conservancy and hydropower industry.
Smart Images

Figure CN120372019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital twin engineering, and particularly to a method and device for constructing a large model of a water conservancy project knowledge graph. Background Art
[0002] The water conservancy large model is oriented to the needs of the water conservancy and hydropower industry. It adopts the "1+2+N" paradigm framework, uses a knowledge base composed of structured and unstructured data as the engine, and takes mechanism models and artificial intelligence models as the core, targeting water conservancy and hydropower business scenarios. It mainly realizes the intelligent empowerment of water conservancy projects and serves the standardization and systematization of high-quality water conservancy development. Here, "1" refers to the water conservancy professional knowledge base, "2" refers to mechanism models and artificial intelligence models, and "N" refers to multiple business scenarios. Following the development trend of artificial intelligence technology and large language models, two universities and research institutions in China have taken the lead in launching water conservancy and hydropower industry large model construction projects. The Department of Hydraulic Engineering of Tsinghua University has formed an interdisciplinary research team and launched a large model project for the hydropower industry based on the independently developed GLM large model in April 2024. The China Institute of Water Resources and Hydropower Research officially released the independently developed large model of water science and technology "SkyLIM" in May 2024. These large models will strongly promote the intelligent transformation of the water conservancy and hydropower industry and develop new water conservancy productive forces.
[0003] Current water conservancy large models are all developed in response to the construction goal of the "digital twin basin". They are mainly oriented to the needs of the basin level and the broad business fields of water conservancy and hydropower. The model data floor can effectively support the application of broad business (L2 level) fields such as water disasters, water resources, water environment, and water ecology, but cannot support the application of typical business (L3 level) fields such as dam safety monitoring and unit fault prediction in water conservancy projects. The knowledge engine of the current water conservancy large model is driven by a large language model (LLM), which performs excellently in knowledge representation, extraction, and fusion, but has the problem of weak knowledge reasoning ability and cannot support the reasoning of multi-objective collaborative schemes and intelligent decision-making in water conservancy projects. The knowledge graph represents the information in the knowledge base as a network of entities (nodes) and relationships (edges), imitating the composition method of human structural knowledge. It can not only capture the basic information of the knowledge base but also capture the high-order relationships of multi-modal information such as Euclidean data and non-Euclidean data in the knowledge base, and has strong reasoning ability.
[0004] In view of this, under the urgent need for the intelligent transformation and upgrading of water conservancy, how to use advanced technical means such as knowledge graphs and large language models to propose a deep fusion model of knowledge graphs and large language models to break through the technical bottleneck of intelligent decision-making in water conservancy projects is the current direction that urgently needs technical research. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a method and device for constructing a large model of a water conservancy project knowledge graph, which can break through the technical bottlenecks of "multi-objective collaborative reasoning" and "intelligent decision-making" of large models in typical business fields of water conservancy projects, promote the intelligent transformation of the water conservancy and hydropower industry, and develop new water conservancy productive forces.
[0006] According to one aspect of the specification of the present invention, there is provided a method for constructing a large model of a water conservancy project knowledge graph, including: Based on the knowledge base of water conservancy project management business, adopt the construction methods of the data layer, concept layer and instance layer of the water conservancy knowledge graph with multi-business knowledge integration to construct the knowledge graph of typical business fields of water conservancy projects, and refine the logical chain among the water conservancy project scheduling, dam safety monitoring, and unit fault prediction services; Utilize the self-iterative water conservancy project typical business field knowledge graph completion technology based on the large language model to improve the knowledge graph of typical business fields of water conservancy projects; Based on the retrieval-enhanced generation technology of the knowledge graph, automatically extract entities, relationships, and their attributes from structured and unstructured knowledge bases, and create a large model of the water conservancy project knowledge graph to support the reasoning application of multi-objective collaborative solutions for water conservancy projects.
[0007] The above technical solution adopts the retrieval-enhanced generation technology based on the knowledge graph to automatically extract entities, relationships, and their attributes from structured and unstructured knowledge bases, create a large model of the water conservancy project knowledge graph for flood control, water supply, power generation scheduling, etc. The constructed large model of the water conservancy project knowledge graph combines the advantages of the knowledge graph and the large model in knowledge representation and knowledge parsing to realize the reasoning application of multi-objective collaborative solutions for flood control, water supply, power generation, etc. in water conservancy projects.
[0008] Aiming at the characteristics of complex water conservancy business knowledge structure and rich potential associations, the above technical solution improves the completeness and accuracy of the knowledge graph of typical water conservancy business fields through the self-iterative water conservancy business knowledge graph completion technology based on the large language model.
[0009] As a further technical solution, constructing the knowledge graph of typical business fields of water conservancy projects includes: Adopt the BERT-BiLSTM-CRF model based on the Transformer architecture to construct the initial water conservancy knowledge graph based on the obtained knowledge base source data. Among them, the pre-trained language model BERT is used to pre-encode the text to capture the deep semantic features of the text; the bidirectional long short-term memory network BiLSTM is used to process the output of the pre-trained language model BERT, and its bidirectional characteristics are used to better capture the context information; the conditional random field model CRF is used to determine the optimal label sequence in the sequence labeling task.
[0010] The above technical solution adopts a construction method for the data layer, concept layer, and instance layer of a water conservancy knowledge graph that integrates multiple business knowledges, constructs a knowledge graph for typical business fields of water conservancy projects, and extracts the logical chain among the business of water conservancy project scheduling, dam safety monitoring, and unit fault prediction, so as to combine the advantages of the knowledge graph and the large model in knowledge representation and knowledge parsing to construct a large model of the water conservancy project knowledge graph.
[0011] As a further technical solution, a self-iterative knowledge graph completion technology for typical business fields of water conservancy projects based on a large language model is used to improve the knowledge graph for typical business fields of water conservancy projects, including: Use a generative language model to perform semantic mining on the text description of the knowledge graph, and generate semantic analysis results in an autoregressive manner as newly mined knowledge; Use a knowledge graph embedding model to perform embedding representations on the newly mined knowledge and the existing knowledge in the knowledge graph respectively, and fuse them through the semantic similarity and structural matching between the new and old knowledge; Through continuous iteration, extract knowledge from new data and update the knowledge graph.
[0012] As a further technical solution, a retrieval-enhanced generation technology based on the knowledge graph is used to automatically extract entities, relationships, and their attributes from structured and unstructured knowledge bases to create a large model of the water conservancy project knowledge graph, including: Use a graph neural network to perform embedding representation learning on the knowledge graph, and retrieve subgraphs or knowledge fragments related to the input question from the knowledge graph; Input the retrieved subgraph or knowledge fragment together with the input question into a reasoning model based on Transformer to generate text output.
[0013] As a further technical solution, the method further includes: Represent the knowledge graph as a graph structure and input it into the GraphRAG-GCN fusion model. After preprocessing the graph structure through GraphRAG technology, use GCN to perform multi-layer graph convolution operations to gradually aggregate neighborhood information to form high-quality feature representations; Based on the formed high-quality feature representations, use a graph neural network to perform embedding representation learning on the knowledge graph, and introduce an attention mechanism during the learning process; Based on the features after embedding representation learning, introduce a contrastive learning method in the pre-training stage of the reasoning model based on Transformer to screen out the knowledge graph data related to the pre-training task.
[0014] As a further technical solution, based on the screened knowledge graph data, perform reasoning in combination with a reasoning model based on Transformer, including: Semantically parse the user's question and convert the question into a query vector; Based on SPARQL and graph topology structure information, GraphRAG optimizes the subgraph retrieval of the knowledge graph to find subgraphs related to the query; Encode the subgraph information into a feature vector and input it together with the query vector into a Transformer-based inference model. Use the self-attention mechanism of Transformer to process the input information and generate an inference result.
[0015] According to one aspect of the specification of the present invention, there is provided a multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph, including: Obtain scheduling requirements; Input the scheduling requirements into the constructed large model of the water conservancy project knowledge graph, and with the help of the multi-objective optimization tool of Langchain, generate a multi-objective collaborative scheduling plan for the water conservancy project.
[0016] According to one aspect of the specification of the present invention, there is provided a device for constructing a large model of a water conservancy project knowledge graph, including: The first main module is used to construct an initial water conservancy knowledge graph based on the obtained knowledge base source data; The second main module is used to introduce a generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph to obtain an optimized water conservancy knowledge graph; The third main module is used to perform knowledge graph embedded learning based on the optimized water conservancy knowledge graph using a graph neural network, and introduce an attention mechanism and a contrast learning method to screen out knowledge graph data related to the pre-training task; The fourth main module is used to perform inference based on the screened knowledge graph data in combination with a Transformer-based inference model to obtain a large model of the water conservancy project knowledge graph with deep integration of the knowledge graph and the large language model.
[0017] Furthermore, in the large model of the water conservancy project knowledge graph, based on the water conservancy project management business knowledge base, a construction method for the data layer, concept layer, and instance layer of the water conservancy knowledge graph with multi-business knowledge integration is adopted to construct a knowledge graph for typical business fields of water conservancy projects, and the logical chain between water conservancy project scheduling, dam safety monitoring, and unit fault prediction services is refined. Among them, the knowledge graph for typical business fields of water conservancy projects aims to improve the completeness and accuracy of the knowledge graph for typical business fields of water conservancy through the self-iterative water conservancy business knowledge graph completion technology based on the large language model in view of the characteristics of complex water conservancy business knowledge structure and rich potential associations.
[0018] According to one aspect of the specification of the present invention, there is provided a multi-objective collaborative scheduling device based on the large model of the water conservancy project knowledge graph, including: A data acquisition module for acquiring scheduling requirements; A collaborative scheduling module for inputting the scheduling requirements into a constructed large model of a water conservancy project knowledge graph and generating a multi-objective collaborative scheduling scheme for water conservancy projects by means of a multi-objective optimization tool of Langchain.
[0019] According to one aspect of the specification of the present invention, an execution device is provided, including a processor and a memory, the processor being coupled to the memory; the memory for storing a program; the processor for executing the program in the memory such that the execution device executes the method for constructing a large model of a water conservancy project knowledge graph or the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph.
[0020] According to one aspect of the specification of the present invention, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the method for constructing a large model of a water conservancy project knowledge graph or the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention combines the advantages of knowledge graphs and large models in knowledge representation and knowledge parsing to construct a large model of a water conservancy project knowledge graph, breaking through the technical bottlenecks of "multi-objective collaborative reasoning" and "intelligent decision-making" of large models in typical business fields of water conservancy projects; constructing a large model of a water conservancy project knowledge graph can not only capture the high-order relationships of multi-modal information such as Euclidean data and non-Euclidean data in the water conservancy knowledge base, but also support the reasoning of multi-objective collaborative schemes and intelligent decision-making for flood control scheduling, power generation scheduling, water supply scheduling, dam safety monitoring, unit fault prediction, etc. of water conservancy projects, promoting the intelligent transformation of the water conservancy and hydropower industry and developing new water conservancy productive forces. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0023] Figure 1 The flowchart of the method for constructing a large model of a water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0024] Figure 2 The flowchart of the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0025] Figure 3 The logic schematic diagram of the construction method of the large model of the water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0026] Figure 4 The block diagram of the composition of the device for constructing the large model of the water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0027] Figure 5 The block diagram of the composition of the multi-objective collaborative scheduling device based on the large model of the water conservancy project knowledge graph provided by an exemplary embodiment is shown.
[0028] Figure 6 The block diagram of the composition of the execution device provided by an exemplary embodiment is shown. Detailed implementation manners
[0029] The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. The flowcharts shown in the drawings are only exemplary descriptions and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. Additionally, the technical features in each embodiment or individual embodiment provided by the present invention can be combined arbitrarily with each other to form a new technical solution. This combination is not restricted by the order of steps and / or the structural composition mode, but must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or is unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0032] Please refer to Figure 1 , in which Figure 1 shows a schematic flowchart of a method for constructing a large knowledge graph of water conservancy projects provided by an exemplary embodiment. The method includes: Step 1: Based on the knowledge base of water conservancy project management business, adopt the construction methods of the data layer, concept layer, and instance layer of the water conservancy knowledge graph with multi-business knowledge integration to construct the knowledge graph of typical business fields of water conservancy projects, and refine the logical chain among water conservancy project scheduling, dam safety monitoring, and unit fault prediction services.
[0033] In the above Step 1, the source data obtained from the knowledge base of water conservancy project management business includes water conservancy professional data, scientific treatises, journal articles, standards and specifications, laws and regulations, images, audio voiceprints, etc.
[0034] The construction of the knowledge graph of typical business fields of water conservancy projects specifically includes: adopting the BERT-BiLSTM-CRF model based on the Transformer architecture to construct the data layer, concept layer, and instance layer of the water conservancy knowledge graph. BERT is used to pre-encode the text and capture the deep semantic features of the text; BiLSTM further processes the output of BERT and better captures the context information by using its bidirectional characteristics; CRF determines the optimal label sequence in the sequence labeling task.
[0035] Preprocess the text data (such as scientific treatises, journal articles, standards and specifications, laws and regulations) in the water conservancy project management business knowledge base through word segmentation, annotation, etc., and convert it into a format acceptable to the model. The data is input into the BERT model to obtain pre-encoded feature representations. These features are then input into the BiLSTM network, where the forward and backward LSTM layers process the sequence simultaneously to capture context information. Finally, the output of the BiLSTM is fed into the CRF layer to calculate the score of each label sequence. In the CRF layer, given the input feature x and the label sequence y, the scoring function S(x, y) is:
[0036] where A is the transition matrix, represents the score for transitioning from label to ; is the emission score for the label at position i. Finally, the optimal label sequence is determined by maximizing the scoring function, thereby extracting entities and relationships and converting them into triples.
[0037] Step 2: Use the self-iterative water conservancy project typical business domain knowledge graph completion technology based on large language models to improve the water conservancy project typical business domain knowledge graph; Initial knowledge graph analysis (LLM technology), using the fine-tuned Qwen model and leveraging its generative ability for semantic mining. Input the text description of the initial water conservancy knowledge graph into the fine-tuned Qwen model. During the fine-tuning process, use the text data in the water conservancy field to further train the pre-trained Qwen model to better adapt to the semantic understanding in the water conservancy field. The model generates the analysis results of the knowledge graph semantics through an autoregressive manner to mine potential semantic associations. During the fine-tuning process, update the model parameters θ by minimizing the loss function L. Commonly used loss functions such as the cross-entropy loss function:
[0038] where N is the number of samples, C is the number of classes, is the true label (0 or 1) of sample i belonging to class j, is the probability that the model predicts sample i belongs to class j.
[0039] Using a method based on semantic similarity and structure matching, combined with knowledge graph embedding technology, the newly mined knowledge and the knowledge in the initial knowledge graph are respectively embedded and represented, and the TransE model is used to map entities and relationships into a low-dimensional vector space. Calculate the semantic similarity between the new and old knowledge embedding vectors, such as cosine similarity. At the same time, consider the structural information of the knowledge graph, such as the degree of nodes, path length, etc., for structure matching. In the TransE model, for the triple (h, r, t), its loss function is:
[0040] where S is the set of true triples, is the set of wrong triples generated by replacing the head entity h or the tail entity t in the true triple, is the margin parameter. By minimizing the loss function, the embedding vectors of entities and relationships are obtained for calculating semantic similarity and structure matching.
[0041] Step 3, based on the knowledge graph retrieval enhanced generation technology, automatically extracts entities, relationships, and their attributes from structured and unstructured knowledge bases, and creates a large model of the water conservancy project knowledge graph to support the inference application of the multi-objective collaborative scheme of the water conservancy project.
[0042] The said Step 3 is the deep integration of the knowledge graph and the large language model: Based on the GraphRAG-GCN fusion model in the graph neural network (GNN), use multi-layer graph convolution to aggregate neighborhood information. Specifically, represent the knowledge graph as a graph structure and input it into the GraphRAG-GCN fusion model. The model first preprocesses the graph structure through the GraphRAG technology to enhance the representation of key nodes and relationships. Then use GCN for multi-layer graph convolution operations to gradually aggregate neighborhood information. Each layer of GCN updates the node features, so that the node features contain more structural information. In the GCN layer, the feature update formula of node i at the l+1 layer is:
[0043] where, is the feature vector of node i at the l layer, is the set of neighborhood nodes of node i, and are the degrees of node i and neighborhood node l respectively, is the weight matrix at the l layer, is the bias term, is the non-linear activation function. Through multiple iterative updates, high-quality feature representations are provided for the subsequent fusion of the knowledge graph and the large language model.
[0044] When the GNN performs embedding representation learning on the knowledge graph (KG), an attention-guided embedding method is adopted. During the pre-training stage of the LLM, the contrastive learning method is used to filter relevant knowledge graph data. In the GNN embedding representation learning, the attention mechanism is introduced to assign higher weights to the key entity relationships in water conservancy (such as the relationship between the dam and the water flow), guiding the GNN to pay more attention to the feature learning of these relationships. During the pre-training stage of the LLM, the knowledge graph data is divided into positive and negative examples. Through the contrastive learning method, the LLM learns the differences between the positive examples (relevant knowledge graph data) and the negative examples (irrelevant data), so as to filter out the knowledge graph data related to the pre-training task. In the attention mechanism, the attention weight of the relationship r is calculated by the formula:
[0045] In the formula, is the similarity score between the query vector q and the relationship r. In the contrastive learning, the loss function can be defined as:
[0046] In the formula, is the similarity between the query vector q and the key vector k, is the positive example key vector, is the negative example key vector, is the temperature parameter.
[0047] GraphRAG optimizes the subgraph retrieval of the knowledge graph based on SPARQL and adopts a Transformer-based inference model. Specifically, the user's question is semantically parsed to convert the question into a query vector. Based on SPARQL and the graph topology structure information, GraphRAG optimizes the subgraph retrieval to find the subgraph related to the query. The subgraph information is encoded into a feature vector and input into the Transformer-based inference model together with the query vector. The inference model uses the self-attention mechanism of Transformer to process the input information and generate the inference result. In the Transformer-based inference model, the calculation process of the self-attention mechanism is:
[0048] In the formula, Q (Query), K (Key), and V (Value) are the input feature matrices, is the dimension of K.
[0049] Based on the large model of the water conservancy project knowledge graph constructed above, the embodiment of the present invention also provides a multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph. Refer to Figure 2, First, obtain the scheduling requirements and convert them into user questions; then, input the user questions into the constructed large model of the water conservancy project knowledge graph, and generate a multi-objective collaborative scheduling plan for water conservancy projects with the help of the multi-objective optimization tool of Langchain.
[0050] The inference process of the multi-objective collaborative scheduling plan specifically includes: Through the interaction between Langchain and the knowledge graph, introduce the constraint conditions in the knowledge graph into the inference process. In flood control scheduling, Langchain obtains constraint information such as the flood control limited water level of the reservoir and the safe discharge of the downstream river channel from the knowledge graph, and automatically verifies whether the generated plan meets these constraints during inference. In water supply scheduling, similarly, based on information such as the water demand of residents and agriculture in the knowledge graph, ensure that the inference results meet the actual needs and safety standards, and guarantee the rationality and reliability of the decision-making.
[0051] Develop application services for water conservancy business topics. For example, provide solutions by combining data and optimization algorithms in water project scheduling, and use anomaly detection algorithms for early warning in dam safety monitoring. Quickly build applications with the help of Langchain. Recommend knowledge services using collaborative filtering algorithms according to user behavior, and optimize the recommendation by associating with the knowledge graph using graphRAG.
[0052] Taking the inference engine of the large model of the knowledge graph as the core, with the help of the multi-objective optimization tool of Langchain, and combining the multi-objective requirements such as flood control, water supply, and power generation, generate the operation mode of water project scheduling. Comprehensively consider the objective weights such as power generation benefits and downstream water supply guarantee levels, and use built-in multi-objective optimization algorithms such as the weighted summation method and the ε-constraint method to generate the optimal scheduling plan that takes into account the needs of all parties through the inference engine.
[0053] For example, using the weighted summation method to linearly combine multiple objective functions into a comprehensive objective function , the formula is:
[0054] In the formula, is the weight of the i-th objective function, and satisfies , . For example, let the objective function of power generation benefits be , and the objective function of water supply guarantee level be . By reasonably allocating the weights and , make the comprehensive objective function achieve the optimal under certain constraint conditions, so as to generate a scheduling plan that takes into account both power generation and water supply needs.
[0055] Preferably, the construction process of the large model of the water conservancy project knowledge graph provided by the embodiments of the present invention can be implemented through steps 2.1 to 2.2.
[0056] In step 2.1, first, based on the python package, install the retrieval enhanced generation software package graphrag based on the knowledge graph.
[0057] pip install graphrag; Then install the existing water conservancy knowledge base Waterknowledgebase and initialize the project Pip install Waterknowledgebase Waterknowledgebase install; Next, create a temporary folder graphrag to store the running data.
[0058] mkdir. / graphrag / input; curl https: / / www.xxx.com / xxx.txt >. / myTest / input / watersupply.txt / / Here is an example code, which is set according to the business to be tested, such as the txt text of water supply scheduling.
[0059] cd. / graphrag; python -m graphrag.index –init; So far, the acquisition of the source data of the water conservancy knowledge base has been completed, including the import of data such as water conservancy professional data, scientific treatises, journal literature, standards and specifications, laws and regulations, images, and audio voiceprints.
[0060] In step 2.2, first modify the configuration file, modify the settings.yam file, and modify the large language model LLM and the corresponding API_base to be used therein. For example: Encoding_model: cl100k_base; Skip_workflows: [ ]; LLM: API_key: ${GRAPHRAG_API_KEY}; Type: openai_chat # or azure_openai_chat; Model: deepseek-chat # Customize your own large model; Model_supports_json: false # recommend if this is available for your model; API_base: http: / / api.agicto.cn / v1 # Domestic open-source large model; The modified configuration file is a global configuration, and there is no need to modify it again when calling the model later.
[0061] Next, build an index for the retrieval-enhanced generation software package based on the knowledge graph: Waterknowledgebase run wkb index –root.. # Waterknowledgebase is a water conservancy knowledge base; Finally, select global query or local query to implement water conservancy project knowledge retrieval.
[0062] Such as global query: It focuses more on retrieving information about the joint water supply scheduling of reservoir groups in a basin, Waterknowledgebase run wkb query –root.. --method global “Joint water supply scheduling information of reservoir groups” Such as local query: It focuses more on retrieving information about the water supply scheduling of a single reservoir Waterknowledgebase run wkb query –root.. --method local “Water supply scheduling information of each reservoir”.
[0063] Please refer to Figure 3 , Figure 3 The logic schematic diagram of the method for constructing a large model of a water conservancy project knowledge graph provided by an embodiment is shown. It is used to support the realization of water conservancy business intelligence through this derivation logic, and support the reasoning applications of multi-objective collaborative schemes such as flood control, water supply, and power generation in water conservancy projects.
[0064] Please refer to Figure 4, which shows a block diagram of the device for constructing a large model of a water conservancy project knowledge graph provided by this application. The device includes: a first main module 102 for constructing an initial water conservancy knowledge graph based on the obtained knowledge base source data; a second main module 104 for introducing a generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph to obtain an optimized water conservancy knowledge graph; a third main module 106 for performing knowledge graph embedded learning using a graph neural network based on the optimized water conservancy knowledge graph, and introducing an attention mechanism and a contrast learning method to screen out knowledge graph data related to the pre-training task; a fourth main module 108 for performing reasoning based on the screened knowledge graph data in combination with a reasoning model based on Transformer to obtain a large model of a water conservancy project knowledge graph with deep integration of the knowledge graph and a large language model.
[0065] Please refer to Figure 5 , which shows a block diagram of the device for constructing a large model of a water conservancy project knowledge graph provided by this application. The model includes: A data acquisition module 202 for acquiring scheduling requirements, including water conservancy professional data, scientific treatises, journal literature, standards and specifications, laws and regulations, images, audio fingerprints, etc. A collaborative scheduling module 204 for inputting the scheduling requirements into the constructed large model of a water conservancy project knowledge graph and generating a multi-objective collaborative scheduling scheme for water conservancy projects with the help of a multi-objective optimization tool of Langchain.
[0066] Current water conservancy large models are all developed in response to the construction goal of "digital twin basin", mainly facing the needs of the basin level and the wide business field of water conservancy and hydropower, and cannot support the application of typical business (L3 level) fields such as dam safety monitoring and unit fault prediction in water conservancy projects. The knowledge engine of current water conservancy large models is driven by a large language model (LLM), and there is a problem of weak knowledge reasoning ability, which cannot support the reasoning of multi-objective collaborative schemes and the current situation of intelligent decision-making in water conservancy projects. Using several modules in the figure can not only capture the high-order relationships of multi-modal information such as Euclidean data and non-Euclidean data in the water conservancy knowledge base, but also support the reasoning of multi-objective collaborative schemes such as flood control scheduling, power generation scheduling, water supply scheduling, dam safety monitoring, and unit fault prediction in water conservancy projects and intelligent decision-making, promoting the intelligent transformation of the water conservancy and hydropower industry and developing new water conservancy productive forces.
[0067] It should be noted that the device embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The difference lies only in setting corresponding functional modules. The principle is basically the same as that of the above device embodiments provided by the present invention. As long as those skilled in the art, on the basis of the above device embodiments, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions constituted by these technical means, and ensure the practicability of the technical solutions, the devices in the above device embodiments can be improved to obtain corresponding device-type embodiments for implementing the methods in other method-type embodiments.
[0068] In view of the fact that the construction of the large model of the water conservancy project knowledge graph and the multi-objective collaborative scheduling method have been discussed in detail in the previous method, the specific execution of the above modules will not be elaborated here.
[0069] Next, a kind of execution device provided by the embodiments of the present invention will be introduced. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the execution device provided by the embodiments of the present invention. The execution device 300 can specifically be an autonomous driving vehicle, a mobile phone, a tablet computer, a laptop computer, a desktop computer, a monitoring data processing device, etc., which is not limited here. Among them, the functions of the execution device corresponding to Figure 1 in the corresponding embodiments are implemented on the execution device 300. Specifically, the execution device 300 includes: a receiver 301, a transmitter 302, a processor 303, and a memory 304 (where the number of processors 303 in the execution device 300 can be one or more, Figure 6 taking one processor as an example here). Among them, the processor 303 can include an application processor 3031 and a communication processor 3032. In some embodiments of the present application, the receiver 301, the transmitter 302, the processor 303, and the memory 304 can be connected through a bus or other means.
[0070] The memory 304 can include a read-only memory and a random access memory, and provide instructions and data to the processor 303. A part of the memory 304 can also include a non-volatile random access memory (NVRAM). The memory 304 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions can include various operation instructions for implementing various operations.
[0071] The processor 303 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, which may include, in addition to the data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.
[0072] The method disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 303. The processor 303 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 303 or the instructions in the form of software. The above-mentioned processor 303 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 303 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 304, and the processor 303 reads the information in the memory 304 and combines its hardware to complete the steps of the above method.
[0073] The receiver 301 can be used to receive input digital or character information and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 302 can be used to output digital or character information through the first interface; the transmitter 302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 302 can also include a display device such as a display screen.
[0074] In the embodiments of the present invention, the processor 303 is used to execute Figure 1 the method for obtaining the system capacity executed by the execution device in the corresponding embodiment. The specific manner in which the application processor 3031 in the processor 303 executes the above steps is the same as that in the present application Figure 1The corresponding method embodiments are based on the same concept, and the technical effects brought by them are the same as those in this application Figure 1 The corresponding method embodiments are the same. For specific content, please refer to the descriptions in the method embodiments shown above in this application, and details will not be repeated here.
[0075] Based on the same technical concept as the foregoing embodiments, an embodiment of the present invention also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method for constructing a large model of a water conservancy project knowledge graph or the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph.
[0076] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0077] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0078] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide for realizing the functions in the processFigure 1 one or more processes and / or boxes Figure 1 steps of functions specified in one or more boxes
[0080] In summary, based on the water conservancy project management business knowledge base, the present invention constructs a knowledge graph for typical business fields of water conservancy projects by adopting a construction method for the data layer, concept layer, and instance layer of a water conservancy knowledge graph that integrates multiple business knowledges, and extracts the logical chain among the water conservancy project scheduling, dam safety monitoring, and unit fault prediction services; aiming at the characteristics of complex water conservancy business knowledge structure and rich potential associations, through the self-iterative water conservancy business knowledge graph completion technology based on a large language model, the completeness and accuracy of the knowledge graph for typical business fields of water conservancy are improved; in order to effectively combine the advantages of the knowledge graph and the large model in knowledge representation and knowledge parsing, the retrieval-enhanced generation technology based on the knowledge graph is adopted to automatically extract entities, relationships, and their attributes from structured and unstructured knowledge bases, create a large model of the water conservancy project knowledge graph, and support the reasoning application of multi-objective collaborative schemes such as flood control, water supply, and power generation in water conservancy projects.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. Construction method of large model for knowledge graph of water conservancy projects, characterized in that Including: Based on the knowledge base of water conservancy project management business, adopting the construction methods of data layer, concept layer and instance layer of water conservancy knowledge graph with multi-business knowledge integration, constructing the knowledge graph of typical business fields of water conservancy projects, and refining the logical chain among the business of water conservancy project scheduling, dam safety monitoring and unit fault prediction; Using the self-iterative knowledge graph completion technology of typical business fields of water conservancy projects based on large language models to improve the knowledge graph of typical business fields of water conservancy projects; Based on the retrieval-enhanced generation technology of knowledge graph, automatically extracting entities, relationships and their attributes from structured and unstructured knowledge bases, and creating a large model of water conservancy project knowledge graph to support the reasoning application of multi-objective collaborative schemes for water conservancy projects.
2. The method for constructing a large model of a water conservancy project knowledge graph according to claim 1, wherein Constructing the knowledge graph of typical business fields of water conservancy projects, including: Adopting the BERT-BiLSTM-CRF model based on the Transformer architecture to construct an initial water conservancy knowledge graph based on the obtained knowledge base source data. Among them, the pre-trained language model BERT is used to pre-encode the text to capture the deep semantic features of the text; the bidirectional long short-term memory network BiLSTM is used to process the output of the pre-trained language model BERT and better capture the context information by using its bidirectional characteristics; the conditional random field model CRF is used to determine the optimal label sequence in the sequence labeling task.
3. The method for constructing a large model of a water conservancy project knowledge graph according to claim 1, wherein Using the self-iterative knowledge graph completion technology of typical business fields of water conservancy projects based on large language models to improve the knowledge graph of typical business fields of water conservancy projects, including: Using the generative language model to conduct semantic mining on the text description of the knowledge graph, and generating the semantic analysis result as the newly mined knowledge in an autoregressive manner; Using the knowledge graph embedding model to perform embedding representations on the newly mined knowledge and the existing knowledge in the knowledge graph respectively, and fusing them through the semantic similarity and structural matching between the new and old knowledge; Through continuous iteration, extracting knowledge from new data and updating the knowledge graph.
4. The method for constructing a large model of a water conservancy project knowledge graph according to claim 1, wherein Based on the retrieval-enhanced generation technology of knowledge graph, automatically extracting entities, relationships and their attributes from structured and unstructured knowledge bases, and creating a large model of water conservancy project knowledge graph, including: Using the graph neural network to conduct embedding representation learning on the knowledge graph, and retrieving subgraphs or knowledge fragments related to the input question from the knowledge graph; Inputting the retrieved subgraph or knowledge fragment together with the input question into the reasoning model based on Transformer to generate text output.
5. The method for constructing a large model of a water conservancy project knowledge graph according to claim 4, wherein, The method further includes: Representing the knowledge graph as a graph structure and inputting it into the GraphRAG-GCN fusion model. After preprocessing the graph structure through the GraphRAG technology, using GCN to perform multi-layer graph convolution operations to gradually aggregate neighborhood information to form high-quality feature representations; Based on the formed high-quality feature representations, using the graph neural network to conduct embedding representation learning on the knowledge graph and introducing the attention mechanism during the learning process; Based on the features after embedding representation learning, introducing the contrastive learning method in the pre-training stage of the reasoning model based on Transformer to screen out the knowledge graph data related to the pre-training task.
6. The multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph according to any one of claims 1-5, characterized in that Including: Obtaining scheduling requirements; Input the scheduling requirements into the constructed large model of the water conservancy project knowledge graph, and generate a multi-objective collaborative scheduling plan for the water conservancy project with the help of the multi-objective optimization tool of Langchain.
7. Device for constructing large model of knowledge graph of water conservancy project, characterized in that, It includes: The first main module is used to construct an initial water conservancy knowledge graph based on the obtained knowledge base source data; The second main module is used to introduce a generative language model and knowledge graph embedding technology to optimize the initial water conservancy knowledge graph to obtain an optimized water conservancy knowledge graph; The third main module is used to perform knowledge graph embedded learning based on the optimized water conservancy knowledge graph using a graph neural network, and introduce an attention mechanism and a contrast learning method to screen out knowledge graph data related to the pre-training task; The fourth main module is used to perform reasoning based on the screened knowledge graph data in combination with a reasoning model based on Transformer to obtain a large model of the water conservancy project knowledge graph with deep integration of the knowledge graph and the large language model.
8. The multi-objective collaborative scheduling device based on the large model of the water conservancy project knowledge graph according to claim 7, characterized in that, It includes: The data acquisition module is used to acquire scheduling requirements; The collaborative scheduling module is used to input the scheduling requirements into the constructed large model of the water conservancy project knowledge graph, and generate a multi-objective collaborative scheduling plan for the water conservancy project with the help of the multi-objective optimization tool of Langchain.
9. An execution device, characterized in that, It includes a processor and a memory, and the processor is coupled with the memory; the memory is used to store programs; the processor is used to execute the programs in the memory, so that the execution device executes the method for constructing the large model of the water conservancy project knowledge graph according to any one of claims 1 to 5 or the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph according to claim 6.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method for constructing the large model of the water conservancy project knowledge graph according to any one of claims 1 to 5 or the multi-objective collaborative scheduling method based on the large model of the water conservancy project knowledge graph according to claim 6.
Citation Information
Patent Citations
Water knowledge concept map model construction method
CN111522960A
Method for constructing domain knowledge graph of hydrological model
CN116383395A
Water conservancy knowledge service system based on knowledge graph technology
CN118152525A
Interactive argument pair identification method based on contrast enhancement and multi-scale semantic perception
CN118467675A
Method for constructing chemical-plastic industry chain knowledge graph by using graph convolutional network
CN119250172A
Cited By
Reservoir scheduling method and device based on intelligent agent, electronic equipment and storage medium
CN121073063A
Flood control emergency plan generation method, system and equipment and storage medium
CN121093173A
Joint modeling method based on knowledge graph and large model retrieval enhancement
CN121117179A