Gravity dam safety risk management and control method and system based on knowledge graph and large model
By combining knowledge graphs and large models to build an intelligent decision-making system, the data fragmentation and decision-making lag problems in the safety risk prevention and control of concrete gravity dams are solved, efficient and accurate risk control is achieved, and intelligent management of dam projects is promoted.
Patent Information
- Application Number
- CN202510398518.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-25
AI Technical Summary
The existing technology has multi-source data fragmentation, static reasoning limitations and decision-making lag in the safety risk prevention and control of concrete gravity dams, which affects the timeliness and accuracy of risk prevention and control.
Combining knowledge graphs and large models, by collecting water conservancy engineering accident case data, building entity and relationship extraction models, generating knowledge graphs, and using large language models to conduct case reasoning and decision-making, combining expert opinions to make secondary fine-tuning to form an intelligent decision-making system.
It improves the decision-making efficiency and accuracy of the prevention and control of gravity dam safety risks, reduces the time cost of traditional methods, and promotes the intelligent and informatization development of dam construction safety management.
Smart Images

Figure CN120373844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of water conservancy projects, and specifically relates to a method and system for gravity dam safety risk control based on knowledge graphs and large models. Background Art
[0002] In the process of safety risk prevention and control management of concrete gravity dams, there are technical bottlenecks in aspects such as multi-source data fragmentation, limitations of static reasoning, decision-making lag, and knowledge inheritance, which seriously affect the safety prevention and control of gravity dams. Although existing research on gravity dam safety has conducted safety assessments of dams by constructing risk assessment index systems, risk assessment models, and using static cross-sectional data, and used expert experience for safety risk prevention and control, these traditional assessment and decision-making methods will inevitably affect the timeliness and accuracy of gravity dam risk prevention and control and decision-making, and even lead to the loss of the best decision-making window for gravity dam risk prevention and control and decision-making. Research on risk case reasoning and decision-making methods for concrete gravity dams using multi-modal knowledge fusion and intelligent reasoning technologies is of great significance for ensuring the safety of dam construction and service. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to improve traditional case reasoning and decision-making technologies by combining knowledge graphs and large model artificial intelligence technologies, and use the powerful language and reasoning capabilities of large models to form a system that integrates intelligent reasoning and decision-making. To solve the above problems, a method and system for gravity dam safety risk control based on knowledge graphs and large models are provided.
[0004] The purpose of the present invention is achieved in the following manner:
[0005] A method for gravity dam safety risk control based on knowledge graphs and large models, the method comprising:
[0006] S1. Collect water conservancy project accident case data, import it into doccano to form an original data set after basic data cleaning, and export the data after entity and relationship annotation using doccano and convert the data format to form a final data set;
[0007] S2. Divide the data set into a training set, a test set, and a validation set, import it into a fine-tuning script, and perform fine-tuning on the relationship and entity extraction model to obtain an entity and relationship extraction model with domain knowledge;
[0008] S3. Use the fine-tuned model with domain knowledge to extract entities and relationships from all case data in the case base to obtain a case data knowledge graph;
[0009] S4. Convert the obtained case data knowledge graph into natural language descriptions, and then use these texts to fine-tune the large language model to finally obtain a large language model with domain knowledge;
[0010] S5. Extract the project overview information using the engineering model, and use the extraction results to form a natural language description of the problem through PromptFormulation and input it into the large language model with domain knowledge in S4;
[0011] S6. The large language model answers the question, and the expert judgment module provides expert opinions based on the reply of the large model. In this way, the expert module can perform secondary fine-tuning on the large model, and at the same time, the correct reply content of the large model can be stored in the case library again as new case data.
[0012] Specifically, in S1, it includes: collecting water conservancy project accident case data, the sources of which cover news reports, enterprise internal documents and relevant books. After basic data cleaning work, including removing duplicate content, unifying terms and standardizing units, import the cleaned text data into doccano, create a NER task, define the entities and relationships in water conservancy project accidents, and let the annotator annotate the entities and relationships of the text to ensure the consistency and accuracy of the annotation. After the annotation is completed, use a python script to convert the data exported from doccano into a json format suitable for model training.
[0013] Specifically, in S2, it includes: dividing the data set into a training set, a test set, and a validation set to ensure the balanced distribution of various types of data, and using Baidu's UIE extraction framework as the basic tool for entity and relationship extraction; import the divided training set into the fine-tuning script, and use the Baidu UIE framework to fine-tune the relationship and entity extraction model to optimize the performance of the model in entity and relationship extraction tasks; finally, use the validation set to evaluate the performance of the fine-tuned model, adjust the hyperparameters to improve the accuracy and generalization ability of the model, and save the fine-tuned entity and relationship extraction model with domain knowledge.
[0014] Specifically, S3 includes: using the fine-tuned model to process all accident case texts in the case library, identifying the entities and categories in the cases, and extracting the relationships between the entities. According to the extracted entities and relationships, construct a structured knowledge graph, which includes entities of multiple categories such as geographical entities, engineering structures, geological entities, time entities, numerical and measurement entities, personnel and organization entities, and event entities and their associated relationships.
[0015] Specifically, S4 includes: converting the knowledge graph into natural language descriptions and fine-tuning the large language model. Using a templating method, entities and relationships in the knowledge graph extracted in S3 are converted into natural language sentences. The converted natural language descriptions are used as training data, and a fine-tuning tool is used to fine-tune the large language model to enhance its understanding and generation capabilities in the field of water conservancy project accidents; evaluating the performance of the fine-tuned large language model in tasks such as water conservancy project accident analysis, prevention suggestions, and emergency response, and deploying it to the actual application environment.
[0016] Therefore, the engineering models in S5 include a factor analysis and risk level determination model and a case similarity calculation model;
[0017] The factor analysis and risk level determination model identifies the main influencing factors of the accident and gives its risk level; the case similarity calculation model finds similar accident cases;
[0018] The factor analysis and risk level determination model includes:
[0019] (1) Combining the construction scenario of gravity dams in cold regions, according to expert opinions and literature reviews, risk evaluation indicators are selected from aspects such as construction personnel safety risks, material and equipment safety risks, construction management safety risks, construction environment safety risks, and construction technology safety risks;
[0020] (2) Select k experts, E = {E1, E2,... E k}], and according to their different experiences, different weights are assigned to the scores of the k experts. The actual performance of the risk factors and their relative importance are evaluated using linguistic variables, and their evaluation results are converted into a triangular fuzzy matrix;
[0021] (3) Comprehensive weighting of risk factors is carried out using FAHP and the maximum deviation method; using the comprehensive weighting method, first calculate the subjective weight α of each index using the AHP method i ; then calculate the objective weight of each index as β using the maximum deviation method i ; finally, the comprehensive weight is obtained as: w i = aα i + bβ i . Where: a is the subjective weight influence factor, b is the objective weight influence factor; based on the VIKOR method, the risk factors and their impacts are sorted, and the sorting results are evaluated and analyzed to obtain the final risk level.
[0022] A gravity dam safety risk control system based on a knowledge graph and a large model, the system includes:
[0023] Domain Model Fine-tuning Module: Used to achieve automated information extraction, obtain structured knowledge, and generate a knowledge graph; including fine-tuning the information extraction model using engineering domain knowledge and engineering cases to enhance the model's information extraction ability in the engineering domain;
[0024] Knowledge Graph Construction Module: Used to structurally represent the data information of dam safety risk accident cases; including visualizing the entities and relationships of dam safety risk accident case data, displaying the accident case entities and relationships extracted by the fine-tuned domain model, and storing and managing this information for subsequent generation of natural language descriptions and reasoning;
[0025] Engineering Model Calculation Module: Used to calculate the highest-risk factors, safety risk assessment levels, and the case with the highest similarity to the case library based on the engineering overview information provided by the user; including a factor analysis and risk level determination model and a case similarity calculation model, and finally forming a question based on the calculation results and a pre-designed Prompt template as the input to the intelligent decision-making module;
[0026] Intelligent Decision-making Module: Used to generate a final decision based on the structured knowledge in the domain model knowledge graph and the questions formed by the Prompt; including large model fine-tuning, Prompt template design, and expert opinion re-fine-tuning. During the system optimization phase, it supports developers to optimize the large model output based on expert opinions in this module.
[0027] The domain model fine-tuning module includes
[0028] S1. Obtain safety accident cases from the network, books, and safety accident reports, and perform basic cleaning operations on them: remove special characters, extra line breaks, and spaces that affect annotation; organize each case into an independent text description;
[0029] S2. Upload the preprocessed text data to doccano, use doccano to complete the sequence annotation task, and create entity tags and relationship tags before completing the annotation task; each case is an independent annotation task; the necessary entities for annotation are: water conservancy project, safety accident, management measures, lessons learned, and the necessary relationships for annotation are: what safety accident does the project "have", what disasters does the safety accident "cause", and what lessons are "derived" from this; after the annotation is completed, export the annotation data;
[0030] S3. Process the annotation data exported in S2, convert the original data exported by doccano into samples that conform to the classification task format, divide the training set, test set, and validation set; then use the divided data set to retrain and fine-tune the general entity relationship extraction model, and finally obtain a model with engineering domain knowledge that can perform engineering domain relationship entity extraction;
[0031] S4. Use the model fine-tuned in S3 to perform entity extraction and relationship extraction on a large number of safety accident cases for knowledge fusion and ablation, construct a water conservancy safety accident knowledge graph, store it in the Node4j graph database and visualize it;
[0032] S5. Convert the structured data in the knowledge graph constructed in S4 into natural language descriptions to construct a dataset for fine-tuning the large model.
[0033] The engineering model calculation module and the intelligent decision-making module include:
[0034] S1. Deploy a local API to implement case reasoning and emergency decision-making. Users can obtain similar case analyses and emergency decision-making processing methods by providing project overview information and problems;
[0035] S2. Import the "project overview information" input by the user in S1 into the "factor analysis and risk level determination model" and the "similarity calculation model" respectively. The two models will respectively obtain the highest-level risk factors in the project overview information and the case name with the highest similarity in the case library;
[0036] S3. Based on the outputs of the two models obtained in S2, design a Prompt template to convert the two "isolated" word descriptions into natural language descriptions, which include information such as risk factors, risk factor levels, similar cases and their corresponding disposal measures, etc.;
[0037] S4. Based on the natural language description converted in S3 and the Prompt engineering, set the role identity for the large language model. Finally, consult the large language model for similar accidents or handling measures in the case of cases in the form of questions. The large language model gives a reply based on the structured knowledge constructed based on the case library through its powerful language ability and reasoning ability;
[0038] The questions answered by the case reasoning and emergency decision-making system based on the LLM in S4 are in the form of natural language descriptions. Experts need to judge the reasoning / output effect according to the collective content of its answers. If it meets the actual situation and is recognized by experts, the reply result will be saved as memory and a new case in the case library. If hallucinations occur or the results are not ideal, experts will provide suggestions and then ask the system again. Repeat this process multiple times to achieve secondary Fine-tune. Make the Base-LLM become a Domain-LLM with professional domain knowledge.
[0039] The beneficial effects of the present invention: 1) The present invention improves the traditional case reasoning and decision-making technology, and uses the powerful language and reasoning abilities of the large model to form an intelligent system that integrates artificial intelligence, reasoning and decision-making.
[0040] 2) By combining the Word2vec model in deep learning with case-based reasoning technology, the present invention proposes an efficient analogical framework for hidden danger governance solutions. This framework can quickly retrieve historical cases similar to the target hidden danger from large-scale hidden danger text data and provide reference solutions for managers, promoting the intelligent and information-based development of dam project construction safety management.
[0041] 3) The present invention constructs a multi-factor risk assessment model applicable to gravity dams and verifies the effectiveness of triangular fuzzy theory, fuzzy AHP, maximum deviation method, and VIKOR in comprehensive decision-making. This model provides a systematic theoretical framework and practical tool for the safety assessment of cold-region concrete gravity dam construction, expanding the application of multi-criteria decision-making methods in cold-region project safety management.
[0042] 4) Traditional dam decision-making and assessment methods inevitably consume a large amount of time, affecting the timeliness and accuracy of gravity dam risk prevention and control and decision-making. In contrast, the present invention can improve decision-making efficiency, enhance the accuracy of risk prevention and control, and reduce the time cost of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic flow chart of the present invention.
[0044] Figure 2 is a schematic diagram of the definition of the ontology when extracting entities and relationships in the present invention.
[0045] Figure 3 is a schematic diagram of the knowledge graph extracted from cases by the fine-tuned domain model of the present invention.
[0046] Figure 4 is a flow chart of the calculation factor analysis and risk level determination model of the present invention.
[0047] Figure 5 is a flow chart of the calculation case similarity model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0049] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same technical meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0050] The present invention provides a dam safety risk case-based reasoning and decision-making method based on a knowledge graph and a large model, including the following steps:
[0051] S1. Collect data on water conservancy project accident cases, with sources covering news reports, enterprise internal documents, and relevant books. After basic data cleaning work, including removing duplicate content, unifying terms (such as "reservoir water level" unified as "reservoir level") and standardizing units, import the cleaned text data into doccano, create an NER task, define entities and relationships in water conservancy project accidents. Entity definitions include "water conservancy project", "accident", "cause", etc., and relationship types include "cause", "lead to", "derive", etc. Two annotators with professional knowledge annotate the entities and relationships in the text to ensure annotation consistency and accuracy. After annotation, use a Python script to convert the data exported from doccano into a JSON format suitable for model training.
[0052] S2. Divide the dataset into a training set, a test set, and a validation set (7:1.5:1.5) to ensure balanced distribution of various types of data. Adopt Baidu's UIE (Universal Information Extraction) extraction framework as the basic tool for entity and relationship extraction. Import the divided training set into the fine-tuning script, and use the Baidu UIE framework to fine-tune the relationship and entity extraction model to optimize the model's performance in entity and relationship extraction tasks. Finally, use the validation set to evaluate the performance of the fine-tuned model, adjust the hyperparameters to improve the model's accuracy and generalization ability, and save the fine-tuned entity and relationship extraction model with domain knowledge for subsequent use.
[0053] S3. Use the fine-tuned model to process all accident case texts in the case library, identify entities and categories in the cases, and extract the relationships between entities. Based on the extracted entities and relationships, construct a structured knowledge graph, including entities of multiple categories such as geographical entities, engineering structures, geological entities, time entities, numerical and measurement entities, personnel and organization entities, and event entities, as well as their associated relationships.
[0054] S4. Convert the knowledge graph into natural language descriptions and fine-tune large language models. Adopt a templated method to convert the entities and relationships in the knowledge graph extracted in S3 into natural language sentences. For example, convert (Vajont Reservoir, located in, the village of Pieve di Cadore, Veneto, Italy) into "Vajont Reservoir is located in the village of Pieve di Cadore, Veneto, Italy". Use the converted natural language descriptions as training data, and use fine-tuning tools such as Hugging Face Transformers to fine-tune large language models such as GPT-4 to enhance their understanding and generation capabilities in the field of water conservancy project accidents. Evaluate the performance of the fine-tuned large language model in tasks such as water conservancy project accident analysis, prevention suggestions, and emergency response, and deploy it to the actual application environment.
[0055] S5. As Figure 3As shown in the figure, the main steps of the factor analysis and risk level determination model are as follows:
[0056] (1) Combining the construction scenarios of gravity dams in cold regions, according to expert opinions and literature reviews, this paper selects risk evaluation indicators from aspects such as construction personnel safety risks (S1), material and equipment safety risks (S2), construction management safety risks (S3), construction environment safety risks (S4), and construction technology safety risks (S5).
[0057] (2) Select k experts, E = {E1, E2, … E k}, according to their different experiences, assign different weights to the scores of the k experts, and use linguistic variables to evaluate the actual performance of risk factors and their relative importance. Convert their evaluation results into a triangular fuzzy matrix.
[0058] (3) Comprehensive weighting of risk factors is carried out using FAHP and the maximum deviation method. In the traditional LEC method, it is defaulted that the risk factors L, E, and C have the same importance, and the weights between risk factors are not considered. Directly determining the weights of risk factors is not scientific. Therefore, this paper considers both subjective and objective factors and uses the comprehensive weighting method. First, use the AHP method to calculate the subjective weight α i of each index; then use the maximum deviation method to calculate the objective weight of each index as β i ; finally, the comprehensive weight is obtained as: w i = aα i + bβ i . Where: a is the subjective weight influence factor, and b is the objective weight influence factor.
[0059] 1) Calculation of subjective weight based on fuzzy AHP. FAHP combines AHP and fuzzy theory, converts the linguistic evaluations of experts into fuzzy numbers, constructs a fuzzy matrix for processing, and then obtains the subjective weights of risk factors. This paper converts the evaluation language into triangular fuzzy numbers. Experts give linguistic evaluations based on their true thoughts during the review and decision-making process. The conversion relationship between triangular fuzzy numbers and the linguistic terms for team members to evaluate the subjective weight vector of risk factors is shown in Table 1:
[0060] Table 1 Linguistic variables for rating risk factor weights
[0061]
[0062] Let the object set be X = {x1, x i , … x n}, and G = {g1, g i , … g n} be a target set. For each object x i, degree analysis is performed on each target, and the degree analysis values under m targets can be obtained, with the symbols as follows:
[0063]
[0064] Among them, is a triangular fuzzy number. The steps to calculate the subjective weights of risk factors L, E, and C are as follows:
[0065] Step1: Calculate the value of the fuzzy comprehensive degree relative to the i-th target according to the following formula.
[0066]
[0067] Among them,
[0068]
[0069] Step2: Compare the two triangular fuzzy numbers S1=(l1, m1, u1) and S2=(l2, m2, u2)
[0070]
[0071] Among them, l1, m1, u1 and l2, m2, u2 are the central value, left fuzzy value, and right fuzzy value of the fuzzy numbers S1 and S2 respectively. The above formula can also be described as:
[0072]
[0073] Among them, point D is located between μs1 and μs2, and is the coordinate of the highest point of the intersection part. Among them, μ takes the values of l, m, and u corresponding to formula 5.
[0074] Step3: The possibility degree that a convex fuzzy number is greater than k convex fuzzy numbers S i (i = 1, …, k) can be calculated by formula (6) as follows:
[0075]
[0076] Step4: Let d'(A i ) = minV(S i ≥ S j ), i = 1, 2, …, k j = 1, 2, …k, k ≠ j. The subjective weight of the risk factor can be calculated by the following formula:
[0077]
[0078] Step5: The normalized subjective weight vector of each risk factor is expressed as:
[0079]
[0080] 2) Objective weight calculation based on the maximum deviation method. To determine the objective weights of risk factors L, E, and C, in this subsection, an objective weight calculation model of risk factors based on the maximum deviation method is established. Let the construction safety risk factors of Linhai Reservoir gravity dam be A i (i = 1, 2, …, m), the evaluation index of risk factors be C j (j = 1, 2, …, n), and the experts be E k (k = 1, 2, …, s). The expert team conducts a linguistic evaluation on the evaluation indexes of each risk factor, and quantifies the evaluation results using triangular fuzzy numbers, as shown in Table 2 Denote the expert E k 's evaluation value of the evaluation index C i in the risk factor A j as p ij k . Use the arithmetic mean algorithm to calculate the group decision value p ij =(p ijl , p ijh , p iju ), and form a fuzzy evaluation matrix R = [p ij ij . Among them, m×n are the central value, left fuzzy value, and right fuzzy value of the triangular fuzzy number respectively
[0081] Table 2 Semantic evaluation table of risk factors
[0082]
[0083] Step1: Defuzzify the evaluation value p ij :
[0084]
[0085] Step2: Normalization of the matrix. The evaluation matrix R = [p ij ij is composed of the elements r m×n . For the evaluation index C j , the highest evaluation value among the m risk factors is The lowest evaluation value is Normalize the matrix R to obtain the standardized matrix NR = [np ij m×n
[0086]
[0087] Step3: According to the maximum deviation method, establish the importance degree IR of the evaluation index C j j Calculation model:
[0088]
[0089] The solution result of this model is:
[0090]
[0091] where p sj is the evaluation value of the s-th evaluation index.
[0092] Step4: According to the above solution results, obtain the objective weight of evaluation index C j as:
[0093]
[0094] 3) Comprehensive weight calculation. This paper adopts the comprehensive weighting method and introduces the risk factor weight adjustment coefficient (usually taking ), and uses and to represent the importance degrees of the subjective weight and the objective weight, and calculate the comprehensive weight as:
[0095]
[0096] (4) Rank the risk factors and their impacts based on the VIKOR method, and evaluate and analyze the ranking results to obtain the final risk level. This part proposes an improved VIKOR method based on fuzzy theory to solve the multi-criteria decision-making problem with uncertainty and ambiguity. This method provides a more reasonable, highly credible and stable decision-making support tool for finding the compromise solution by calculating the fuzzy distance between the risk factors and the ideal solution.
[0097] Step1: Summarize the expert opinions and construct the comprehensive fuzzy evaluation matrix as follows:
[0098] x ij =(x ij1 ,x ij2 ,x ij3 ) (16)
[0099] where:
[0100]
[0101] where xij is the score of risk factor A i relative to criterion C j , and x ij =(x ij1 ,xij2 , x ij3 ), where \(i = 1, 2, \ldots, m\) and \(j = 1, 2, \ldots, n\).
[0102] Step 2: Determine the fuzzy optimal and fuzzy worst values, where \(j = 1, 2, \ldots, n\)
[0103]
[0104] Step 3: Calculate the normalized fuzzy distance \(d\) ij \((i = 1, 2, \ldots, m\) and \(j = 1, 2, \ldots, n)\) as follows:
[0105]
[0106] Step 4: Calculate the maximum group utility \(S\) i and the minimum individual regret \(R\) i , where \(i = 1, 2, \ldots, m\),
[0107]
[0108] where: is the comprehensive weight of the risk factors. represents decision-making based on maximizing the group utility, represents decision-making based on minimizing the individual regret mechanism.
[0109] Step 5: Calculate \(Q\) i
[0110]
[0111] where: \(v\) is the weight of the maximum group utility strategy, and \(1 - v\) is the weight of the individual regret.
[0112] Step 6: Sort the risk factors in descending order according to the values of \(S\), \(R\), and \(Q\). The smaller the value, the lower the risk.
[0113] Step 7: Determine the compromise risk factor risk ranking. That is, for a risk factor, if it satisfies the following two conditions, then the risk ranking is considered the best measured by \(Q\) (the maximum value).
[0114] Condition 1: where \(n\) is the total number of risk factors;
[0115] Condition 2: If the risk factor risk ranking \(A\) (1) is also the optimal risk ranking according to \(S\) i , \(R\) i , then it is determined that \(A\) (1)Sorting for the stable maximum risk.
[0116] If the above two conditions cannot be satisfied simultaneously, then two compromise solution risk sorting results are obtained, including two cases:
[0117] (1) If condition 1 is satisfied but condition 2 is not, then there are two compromise solution risk sortings: A (1) 、A (2) ;
[0118] (2) If condition 1 is not satisfied while condition 2 is satisfied, then there are M compromise solution risk sortings: A (1) 、A (2) 、…、M, where M is the maximized M value determined according to Q(A (1) ) - Q(A (M) ) < 1 / (m - 1).
[0119] S6. As Figure 4 shown, the main steps of the case similarity calculation model are as follows:
[0120] (1) Data processing. To obtain effective and available data related to the hidden danger treatment of gravity dam projects, combined with the "Supervision and Management Measures for the Hidden Danger Treatment of Hydropower Station Dam Projects" issued by the China National Energy Administration and the construction experience of actual projects, the data is processed as follows: 1) Filter the original data, and use mathematical statistics, data mining or other rules to complement relevant missing information or delete duplicate and incorrect information to make the data in the database correct and complete; 2) Remove the stop words in the text description of hidden danger problems, as well as the information that is repeated with relevant attribute descriptions and the information inferred by the supervisor; 3) Use the Jieba word segmentation technology to perform word segmentation on the text description of hidden danger problems.
[0121] (2) Design of hidden danger case representation. Commonly used case representation methods include Predicate Logic Representation, Production Rule Representation, Frame Representation, Semantic Network Representation, Ontology Representation, and Case-Based Representation. Since the frame representation method represents a case as a frame with a hierarchical structure, and the frame contains the attributes of the concept and their corresponding values. This is a representation method similar to object-oriented programming, which can represent different categories, instances, and their characteristics in knowledge. Therefore, the frame representation method is used to represent the hidden danger records of gravity dam projects, and its attributes are modified and supplemented.
[0122] The framework consists of several slots with slot names and slot values. For some complex frameworks, a general framework can be divided into multiple sub - frameworks. For some complex slots, the slot can also be divided into several facets first, and then these facets are assigned values to obtain facet values.
[0123] Using the framework method to represent the case of the construction safety accident of the gravity dam project, as shown in Table 1 specifically, its general framework is:
[0124] S = <I, D, P> (23)
[0125] In the formula, S represents the case of the construction safety accident of the gravity dam project; I represents the sub - framework of the theme information, that is, the description of the accident case; P represents the sub - framework of the early warning plan. The specific hidden danger attributes and their overviews are shown in Table 3.
[0126] Table 3 General framework of the case of the construction safety accident of the gravity dam project
[0127]
[0128] The sub - framework of the theme information of the gravity dam accident includes 11 slots, namely the case location, project name, category of sudden safety incidents during the operation of the gravity dam, dam height, reservoir capacity, etc., as shown in Table 4 specifically.
[0129] Table 4 Attributes and overviews of the gravity dam accident
[0130]
[0131] (3) Case - based reasoning framework. The safety management of the gravity dam project belongs to a knowledge - intensive task. Traditional safety management methods for gravity dam projects have deficiencies in concepts, risk management, technical support, emergency management, etc. It is necessary to improve and enhance through measures such as introducing modern information technology, strengthening risk management, and enhancing technical support capabilities. Based on case - based reasoning, a framework for the case - based reasoning method for the treatment of hidden dangers in the safety of gravity dam projects is constructed. The specific steps are as follows:
[0132] 1) The user inputs the hidden danger feature information of the target case into the system, and conducts clustering analysis on the cases to form a case set
[0133] 2) In the case retrieval stage, according to the optional hidden danger features, determine the case set to which the target case belongs, and use Word2vec in deep learning to calculate the similarity of the description of the hidden danger problems of the cases in the corresponding case set, and obtain the similarity with all cases in the case base;
[0134] 3) Select the source case with the highest similarity to the target case. Combining the hidden danger rectification measures, costs, personnel, time, lessons learned, etc. of the source case, the user revises the hidden danger treatment plan, and finally obtains the hidden danger treatment plan;
[0135] 4) Finally, judge whether this case is saved to the case library according to the threshold, and update the case library data.
[0136] The text analogical method of the hidden danger treatment plan based on the case-based reasoning technology can make full use of the text information of the hidden danger cases that have been successfully rectified, avoiding the waste of information resources; it is beneficial for safety management personnel to overcome the problems of many and complex types of hidden dangers in hydropower engineering construction, and reduce the requirements for the completeness of users' hydropower construction professional knowledge; when dealing with similar hidden danger problems, a reference plan can be directly obtained without having to analyze, calculate, and make decisions from scratch, which helps to improve the safety management efficiency.
[0137] (4) Case retrieval. Case retrieval is the core link of intelligent inference of gravity dam safety. The speed and accuracy of this process directly affect the timeliness and accuracy of the hidden danger treatment plan. By calculating the Jaccard coefficient to screen the case set, the retrieval range can be effectively reduced, thereby reducing the required retrieval time; at the same time, using the Word2vec model based on deep learning to calculate the similarity of the hidden danger problem description can completely retain the semantic information of the hidden danger problem, enhance the accuracy of case similarity calculation, and achieve efficient and accurate case retrieval.
[0138] (5) Determine the case set based on the Jaccard coefficient calculation. Use the Jaccard coefficient to calculate the similarity between different samples, and draw a conclusion of "whether they are the same" by comparing two sets of data. By calculating the Jaccard coefficient of the characteristics of the dangerous case options, it is judged whether the characteristics of the dangerous information match, so as to judge the case set to which the case belongs. When the Jaccard coefficient is large, the similarity between the samples is large.
[0139]
[0140] Where A = {a1, a2, …, a n} and B = {b1, b2, …, b n}; J(A, B) is the Jaccard similarity of sets A and B, which is defined as the ratio of the size of the intersection of sets A and B to the size of the union, and is used to measure the overlap of sets A and B in common elements; the vertical bar || represents the number of elements in the set.
[0141] The case feature set can be expressed as h = {closed and completed, on-site management, personal protection, others}. The conditions for judging whether a hidden danger belongs to this case set are as follows.
[0142]
[0143] If the case set H1 = {"Closed Completion", "On-site Management", "Personal Protection", "Others"}, then J(h, H) = 1, which means the case h belongs to the case set H1. If the case set H2 = {"Closed Completion", "On-site Management", "Facilities and Equipment", "Inadequate Safety Facilities Management"}, then J(h, H2) = 1 / 3 ≠ 1, which means the case h does not belong to the case set H2.
[0144] (6) Calculate the similarity of potential hazard texts based on Word2vec. Use short texts to represent the feature information of "problem descriptions of potential hazard cases". Texts belong to unstructured data information and are difficult to be simplified into computer language. Therefore, their similarity cannot be directly calculated. Based on the Word2vec model in deep learning, preprocess the data, build a model imitating the processing method of the human brain, and train the data. Integrate statistical methods to convert potential hazard text information into vectors in the semantic space, represent all words as low-dimensional dense vectors, and thus qualitatively measure the similarity between words in the word vector space. There are two training models, namely Continuous Bag of Words (CBOW) and Skip-gram. Among them, CBOW predicts the current value through the context, and Skip-gram predicts the context through the current value. The model consists of an input layer, a projection layer, and an output layer. In this study, the CBOW model is selected. Input the words related to the context of the target word, train the word vector through the neural network, and finally output the word vector of the target word. The specific principle is as follows:
[0145] Given a word sequence C = {w1, w2, …, w m}, calculate the objective function using maximum likelihood estimation as follows.
[0146]
[0147] In the formula, w i is a certain central word, t is the size of the window on the left and right of the central word w i ; P(w i |w i-t , …, w i+t ) represents the probability of the central word w i given the context of the central word w i , and is calculated through the softmax regression function training.
[0148]
[0149] In the formula, W is the dictionary library, v i is the word vector representation of the central word w i ;
[0150] v0 is the mean of the word vectors of the context words of w i , and the formula is
[0151]
[0152] Maximize the objective function using the stochastic gradient ascent method, and calculate the probability P(w i appearing in the context, i |w i-t ,…,w i+t ), to achieve the prediction of a specific word; aiming at maximizing the similarity between the predicted specific word and the measured specific word, feedback and correct the word vector of the input word to obtain the final vector word vector v i (w i ).
[0153] Calculate the cosine distance between the final word vectors, which is the semantic similarity of the specific description of the hidden danger. The formula is
[0154]
[0155] In the formula, the vertical bar || represents the norm of the vector.
[0156] S7. Use the engineering model (factor analysis and risk level determination model and case similarity calculation model) to extract the project overview information. For example, use the factor analysis and risk level determination model to identify the main influencing factors of the accident and give their risk levels, and use the case similarity calculation model to find similar accident cases. Borrow the extraction results and use PromptFormulation (prompt construction) to form a natural language description of the problem and input it into the large language model with domain knowledge in S4.
[0157] S8. After receiving the problem described in natural language, the large language model generates a corresponding reply. The reply result is judged by experts in the field of water conservancy projects, who provide professional opinions and improvement suggestions. According to the expert opinions, the large model is fine-tuned twice to further improve the accuracy and professionalism of the model. Store the correct reply content of the large model as new case data in the case library again to enrich the knowledge graph and training data.
[0158] A gravity dam safety risk control system based on a knowledge graph and a large model, comprising:
[0159] Domain model fine-tuning module: used to achieve automated information extraction, obtain structured knowledge, and generate a knowledge graph. It includes fine-tuning the information extraction model using engineering domain knowledge and engineering cases to improve the model's ability to extract information in the engineering domain.
[0160] Knowledge Graph Construction Module: Used to structurally represent the data information of gravity dam safety risk accident cases. It includes the visualization of entities and relationships in the gravity dam safety risk accident case data, showing the accident case entities and relationships extracted from the fine-tuned domain model. At the same time, it stores and manages this information for subsequent generation of natural language descriptions and reasoning.
[0161] Engineering Model Calculation Module: Used to calculate the highest-risk factor, safety risk assessment level, and the case with the highest similarity to the case library based on the engineering overview information provided by the user. It includes a factor analysis and risk level determination model and a case similarity calculation model. Finally, based on the calculation results and a pre-designed Prompt template, a question is formed as the input to the intelligent decision-making module.
[0162] Intelligent Decision-making Module: Used to generate the final decision based on the structured knowledge in the domain model knowledge graph and the questions formed by the Prompt. It includes large model fine-tuning, Prompt template design, and expert opinion re-fine-tuning. During the system optimization phase, it supports developers to optimize the large model output based on expert opinions in this module.
[0163] The domain model fine-tuning module includes:
[0164] S1. Obtain safety accident cases from the network, books, safety accident reports, etc., and perform basic cleaning operations on them: remove special characters, extra line breaks, spaces, etc. that affect annotation. Organize each case into an independent text description.
[0165] S2. Upload the pre-processed text data to doccano, and use doccano to complete the sequence labeling task. Before completing the labeling task, entity labels and relationship labels need to be created. Each case is an independent labeling task. The necessary entities for labeling are: water conservancy projects, safety accidents, management measures, lessons learned, etc. The necessary relationships for labeling are: what safety accidents does the project "occur", what disasters does the safety accident "cause", and what lessons are "derived" from this, etc. After the labeling is completed, export the labeled data.
[0166] S3. Process the labeled data exported in S2, convert the original data exported from doccano into samples that conform to the classification task format, and divide the training set, test set, and validation set. Then use the divided data set to retrain and fine-tune the general entity relationship extraction model, and finally obtain a model that can perform engineering domain relationship entity extraction with engineering domain knowledge.
[0167] S4. Use S3 to perform entity extraction and relationship extraction on a large number of safety accident cases in the trained and fine-tuned model for knowledge fusion and ablation, construct a water conservancy safety accident knowledge graph, store it in the Node4j graph database, and visualize it.
[0168] S5. Convert the structured data in the knowledge graph constructed in S4 into natural language descriptions to construct a dataset for fine-tuning the large model.
[0169] The engineering model calculation module and the intelligent decision-making module include the following steps:
[0170] S1. Deploy a local API to implement case reasoning and emergency decision-making. Users can obtain similar case analyses and emergency decision-making processing methods by providing project overview information and problems.
[0171] S2. Import the "project overview information" input by the user in S1 into the "factor analysis and risk level determination model" and the "similarity calculation model" respectively. The two models will respectively obtain the highest-level risk factors in the project overview information and the name of the case with the highest similarity in the case library.
[0172] S3. Based on the outputs of the two models obtained in S2, design a Prompt template to convert the two "isolated" word descriptions into natural language descriptions, which include information such as risk factors, risk factor levels, similar cases, and their corresponding disposal measures.
[0173] S4. Based on the natural language description converted in S3 and Prompt engineering, set the role identity for the large language model. Finally, consult the large language model about similar accidents or handling measures in the case library in the form of questions. The large language model gives answers based on the structured knowledge constructed based on the case library through its powerful language ability and reasoning ability.
[0174] The questions answered by the case reasoning and emergency decision-making system based on the LLM in S4 are in the form of natural language descriptions. Experts need to judge the reasoning / output effect based on the collective content of its answers. If it conforms to the actual situation and is recognized by the experts, the reply result will be saved as memory and a new case in the case library. If hallucinations occur or the results are not ideal, the experts will provide suggestions and then ask the system again. Repeat this process multiple times to achieve secondary Fine-tune. Make the Base-LLM become a Domain-LLM with professional domain knowledge.
[0175] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the overall concept of the present invention, several changes and improvements can be made, which should also be regarded as the protection scope of the present invention.
Claims
1. A safety risk control method for gravity dams based on knowledge graphs and large models, characterized in that: The method includes S1. Collect data of water conservancy project accident cases, import them into doccano to form an original dataset after basic data cleaning work, and export the data after entity and relationship annotation using doccano and convert the data format to form a final dataset; S2. Divide the dataset into a training set, a test set, and a validation set, import them into the fine-tuning script, and perform fine-tuning on the relationship and entity extraction model to obtain an entity and relationship extraction model with domain knowledge; S3. Use the fine-tuned model with domain knowledge to extract entities and relationships from all case data in the case library to obtain a case data knowledge graph; S4. Convert the obtained case data knowledge graph into natural language descriptions, and then use these texts to fine-tune the large language model to finally obtain a large language model with domain knowledge; S5. Use the engineering model to extract information from the project overview information, and use the extraction results to form natural language description questions through PromptFormulation and input them into the large language model with domain knowledge in S4; S6. The large language model answers the questions, and the expert evaluation module provides expert opinions based on the replies of the large model. In this way, the expert module can perform secondary fine-tuning on the large model, and at the same time, the correct reply content of the large model can be stored as new case data in the case library again.
2. The gravity dam safety risk control method based on a knowledge graph and a large model according to claim 1, wherein: Specifically included in S1 are: collecting data of water conservancy project accident cases, the sources covering news reports, enterprise internal documents and relevant books. After basic data cleaning work, including removing duplicate content, unifying terms and standardizing units, import the cleaned text data into doccano, create an NER task, define the entities and relationships in water conservancy project accidents, and let annotators perform entity and relationship annotation on the text to ensure the consistency and accuracy of annotation. After annotation, use a python script to convert the data exported from doccano into a json format suitable for model training.
3. The gravity dam safety risk control method based on the knowledge graph and the large model according to claim 1, characterized in that: Specifically included in S2 are: dividing the dataset into a training set, a test set, and a validation set to ensure the balanced distribution of various types of data, and adopting Baidu's UIE extraction framework as the basic tool for entity and relationship extraction; import the divided training set into the fine-tuning script, and use the Baidu UIE framework to fine-tune the relationship and entity extraction model to optimize the performance of the model in entity and relationship extraction tasks; finally, use the validation set to evaluate the performance of the fine-tuned model, adjust the hyperparameters to improve the accuracy and generalization ability of the model, and save the fine-tuned entity and relationship extraction model with domain knowledge.
4. The gravity dam safety risk control method based on a knowledge graph and a large model according to claim 1, characterized in that: Specifically included in S3 are: using the fine-tuned model to process all accident case texts in the case library, identifying the entities and categories in the cases, and extracting the relationships between entities. According to the extracted entities and relationships, construct a structured knowledge graph, including entities of multiple categories such as geographical entities, engineering structures, geological entities, time entities, numerical and measurement entities, personnel and organization entities, and event entities and their associated relationships.
5. The gravity dam safety risk control method based on a knowledge graph and a large model according to claim 1, characterized in that: S4 specifically includes: converting the knowledge graph into natural language descriptions and fine-tuning the large language model. Using a templated method, convert the entities and relationships in the knowledge graph extracted in S3 into natural language sentences. Use the converted natural language descriptions as training data, and use a fine-tuning tool to fine-tune the large language model to enhance its understanding and generation capabilities in the field of water conservancy project accidents; evaluate the performance of the fine-tuned large language model in tasks such as water conservancy project accident analysis, prevention suggestions, and emergency response, and deploy it to the actual application environment.
6. The gravity dam safety risk control method based on the knowledge graph and the large model according to claim 1, wherein: Therefore, the engineering models in S5 include a factor analysis and risk level determination model and a case similarity calculation model; The factor analysis and risk level determination model identifies the main influencing factors of the accident and gives its risk level; The case similarity calculation model finds similar accident cases.
7. The gravity dam safety risk control method based on a knowledge graph and a large model according to claim 1, wherein: The factor analysis and risk level determination model includes: (1) Combining the construction scenario of cold region gravity dams, according to expert opinions and literature reviews, select risk evaluation indicators from aspects such as construction personnel safety risks, material and equipment safety risks, construction management safety risks, construction environment safety risks, and construction technology safety risks; (2) Select k experts, E = {E1, E2, … E k} According to their different experiences, score the k experts and assign different weights. Use linguistic variables to evaluate the actual performance of risk factors and their relative importance, and convert the evaluation results into a triangular fuzzy matrix; (3) Comprehensive weighting of risk factors using FAHP and the maximum deviation method; using the comprehensive weighting method, first calculate the subjective weight α of each indicator using the AHP method i ; then calculate the objective weight of each indicator as β using the maximum deviation method i ; finally, the comprehensive weight is obtained as: w i = aα i + bβ i . Where: a is the subjective weight influence factor, b is the objective weight influence factor; rank the risk factors and their influences based on the VIKOR method, and evaluate and analyze the ranking results to obtain the final risk level.
8. A gravity dam safety risk control system based on a knowledge graph and a large model, characterized in that: The system includes: Domain model fine-tuning module: used to achieve automated information extraction, obtain structured knowledge, and generate a knowledge graph; including using engineering domain knowledge and engineering cases to fine-tune the information extraction model to improve the model's information extraction ability in the engineering domain; Knowledge graph construction module: used to structurally represent the data information of dam safety risk accident cases; including visualizing the entities and relationships of dam safety risk accident case data, displaying the accident case entities and relationships extracted by the fine-tuned domain model, and storing and managing this information for subsequent generation of natural language descriptions and reasoning; Engineering model calculation module: used to calculate the highest risk factor, safety risk evaluation level, and the case with the highest similarity to the case library according to the project overview information given by the user; including a factor analysis and risk level determination model and a case similarity calculation model, and finally form a question according to the calculation results and a pre-designed Prompt template as the input to the intelligent decision-making module; Intelligent decision-making module: used to generate a final decision based on the structured knowledge in the domain model knowledge graph and the question formed by the Prompt; including large model fine-tuning, Prompt template design, and expert opinion re-fine-tuning. In the system optimization stage, support developers to optimize the large model output according to expert opinions in this module.
9. The gravity dam safety risk control system based on a knowledge graph and a large model according to claim 8, characterized in that: The domain model fine-tuning module includes S1. Obtain safety accident cases from the network, books, and safety accident reports, and perform basic cleaning operations on them: remove special characters, extra line breaks, and spaces that affect annotation; organize each case into an independent text description; S2. Upload the preprocessed text data to doccano, and use doccano to complete the sequence annotation task. Entity tags and relationship tags need to be created before completing the annotation task; each case is an independent annotation task; the necessary entities for annotation are: water conservancy projects, safety accidents, management measures, lessons learned, and the necessary relationships for annotation are: what safety accidents "occur" in the project, what disasters are "caused" by the safety accidents, and what lessons learned are "derived" from this; after the annotation is completed, export the annotation data; S3. Process the annotation data exported in S2, convert the original data exported from doccano into samples that conform to the classification task format, and divide the training set, test set, and validation set; then use the divided dataset to retrain and fine-tune the general entity relationship extraction model, and finally obtain a model that can perform engineering domain relationship entity extraction with engineering domain knowledge; S4. Use the model trained and fine-tuned in S3 to perform entity extraction and relationship extraction on a large number of safety accident cases for knowledge fusion and ablation, construct a water conservancy safety accident knowledge graph, store it in the Node4j graph database, and visualize it; S5. Convert the structured data in the knowledge graph constructed in S4 into natural language descriptions to construct a dataset for fine-tuning the large model. The engineering model calculation module and the intelligent decision-making module include:
10. The gravity dam safety risk control system based on the knowledge graph and large model according to claim 8, characterized in that: S1. Deploy a local API to implement case reasoning and emergency decision-making. Users can obtain similar case analysis and emergency decision-making processing methods by providing project overview information and questions; S2. Import the "project overview information" input by the user in S1 into the "factor analysis and risk level determination model" and the "similarity calculation model" respectively. The two models will respectively obtain the highest-level risk factors in the project overview information and the name of the case with the highest similarity in the case library; S3. Based on the outputs of the two models obtained in S2, design a Prompt template to convert the two "isolated" word descriptions into natural language descriptions, which include information such as risk factors, risk factor levels, similar cases, and their corresponding disposal measures; S4. Based on the natural language description converted in S3 and the Prompt project, set the role identity for the large language model. Finally, consult the large language model for similar accidents or handling measures in case scenarios in the form of questions. The large language model gives a reply based on the structured knowledge constructed based on the case library through its powerful language ability and reasoning ability; S5. The questions answered by the case reasoning and emergency decision-making system based on the LLM in S4 are in the form of natural language descriptions. Experts need to judge the reasoning / output effect based on the collective content of its answers. If it conforms to the actual situation and is recognized by experts, the reply result will be saved as memory and a new case in the case library. If hallucinations or unsatisfactory results occur, experts will provide suggestions and then ask the system again. Repeat this process multiple times to achieve secondary Fine-tune. Make the Base-LLM become a Domain-LLM with professional domain knowledge.
Citation Information
Cited By
Intelligent herdsman decision-making method and system based on large language model and retrieval enhancement generation
CN121032259A
Construction scheme auditing method, system and equipment, storage medium and program product
CN121213299A
Large language model prompt learning method and system for engineering emergency response
CN121365128A
Urban highway tunnel operation period toughness safety evaluation method
CN121563312A
Safety production knowledge graph construction system and method based on large language model
CN121599069A