Water conservancy project construction hidden danger question and answer method and device
By constructing a knowledge graph of construction hazards in water conservancy projects and a multi-task learning model, the shortcomings of traditional models in identifying construction hazard terminology and adapting to small sample learning are solved, achieving efficient hazard question answering and governance decision support.
Patent Information
- Application Number
- CN202511500136.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-30
AI Technical Summary
Traditional named entity extraction models suffer from the problem of losing local features in question answering about construction hazards in water conservancy projects, which makes it impossible to effectively identify construction hazard terms. Furthermore, existing methods cannot adapt to small sample multi-task learning, affecting the accuracy of question answering.
A short text-enhanced named entity recognition model is built, high-frequency coefficients are introduced as feature probability values, a knowledge graph of construction hazards in water conservancy projects is constructed, meta-learning algorithms and prompting learning strategies are integrated, and background information of hazards is predicted through a multi-task learning model.
It improves the accuracy of Q&A on potential hazards in water conservancy project construction, addresses the issue of sparse features in short texts, accelerates model convergence, and provides more scientific and intelligent decision support for hazard management.
Smart Images

Figure CN121434342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and apparatus for answering questions about potential hazards, and more particularly to a method and apparatus for answering questions about potential hazards in water conservancy engineering construction. Background Technology
[0002] As a crucial infrastructure for national economic and social development, the safe and stable operation of water conservancy projects directly impacts the rational utilization of national water resources, the protection of the ecological environment, and the safety of life and property for the general public. However, with the rapid development of water conservancy project construction, hidden dangers have gradually emerged, becoming a major factor hindering the healthy development of water conservancy projects. Q&A regarding hidden dangers in water conservancy project construction is of great significance for improving the quality of water conservancy project construction and ensuring national water resource security and people's well-being.
[0003] Traditional named entity extraction models suffer from the drawback of losing local features in short texts. Therefore, this invention proposes a method that constructs a short text-enhanced named entity recognition model to jointly extract named entities and entity relationships. High-frequency coefficients are introduced at the feature extraction layer to enhance text features, and a knowledge graph of potential hazards in water conservancy engineering construction is built, thus improving the feature sparsity problem of short text logs related to potential hazards in water conservancy engineering construction. To better suit multi-task learning with small sample sizes, this invention proposes a hazard knowledge exploration model that integrates meta-learning algorithms and prompting learning strategies, improving the model's accuracy in multi-task prediction. Ultimately, this results in an intelligent question-answering device for the management of potential hazards in water conservancy engineering construction, providing a more scientific and intelligent technical means for construction safety management. Summary of the Invention
[0004] Purpose of the invention: The first purpose of the present invention is to provide a question-and-answer method for water conservancy engineering construction hidden dangers that improves the accuracy of question and answer, and the second purpose of the present invention is to provide a question-and-answer device for water conservancy engineering construction hidden dangers.
[0005] Technical solution: The present invention provides a water conservancy project construction hazard question and answer device, comprising a construction hazard dictionary unit, a hazard knowledge graph unit, a hazard knowledge exploration unit, and a construction hazard question and answer unit;
[0006] The construction hazard dictionary unit extracts keywords from the text corpus of construction hazard logs, calculates the correctness and effect of keyword combinations, and constructs a hazard dictionary.
[0007] The hidden danger knowledge graph unit builds a short text-enhanced named entity recognition model to jointly extract named entities and entity relationships, and constructs a knowledge graph of hidden dangers in water conservancy project construction.
[0008] The Hazard Knowledge Exploration Unit builds a multi-task learning knowledge exploration model, which integrates meta-learning algorithms and prompting learning strategies to predict hazard background information and explore new hazard knowledge.
[0009] The construction hazard Q&A unit organizes relevant information by searching a graph database and answers hazard-related questions in both knowledge graph and text formats.
[0010] Furthermore, the construction hazard dictionary unit extracts candidate hazard keywords from the text corpus of construction hazard logs, mines high-frequency feature words in the text, calculates the correctness and effect of the combination of hazard keywords, and constructs a hazard dictionary.
[0011] The hidden danger knowledge graph unit builds a short text-enhanced named entity recognition model, introduces high-frequency coefficients as feature probability values to enhance text features, and extracts named entities; adopts a joint extraction model based on parameter sharing to complete entity relationship extraction, constructs a knowledge graph of hidden dangers in water conservancy project construction, and improves the problem of sparse features in short texts of water conservancy project construction hidden danger logs.
[0012] The hazard knowledge exploration unit builds a multi-task learning knowledge exploration model, introduces a prompt learning strategy to automatically construct prompt learning templates to achieve prediction, and obtains the loss and minimum weight parameters through the inner and outer loops of the meta-learning algorithm to predict hazard background information, explore new hazard knowledge, improve the accuracy of the model for multi-task prediction, and accelerate the convergence speed of the model.
[0013] The construction hazard Q&A unit organizes relevant information by searching a graph database and answers hazard-related questions in both knowledge graph and text formats.
[0014] In summary, the construction hazard dictionary unit extracts candidate hazard keywords from the text corpus of construction hazard logs, calculates the correctness and effectiveness of keyword combinations, and constructs a hazard dictionary. The hazard knowledge graph unit builds a short text-enhanced named entity recognition model, introduces high-frequency coefficients as feature probability values to enhance text features, extracts named entities, and adopts a parameter-sharing-based joint extraction model to complete entity relationship extraction, thus constructing a knowledge graph of water conservancy project construction hazards. The hazard knowledge exploration unit builds a multi-task learning knowledge exploration model, integrating meta-learning algorithms and prompting learning strategies to predict hazard background information and explore new hazard knowledge. The construction hazard question-and-answer unit retrieves all related information from the graph database, providing auxiliary decision-making basis for water conservancy project safety hazard management in both knowledge graph and text formats.
[0015] The present invention provides a method for answering questions about potential construction hazards in water conservancy projects, comprising the following steps:
[0016] (1) The construction hazard dictionary unit preprocesses and segments the corpus based on the hazard corpus, obtains candidate hazard keywords, mines high-frequency feature words in the text, calculates the correctness and effect of the combination of hazard keywords, and constructs a hazard dictionary;
[0017] (2) The hidden danger knowledge graph unit builds a short text-enhanced named entity recognition model based on the hidden danger dictionary constructed in step (1). A feature enhancement layer is added to the model, and the features within the weighted calculation area are used as the final regional feature value. The enhanced local features and global semantic features are fused and represented as the final feature vector of the model, and named entities are extracted.
[0018] (3) Adopt a joint extraction model based on parameter sharing, classify the candidate relationships between entity pairs through an end-to-end bidirectional dependency tree structure, complete entity relationship extraction, and construct a knowledge graph of hidden dangers in water conservancy project construction.
[0019] (4) The knowledge exploration unit for hidden dangers builds a knowledge exploration model for multi-task learning based on the knowledge graph of hidden dangers in water conservancy projects constructed in steps (2) and (3), designs a template according to the downstream attribute prediction task, and modifies the format of the input hidden danger text using a prompt function;
[0020] (5) During the training process of the model, a multi-task prediction algorithm based on meta-learning is adopted. The multi-task loss and the minimum weight parameters are obtained through inner and outer loops to predict the background information of hidden dangers and explore new hidden danger knowledge.
[0021] (6) The construction hazard Q&A unit answers users’ questions in both knowledge graph and text formats by using structured data and semantic relationships in the knowledge graph.
[0022] Furthermore, in step (1), the hazard corpus is first preprocessed and segmented to obtain candidate hazard keywords. Then, the correctness and effectiveness of the combination of hazard keywords are calculated according to formula (1). i :
[0023]
[0024] in The number of times the keyword combination p1 appears is the potential hazard; N is the total number of times all phrases appear. For the combination of keywords for hidden dangers p n Number of times it appears The value represents the frequency of occurrence of the keyword w1, μ is the collocation coefficient, D is the collocation correctness, and M is the collocation effect.
[0025] By using the matching coefficient μ, the effect value M is controlled within the range of 0 to 1. Finally, a threshold is set to measure whether the combination of hidden danger keywords can be used as a terminology for construction hidden dangers, while eliminating hidden danger keyword combinations with incorrect matching.
[0026] In step (2), a new feature enhancement layer is added, which performs a convolution operation on the input vector and the introduced construction hazard dictionary, and introduces high-frequency coefficients as the regional feature probability values ω. iAfter convolution of the text vectors, a weighted pooling operation is performed on the feature matrix obtained to enhance local features according to formula (2):
[0027] C i =Conv(W w ·w i:i+h-1 )·ω i (2)
[0028] Where Conv represents the convolution operation, W w The weight matrix of the convolution kernel, w i:i+h-1 Let be the window consisting of the i-th row to the (i-h+1)-th row of the input matrix, i = 1, 2, ..., s-h+1, where s is the number of words in the input text, h is the width of the convolution kernel, and ω... i For the region feature probability value, C i These are local features.
[0029] Meanwhile, global semantic features are extracted through a bidirectional long short-term memory network model, and the two feature vectors are fused into the final feature vector of the model according to formula (3):
[0030]
[0031] Where L i For global semantic features, C i For local features, F i For the final feature vector, w i Let i = 1, 2, ..., n, where n is the total number of keywords in the hazard dictionary, and Bi-LSTM is a bidirectional long short-term memory network.
[0032] This approach integrates enhanced local and global semantic features to improve the sparse textual features of potential hazards in water conservancy projects.
[0033] In step (3), a joint extraction method based on parameter sharing is used to complete the entity relation extraction task. The same embedding layer and encoding layer are used as in the named entity recognition task, only the decoder needs to be used. During the decoding process, the last Chinese character predicted as an entity is selected, all possible combinations are incrementally constructed, relation candidate vectors are used to predict relation labels, and the relations between entities are classified.
[0034] In step (4), to better suit learning with small samples, cue learning is integrated into the model embedding layer. First, a template is designed based on the downstream hazard attribute prediction task, and a cue function f is used. prompt (x) Modify the format of the input hazard text x, as shown in formula (4):
[0035]
[0036] Then, the revised hidden danger text The input is fed into the training model, and based on the learned semantic representation, the model obtains the highest score in a cloze test task. By automatically constructing prompt templates to change the text input format, the difference between the pre-training task and the downstream prediction task is reduced, thus improving the model's running efficiency.
[0037] After completing the mask prediction task in the model's encoding layer, each input text will obtain a predicted word that is semantically similar to the correct answer, but this predicted word does not belong to the category label set. Therefore, it is necessary to convert the predicted word into a category label through the label mapping function f, as shown in formula (5):
[0038]
[0039] Wherein, for a given input text sequence x = {x1, x2, ..., x...} n After the MASK prediction task, a set of predicted words V will be output. y ={v1, v2, ..., v n Each predicted word can be transformed into a category label y∈Y={y1, y2, ..., y} through a label mapping function. m}, where m is the length of the category label sequence, then the category prediction problem is successfully transformed into selecting the category with the highest probability.
[0040] In step (5), a multi-task prediction algorithm based on meta-learning is used during the model training process to divide the collected data into training set D. train and support set D support The model is trained by iterating through inner and outer loops on the attribute prediction task of hidden danger background information. The training set data D is used in the inner loop. train Gradient descent is performed for each attribute prediction task. The model parameters are updated by using the construction hazard source marking data in the "Safety Risk Library". The initial values W of the weight matrices of the shared model for the four prediction tasks "source of safety hazard", "potential accident type", "level of safety hazard" and "whether it is a major accident" are learned and shared, and are denoted as W1, W2, W3 and W4 respectively. The initial values of the prediction task weight matrices are calculated according to formula (6).
[0041]
[0042] Among them W i For Task i A better weight matrix, W is the initial value of the prediction task weight matrix, α is the learning rate of the task parameters, and Loss is... i (f W ) for Task iThe task loss function.
[0043] In the outer loop, the support set data D support The loss is minimized iteratively using the stochastic gradient descent method. By using the shared gradient experience information W1, W2, W3, and W4 of the model, the model's weight matrix is updated based on the original initial parameters W, as shown in formula (7).
[0044]
[0045] Among them W min The weight matrix that minimizes the prediction task loss, W is the initial value of the prediction task weight matrix, β is the meta-parameter learning rate, and Loss is... i (f w ) for Task i The task loss function.
[0046] This enables the model to achieve rapid convergence in multi-task learning scenarios, allowing the model to obtain the weight matrix W that minimizes the loss of attribute prediction tasks within the hazard background information. min .
[0047] In step (6), the hidden danger knowledge graph displays knowledge and relationships related to construction hidden dangers in the water conservancy field. Users raise questions, and the knowledge graph answers the users' questions in both knowledge graph and text forms through structured data and semantic relationships in the knowledge graph.
[0048] Beneficial Effects: Compared with existing water conservancy hazard question-and-answer methods, this invention has the following significant advantages: By calculating the correctness and effect of the combination of hazard keywords, it solves the defect of being unable to identify construction hazard terms, providing data support for the subsequent construction of a hazard knowledge graph using deep learning models; It builds a short text-enhanced named entity recognition model to jointly extract named entities and entity relationships, and introduces high-frequency coefficients as feature probability values on the basis of the hazard dictionary, emphasizing hazard terms in the business domain and improving the problem of sparse short text features; It builds a multi-task learning knowledge exploration model, which quickly obtains model parameters for multi-task learning scenarios through the inner and outer loops of the meta-learning algorithm, and introduces a prompt learning strategy to automatically construct prompt learning templates, improving the accuracy of predicting water conservancy hazard background information. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the specific implementation of the present invention.
[0050] Figure 2 This is a diagram illustrating the overall architecture of the short text-enhanced named entity recognition model of the present invention.
[0051] Figure 3This is the overall architecture diagram of the multi-task learning knowledge exploration model of the present invention. Detailed Implementation
[0052] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0053] The core of this invention is a question-and-answer method and device for potential construction hazards in water conservancy projects. By calculating the correctness and effectiveness of keyword combinations for potential hazards, it solves the problem of being unable to identify construction hazard terminology. A short-text-enhanced named entity recognition model is built to jointly extract named entities and entity relationships, introducing high-frequency coefficients as feature probability values to improve the sparsity problem of short-text features. A multi-task learning knowledge exploration model is built, integrating meta-learning algorithms and prompting learning strategies, improving the model's accuracy in multi-task prediction and accelerating its convergence speed. Figure 1 As shown in the flowchart, the water conservancy project construction hazard question and answer device of the present invention includes a construction hazard dictionary unit, a hazard knowledge graph unit, a hazard knowledge exploration unit, and a construction hazard question and answer unit.
[0054] The construction hazard dictionary unit extracts candidate hazard keywords from the text corpus of construction hazard logs, calculates the correctness and effect of keyword combinations, and constructs a hazard dictionary.
[0055] The hidden danger knowledge graph unit builds a short text-enhanced named entity recognition model, introduces high-frequency coefficients as feature probability values to enhance text features, and extracts named entities; adopts a joint extraction model based on parameter sharing to complete entity relationship extraction, constructs a knowledge graph of hidden dangers in water conservancy project construction, and improves the problem of sparse features in short texts of water conservancy project construction hidden danger logs.
[0056] The hazard knowledge exploration unit builds a multi-task learning knowledge exploration model that integrates meta-learning algorithms and prompting learning strategies to predict hazard background information, explore new hazard knowledge, improve the model's accuracy in multi-task prediction, and accelerate the model's convergence speed.
[0057] Figure 1 This is a flowchart illustrating a specific implementation of the present invention, including the following steps:
[0058] Step 1: The construction hazard dictionary unit extracts candidate hazard keywords from the text corpus of construction hazard logs, uses adjacent word frequency phrase analysis to calculate the correctness and effect of the combination of hazard keywords, and constructs a hazard dictionary.
[0059] This example uses the construction hazard log of a water conservancy project, as shown in Table 1.
[0060] The preprocessed corpus of hazard log text is read, and all possible hazard keywords are identified using the Jieba word segmentation prefix dictionary. The word frequencies of all hazard keywords are accumulated to generate a directed acyclic graph. Based on the accumulated word frequency values, the maximum path probability is calculated bottom-up to read the candidate hazard keyword set. The TF-IDF algorithm is used to obtain high-frequency construction hazard keywords, and the TF value of candidate phrases is calculated to determine whether they are high-frequency words.
[0061] Table 1 Construction Log of Potential Hazards
[0062]
[0063] Whether a candidate hazard keyword is a relevant term in the hazard field is determined by calculating its inverse document frequency (IDF) value, as shown in formulas (8) to (11):
[0064]
[0065] TF-IDF = TF i *IDF i (11)
[0066] Where n i The number of times the word phrase i appears, ∑ k n k The sum of the occurrences of all phrases is represented by N, where N is the sum of the number of documents in the document library, and phrase i appears in N. i TF appeared in the document i This represents the importance of term i in the existing document, IDF i This represents the frequency of term i. To prevent the denominator from being 0, formula (9) has been smoothed, as shown in formula (10).
[0067] The sentence is divided into n short word sequences S, as shown in formula (12); the positive maximum matching algorithm is used to combine the hidden danger keyword w1 with other hidden danger keywords (such as w2, w2 and w3) to form a candidate hidden danger keyword sequence P, as shown in formula (13), where p1 = w1, p2 = w1 + w2, p3 = w1 + w2 + w3, and so on.
[0068] S = {w1, w2, w3, ..., w} n} (12)
[0069] P = {p1, p2, p3, ..., p} n} (13)
[0070] Then, the correctness of the keyword combination is analyzed by counting the number of times it appears in the corpus; the collocation effect is evaluated by analyzing the word frequency ratio before and after the keyword combination, and the collocation correctness and collocation effect E of the keyword combination are calculated. i .
[0071] By controlling the effect value M within the range of 0 to 1 using the matching coefficient μ, a threshold is finally set to measure whether the combination of hidden danger keywords can be used as a terminology for construction hidden dangers. At the same time, hidden danger keyword combinations with incorrect matching are eliminated, and high-frequency hidden danger keywords and construction hidden danger terms are extracted to construct a dictionary of construction hidden dangers in the field of water conservancy engineering.
[0072] Step 2: The Hidden Danger Knowledge Graph Unit builds a short text-enhanced named entity recognition model based on the hidden danger dictionary constructed in Step 1. A feature enhancement layer is added to the model, and high-frequency coefficients are introduced as feature probability values. The enhanced local features and global semantic features are fused and represented to extract named entities.
[0073] The knowledge graph schema layer was designed, and entity concept categories were defined as shown in Table 2. At the same time, entity relationship categories were defined, connecting entities that were independently scattered in the graph structure into a directed graph structure for the construction hazard knowledge graph. The construction operation entity "excavation construction" was used as the named entity h, as shown in Table 3.
[0074] Table 2 Examples of Entity Concept Categories
[0075]
[0076] Table 3 Examples of Entity Relationship Categories
[0077]
[0078] like Figure 2 As shown in the model architecture diagram, the short text-enhanced named entity recognition model includes an embedding layer, an encoding layer, and a decoding layer. The embedding layer consists of two modules: a text encoder (T-Encoder) and a knowledge encoder (K-Encoder). The former is responsible for structured encoding to capture semantic information from the text, while the latter is responsible for fusing heterogeneous information to ultimately obtain a token sequence and an entity sequence. The encoding layer extracts enhanced local features and global semantic features through two modules. The former performs scanning convolution operations on the token input vector and the introduced construction hazard dictionary, both using convolution kernels of size (2, 3, 4). The convolution kernel used for the introduced dictionary is W. t = [1, 1, 1, ..., 1], as shown in formula (14), if w in the token sequence i If the word or phrase is a keyword in the hazard dictionary, then T iThe value is the TF-IDF value; if it is not a keyword, then T i The value is 0.5, and the threshold of high-frequency words is taken as the basic value.
[0079]
[0080] in For high-frequency coefficients, Conv represents the convolution operation, and W represents the high-frequency coefficients. t For convolution kernel, T i:i+h-1 Let s be the window consisting of the i-th row to the (i-h+1)-th row of the input matrix, i = 1, 2, ..., s-h+1, where s is the number of words in the input text and h is the width of the convolution kernel.
[0081] The high-frequency coefficients obtained after processing the above convolution operation Normalization is performed to obtain the region feature probability value ω. i As shown in formula (15), where The high-frequency coefficients are derived by convolving the TF-IDF values of potential hazard keywords, ω. i This represents the probability value of regional features.
[0082]
[0083] Introducing high-frequency coefficients as the regional feature probability value ω i Weighted pooling is performed on the feature matrix obtained after convolution of text vectors to improve the model's computation speed during downsampling, resulting in enhanced local features C. i .
[0084] Simultaneously, global semantic features L are extracted using a bidirectional long short-term memory network model. i The two feature vectors are fused together to form the final feature vector F of the model. i .
[0085] The encoding layer outputs a feature vector, and the decoding layer predicts the corresponding label using the feature vector. The model uses a Conditional Random Field (CRF) as the final output function. Conditional probabilities are used to add constraints to the final label classification to ensure that the predicted entity label sequence is valid.
[0086] Step 3: Adopt a joint extraction model based on parameter sharing, classify candidate relationships between entity pairs through an end-to-end bidirectional dependency tree structure, complete entity relationship extraction, and construct a knowledge graph of hidden dangers in water conservancy project construction.
[0087] Similar to the named entity recognition task, this method employs the same embedding and encoding layers, using a bidirectional tree structure to acquire nearby relevant dependency information to determine candidate relationships between entity pairs. During decoding, the last Chinese character predicted as an entity is selected, and all possible combinations are incrementally constructed. Relationship candidate vectors are then used to predict relationship labels, and the relationships between entities are classified.
[0088] Finally, named entities and entity relationships are stored in the Neo4j graph database to realize the construction of a knowledge graph of potential hazards in water conservancy engineering construction.
[0089] Step 4: The Hazard Knowledge Exploration Unit builds a multi-task learning knowledge exploration model based on the water conservancy project hazard knowledge graph constructed in Steps 2 and 3. It designs a template according to the downstream attribute prediction task and uses a prompt function to modify the format of the input hazard text. To assist project safety managers in achieving more scientific and intelligent hazard management, the background information attributes of hazards are defined by analyzing the hazard log storage structure and referring to construction monitoring parameter settings.
[0090] like Figure 3 As shown in the model architecture diagram, the knowledge exploration model for multi-task learning includes a prompt learning embedding layer, an MLM encoding layer, and a label mapping layer. The embedding layer replaces traditional text with embedding vectors as prompt templates, transforming the template construction task into a continuous parameter optimization task. It uses the BERT pre-trained model to select the set of embedding vectors with the highest confidence as prompt templates, as shown in formula (16):
[0091] E(P 0:n ) = argmin p Loss(BERT(E(x), E(y))) (16)
[0092] Where E(X) represents the vector representation of the hazard log description, and the solvable template representation is E(P) = {P0, P1, ..., P}. n}, where E(y) is the classification label sequence to be predicted.
[0093] By selecting P with the minimum loss i The system automatically constructs a prompt template vector using a loss function. Then, it modifies the format of the input potential hazard text using a prompt function and inputs it into the training model. Based on the learned semantic representation, it obtains the highest-scoring output in the prediction task of the coding layer mask.
[0094] After completing the MASK prediction task in the model encoding layer, each input text will obtain a predicted word that is semantically similar to the correct answer. In the label mapping layer, the predicted word needs to be converted into a category label through the label mapping function f.
[0095] Step 5: During the training process of the hazard knowledge exploration model, a multi-task prediction algorithm based on meta-learning is adopted. By obtaining the multi-task loss and the weight parameters with the minimum through inner and outer loops, the background information of hazards is predicted and new hazard knowledge is explored.
[0096] The collected data is divided into training set D. train and support set D support The model is trained by iterating through inner and outer loops on the attribute prediction task of hidden danger background information. The training set data D is used in the inner loop. train Gradient descent is performed for each attribute prediction task. The model parameters are updated by using the construction hazard source marking data in the "Safety Risk Library". The initial values W of the weight matrices of the shared model for the four prediction tasks "source of safety hazard", "potential accident type", "level of safety hazard" and "whether it is a major accident" are learned and shared, and are denoted as W1, W2, W3 and W4 respectively. The initial values of the prediction task weight matrices are calculated.
[0097] In the outer loop, the support set data D support The loss is minimized iteratively using the stochastic gradient descent method. The model's weight matrix is updated based on the original initial parameters W by using the shared gradient empirical information W1, W2, W3, and W4.
[0098] This enables the model to achieve rapid convergence in multi-task learning scenarios, allowing the model to obtain the weight matrix W that minimizes the loss of attribute prediction tasks within the hazard background information. min .
[0099] Step 6: The front-end of the construction hazard Q&A unit uses the React framework, and the back-end uses the Spring Boot framework. Both the front-end and back-end communicate using JSON format. The databases used are Neo4j graph database and MySQL relational database. The former can display hazard knowledge in the form of a directed graph, clearly showing the relationships between hazard knowledge points; the latter can store other related data. The Neo4j graph database is used to visualize the domain knowledge graph. All related information is retrieved from the graph database, and questions are answered in both knowledge graph and text formats.
[0100] Taking the question "What precautions should be taken during slope excavation?" as an example, the system's business logic layer matches the keyword "slope excavation" in the question, searches the graph database to find all relevant hazard knowledge nodes, and exports an image to display to the project safety manager. Simultaneously, it further reads the hazard warning knowledge and background information contained in the nodes, compiling them into a text list to provide users with a basis for hazard management. Hazards predicted to potentially cause major accidents are marked in red, providing users with warnings, such as intelligent suggestions: Slope excavation operations require warning signs; may lead to collapse accidents; risk level is I; pay special attention to areas without warning signs during safety inspections.
Claims
1. A water conservancy construction hidden danger question and answer device, characterized in that, The device comprises a construction hidden danger dictionary unit, a hidden danger knowledge graph unit, a hidden danger knowledge exploration unit and a construction hidden danger question and answer unit; The construction hidden danger dictionary unit extracts keywords from the text corpus of the water conservancy project construction hidden danger log, calculates the collocation correctness and collocation effect of hidden danger keyword combinations, and constructs a hidden danger dictionary; The hidden danger knowledge graph unit builds a short text enhanced named entity recognition model to jointly extract named entities and entity relationships, and constructs a water conservancy project construction hidden danger knowledge graph; The hidden danger knowledge exploration unit builds a multi-task learning knowledge exploration model that combines meta-learning algorithms and prompt learning strategies to predict hidden danger background information and explore new hidden danger knowledge; The construction hidden danger question and answer unit answers hidden danger knowledge questions in the form of knowledge graph and text by retrieving related information from the graph database.
2. The water conservancy construction hidden danger question and answer device according to claim 1, characterized in that: The construction hidden danger dictionary unit extracts candidate hidden danger keywords from the text corpus of the construction hidden danger log, mines high-frequency feature words of the text, calculates the collocation correctness and collocation effect of hidden danger keyword combinations, and constructs a hidden danger dictionary, solving the defect that professional terms cannot be recognized in text mining.
3. The water conservancy construction hidden danger question and answer device according to claim 1, characterized in that: The hidden danger knowledge graph unit builds a short text enhanced named entity recognition model, introduces high-frequency coefficients as feature probability values to enhance text features, and extracts named entities; a joint extraction model based on parameter sharing is adopted to complete entity relationship extraction, and a water conservancy project construction hidden danger knowledge graph is constructed to improve the problem of sparse short text features.
4. The water conservancy construction hidden danger question and answer device according to claim 1, characterized in that: The hidden danger knowledge exploration unit builds a multi-task learning knowledge exploration model, introduces a prompt learning strategy to automatically construct a prompt learning template for prediction, and obtains multi-task loss and minimum weight parameters through inner and outer loops of the meta-learning algorithm to predict hidden danger background information and explore new hidden danger knowledge.
5. The water conservancy construction hidden danger question and answer device according to claim 1, characterized in that: The construction hidden danger question and answer unit answers hidden danger knowledge questions in the form of knowledge graph and text by retrieving related information from the graph database.
6. A method for asking and answering hidden troubles in water conservancy construction, characterized in that, The method comprises the following steps: (1) The construction hidden danger dictionary unit preprocesses and segments the corpus according to the hidden danger corpus, obtains candidate hidden danger keywords, mines high-frequency feature words of the text, calculates the collocation correctness and collocation effect of hidden danger keyword combinations, and constructs a hidden danger dictionary; (2) The hidden danger knowledge graph unit builds a short text enhanced named entity recognition model according to the hidden danger dictionary constructed in step (1), adds a feature enhancement layer to the model, introduces high-frequency coefficients as feature probability values, and calculates the weighted features in the region as the final regional feature values. The enhanced local features and global semantic features are fused to represent the final feature vector of the model, and the named entities are extracted; (3) A joint extraction model based on parameter sharing is adopted to classify the candidate relationships between entity pairs through an end-to-end bidirectional dependency tree structure, complete entity relationship extraction, and construct a water conservancy project construction hidden danger knowledge graph; (4) The hidden danger knowledge exploration unit builds a multi-task learning knowledge exploration model according to the water conservancy project hidden danger knowledge graph constructed in steps (2) and (3), designs a template according to the downstream attribute prediction task, and modifies the format of the input hidden danger text with a prompt function; (5) In the training process of the model, a multi-task prediction algorithm based on meta-learning is adopted to obtain multi-task loss and minimum weight parameters through inner and outer loops, predict hidden danger background information, and explore new hidden danger knowledge; (6) The construction hidden danger question and answer unit is based on the hidden danger knowledge sorted by the construction hidden danger knowledge graph in the field of water conservancy engineering, which contains hidden danger early warning knowledge and hidden danger background information. Through searching the graph database, all related information is sorted to provide the basis for auxiliary decision-making for water conservancy engineering safety hidden danger governance in the form of knowledge graph and text.
7. The method according to claim 6, wherein In step (2), in order to improve the problem of sparse features of hidden danger text, a feature enhancement layer is added, and convolution operation is performed on the input vector and the introduced construction hidden danger dictionary, and high frequency coefficients are introduced as regional feature probability values ω i The feature matrix obtained after the convolution operation on the text vector is subjected to weighted pooling operation, and the local features are enhanced according to formula (1): C i = Conv(W w ·w i:i+h-1 )·ω i (1) wherein Conv denotes a convolution operation, W w is a convolution kernel, w i:i+h-1 is a window consisting of the i-th to i-h+1-th rows of the input matrix, i = 1, 2,..., s-h+1, s is the number of words of the input text, h is the width of the convolution kernel, ω i is a region feature probability value, C i is a local feature; At the same time, the global semantic features are extracted by the bidirectional long short-term memory network model, and the two feature vectors are fused into the final feature vector of the model according to formula (2): where L i is the global semantic feature, C i is the local feature, F i is the final feature vector, w i is the keyword in the hidden danger dictionary, i = 1, 2, …, n, n is the total number of keywords in the hidden danger dictionary, and Bi-LSTM is a bidirectional long short-term memory network.
8. The method according to claim 6, wherein In step (4), in order to be more suitable for small sample learning, prompt learning is fused in the model embedding layer. First, a template is designed according to the downstream hidden danger attribute prediction task, and a prompt function f prompt (x) modifying the format of the input hidden danger text x, as shown in formula (3): Then, the modified hazard text is input into the training model, the output with the highest score in the completion task is obtained according to the learned semantic representation, and finally the predicted word is converted into a category label through a label mapping function.
9. The method according to claim 6, wherein In step (5), in the training process of the model, a meta-learning-based multi-task prediction algorithm is used to divide the collected data into a training set D train and a support set D support , and the model is trained through inner loop and outer loop iterations on the attribute prediction task of the hidden danger background information. Using the training set data D in the inner loop train Gradient descent is performed for each attribute prediction task, learning the initial values of the shared model weight matrices W1, W2, W3, W4 for the 4 prediction tasks, respectively, according to formula (4): where W i is the optimal weight matrix for Task i , W is the initial value of the predicted task weight matrix, a is the task parameter learning rate, Loss i (f w ) is the task loss function for Task i . In the outer loop, support set data D support The random gradient descent method is used to obtain the minimum value of the loss iteratively, and the weight matrix of the model is updated on the basis of the original initial parameters W by using the gradient experience information W1, W2, W3, W4 of the shared model, and the weight matrix of the model is updated on the basis of the original initial parameters, as shown in formula (5): where W min is the predicted task loss and the minimum weight matrix, W is the initial value of the predicted task weight matrix, β is the meta-parameter learning rate, Loss i (f w ) is the task loss function of Task i .