Mountain flood disaster risk level prediction method based on knowledge graph and graph neural network
By constructing a mountain torrent disaster risk level prediction model based on knowledge graphs and graph neural networks, the problems of high data quality requirements and difficult to determine model parameters in the existing technology are solved, and more accurate and reliable risk level prediction is achieved, reducing the risk loss of mountain torrent disasters.
Patent Information
- Application Number
- CN202510222020.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing methods for predicting the risk level of the mountain torrent disasters have problems such as high data quality requirements and difficult to determine the model parameters, which leads to deviations in the risk level prediction and increases the risk loss of the mountain torrent disasters.
Using a method based on knowledge graph and graph neural network, a risk level prediction model is constructed that combines risk knowledge graph and graph neural network, entities and relationships in the data are extracted, and the graph neural network is used to learn and aggregate node features to achieve risk level prediction.
It improves the accuracy and reliability of the prediction of mountain torrent disaster risk level, reduces risk losses, and provides scientific basis and decision-making support for disaster prevention and mitigation work.
Smart Images

Figure CN119721722B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of risk level prediction, and specifically to a method for predicting the risk level of mountain flood disasters based on a knowledge graph and a graph neural network. Background Art
[0002] As a natural disaster with strong suddenness and great destructiveness, mountain flood disasters pose a huge threat to people's lives and property safety. Accurately predicting the risk level of mountain flood disasters is of great significance for taking effective disaster prevention and mitigation measures in advance and reducing disaster losses. In recent years, many scholars have been committed to studying methods for risk assessment of mountain flood disasters. Traditional methods mainly include methods based on hydrological models, methods based on empirical formulas, and methods based on index systems. These methods can evaluate the risk of mountain flood disasters to a certain extent, but there are problems such as large data requirements, complex calculations, and difficulty in considering the interaction of multiple factors.
[0003] The reference patent is titled: Method for Risk Zoning and Prediction of Mountain Flood Disasters Based on GIS-Neural Network Integration (Patent Publication No.: CN108280553A, Patent Publication Date: July 13, 2018), which includes: S1. Using association rules to mine the association relationship between risk factors and risk levels in mountain flood disasters, identifying risk factors, and constructing a quantitative risk assessment index system for mountain flood disasters; S2. Using the analytic hierarchy process to determine the hazard and vulnerability index systems and their weights, and generating each element layer; S3. Using ArcGIS to overlay the mountain flood disaster hazard and vulnerability distribution layers to obtain a mountain flood disaster risk distribution map; S4. Using the ISO maximum likelihood method for clustering and a method combining bottom-up regional merging and top-down qualitative analysis to form a mountain flood disaster risk zoning; S5. Using an Elman neural network to analyze the non-linear relationship between evaluation indicators, risk levels, and disaster situation data, and constructing a mountain flood disaster risk assessment and loss prediction model.
[0004] Based on the description of the above document, although the existing risk level prediction methods can obtain and analyze data more accurately and improve the accuracy of risk assessment, there are still some limitations, such as high requirements for data quality and difficulty in determining model parameters, so that there are deviations in the operation of risk level prediction, resulting in greater risk losses of mountain flood disasters. Therefore, the present invention provides a method for predicting the risk level of mountain flood disasters based on a knowledge graph and a graph neural network. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a method for predicting the risk level of mountain flood disasters based on a knowledge graph and a graph neural network, which solves the problems that although the existing risk level prediction methods can obtain and analyze data more accurately and improve the accuracy of risk assessment, there are still some limitations, such as high requirements for data quality and difficulty in determining model parameters, resulting in deviations in the risk level prediction operation and greater risk losses of mountain flood disasters.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for predicting the risk level of mountain flood disasters based on a knowledge graph and a graph neural network, specifically including the following steps:
[0007] A1. Collect data related to the risk of mountain flood disasters and perform data preprocessing operations;
[0008] A2. Construct and train a risk level prediction model combining a knowledge graph and a graph neural network. The specific operations are as follows:
[0009] a21. Extract entities and relationships in the data, and construct a risk knowledge graph through operations such as node definition, edge definition, knowledge fusion, and graph update;
[0010] a22. Based on the constructed knowledge graph, use a graph neural network to learn and aggregate node features, and extract feature information for risk level prediction;
[0011] a23. Use operations such as forward propagation, loss calculation, parameter update, and result evaluation to obtain the prediction result of the risk level and achieve multiple trainings;
[0012] A3. Introduce parameter data of the area to be measured into the risk level prediction model for risk level prediction.
[0013] Preferably, the data preprocessing operation in A1 is as follows:
[0014] a11. Clean the data, remove missing values and outliers in the data, and complete the compensation for missing values and outliers by filling or deleting;
[0015] a12. Convert the categorical variables in the remaining data into numerical features that can be used for subsequent model training through label encoding;
[0016] a13. Standardize the numerical features so that the processed numerical features conform to the standard normal distribution;
[0017] a14. Select numerical features with high contribution to model training through feature selection to optimize the model performance.
[0018] Preferably, the formula for standardizing the numerical features in a13 is:
[0019] ;
[0020] Among them, X norm represents the standardized features, X b represents the original feature, α and λ represent the mean and standard deviation values of the feature, respectively, to determine the reduction in variance of the numerical feature.
[0021] Preferably, the operation of constructing the risk knowledge graph in a21 is:
[0022] B1. Identify each entity in the data and assign a unique identifier to define it as a node in the graph;
[0023] B2. Define the edges between nodes, implement the association operations between nodes, and form the infrastructure of the risk knowledge graph by constructing formulas;
[0024] B3. Through the knowledge fusion step, extract external impact data and integrate it into the knowledge graph to improve the risk knowledge graph architecture;
[0025] B4. Add the preprocessed new data to the risk knowledge graph, update the attributes of the nodes and the connection information of the edges, delete outdated data or no longer relevant relationships, and realize real-time update of the graph.
[0026] Preferably, the formula constructed in B2 is:
[0027] G = (V, E);
[0028] G represents the risk knowledge graph, V represents the node set, and E represents the edge set;
[0029] For the association operation between nodes, it is necessary to determine the relationship similarity between the nodes to determine the edge weight and the association direction, and the specific determination operation is:
[0030] b21. Select any two nodes and first determine whether the two nodes are of the same type of regional location features. If they are of the same type, it is necessary to determine whether the regional locations are adjacent, and then introduce geographic information to determine the risk flow direction to point to the flow node;
[0031] b22. If the features are not of the same type and one node is a regional location feature, the association direction is transmitted from the current node to another node, and the nodes that need to be constructed related to the current node are summarized and marked as v n {v 1 , v 2 , …, v n}, determine the contribution impact degree based on the economic loss value caused by each individual node feature in the historical data, determine the influence order of each node feature according to the magnitude of the economic loss value, and the one with the largest value has the greatest influence. Then, determine the edge weight according to the sorted node features;
[0032] b23. Thus, a complete risk knowledge graph architecture is constructed.
[0033] Preferably, the calculation formula for the edge weight in b22 is:
[0034] W n(n-1) = sim(v n , v n-1 );
[0035] W n(n-1) represents the edge weight between node v n and node v n-1 , v n represents the nth node, and sim(v n , v n-1 ) represents the similarity between nodes;
[0036] And sim(v n , v n-1 ) = [(k 1 + k 2 +…+ k n ) / n] / k n ;
[0037] k n represents the nth loss value after sorting by the magnitude of the economic loss value, and (k 1 + k 2 +…+ k n ) / n is the average economic loss value of all influencing node features.
[0038] Preferably, the operation of using a graph neural network to learn and aggregate node features in a22 is as follows:
[0039] C1. Use the node feature matrix and the adjacency matrix as the input layer of the risk level prediction model to ensure that the graph neural network can capture the relationships between nodes;
[0040] C2. Set the GCN layer to perform layer-by-layer aggregation and update of node features. Each layer of GCN uses the adjacency matrix and the degree matrix to aggregate the information of neighbor nodes;
[0041] C3. The graph neural network learns the local features of nodes and maps the updated node feature matrix to a specific risk level in the output layer.
[0042] Preferably, the expression for updating the node features in C2 is:
[0043] H (l+1) = σ(D 1 / 2 AD 1 / 2 H (l) W (l) );
[0044] And H (l+1) represents the node feature matrix of the (l + 1)-th layer, H (l) represents the node features of the adjacent l-th layer, the matrix A represents the adjacency matrix, D represents the degree matrix, W (l) represents the weight matrix of the l-th layer, and σ represents the activation function;
[0045] And the node feature update formula is:
[0046] ;
[0047] where, h i represents the updated feature of node i, N(i) represents the set of neighbor nodes of node i, d i and d j represent the degrees of node i and node j respectively;
[0048] And the mapping formula for the risk level in C3 is:
[0049] Y = softmax(H (l) W (l+1) );
[0050] And W (l+1) is the weight matrix of the output layer, Y is the risk level prediction result of the node, each row represents the predicted probability distribution of a node, and softmax is a combination of a fully connected layer and a classifier to implement the mapping operation of the risk level.
[0051] Preferably, the specific operation of the risk level prediction result in a23 is:
[0052] D1. Output the risk prediction result based on the graph neural network and perform backpropagation;
[0053] D2. Use the cross-entropy loss function to calculate the error between the prediction result and the true label, and implement the compensation operation between the prediction result and the true label. And the specific loss function formula is:
[0054] ;
[0055] N represents the number of samples, C represents the number of categories, y gc represents the true label of sample g, represents the predicted probability of sample g;
[0056] Then, the model parameters are updated through the gradient descent algorithm to minimize the loss function, which is formulated as:
[0057] θ new =θ old -η(▽ θ )L;
[0058] θ new is the new model parameter, θ old is the old model parameter, η represents the learning rate, (▽ θ )L represents the gradient of the loss function with respect to the parameters;
[0059] D3. Use accuracy, precision, recall and F1 score to evaluate the performance of the risk level prediction model.
[0060] Preferably, the operation of transferring the parameter data of the area to be tested in A3 to the risk level prediction model for risk level prediction is:
[0061] a31. After collecting parameter data of the area to be tested, the data is cleaned and processed by the risk level prediction model, and then a current risk level prediction model for the area to be tested is established based on the knowledge graph and graph neural network;
[0062] a32. Analyze and output the risk results of the area to be tested based on the risk level prediction model, and adjust the prevention strategy based on the predicted risk results.
[0063] The present invention provides a method for predicting the risk level of flash flood disasters based on knowledge graph and graph neural network. Compared with the prior art, it has the following beneficial effects:
[0064] 1. This method for predicting the risk level of flash flood disasters based on knowledge graph and graph neural network extracts entities and relationships in the data, constructs a risk knowledge graph using node definition, edge definition, knowledge fusion and graph update operations, and uses graph neural network to learn and aggregate node features based on the constructed knowledge graph, and extracts feature information for risk level prediction. Therefore, the knowledge graph can represent entities, relationships and attributes in the form of a graph, and can effectively integrate and express complex knowledge and information. The graph neural network can effectively process graph structured data, capture the relationships and dependencies between nodes, and realize the modeling and analysis of complex systems. The introduction of knowledge graph and graph neural network into the prediction of flash flood disaster risk level aims to make full use of multi-source data, explore potential risk patterns and laws, improve the accuracy and reliability of prediction, and provide a scientific basis and decision-making support for disaster prevention and mitigation.
[0065] 2. The flash flood disaster risk level prediction method based on knowledge graph and graph neural network defines the edges between nodes to implement the association operation between nodes, constructs the basic architecture of the risk knowledge graph by building formulas, determines the contribution influence degree according to the economic loss value caused by the characteristics of each single node in historical data, determines the influence order of each node feature according to the magnitude of the economic loss value, and the one with the largest value has the greatest influence. Then, the edge weights are determined according to the sorted node features, so that the knowledge graph and the relationship between nodes in the associated knowledge graph can be quickly established, and the edge weight values are obtained through the judgment and training of historical data. With the continuous influx of new data, the graph structure is dynamically updated through the graph update mechanism to ensure its timeliness and accuracy.
[0066] 3. The flash flood disaster risk level prediction method based on knowledge graph and graph neural network takes the node feature matrix and the adjacency matrix as the input layer of the risk level prediction model, sets the GCN layer to aggregate and update the node features layer by layer. Each layer of GCN uses the adjacency matrix and the degree matrix to aggregate the information of neighbor nodes. The graph neural network learns the local features of nodes, and maps the updated node feature matrix to the specific risk level in the output layer, so as to gradually adapt to the complexity of the data through multiple adjustments during the training process, the model parameters are gradually converged and finally tend to be stable, so that the model reaches the optimal state and better performs the risk level prediction operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is the operation flow chart of the disaster risk level prediction method of the present invention;
[0068] Figure 2 is the operation flow chart of constructing the risk knowledge graph of the present invention;
[0069] Figure 3 is the model training loss curve graph of the present invention;
[0070] Figure 4 is the ROC curve graph of the model training process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0072] Please refer to Figures 1 - 4 , the present invention provides two technical solutions:
[0073] Embodiment 1: A method for predicting flash flood disaster risk level based on knowledge graph and graph neural network, specifically comprising the following steps:
[0074] A1. Collect data related to flash flood disaster risks and perform data preprocessing operations;
[0075] A2. Build and train a risk level prediction model combining knowledge graph and graph neural network. The specific operations are as follows:
[0076] a21. Extract entities and relationships from the data, and construct a risk knowledge graph using node definition, edge definition, knowledge fusion, and graph update operations;
[0077] a22. Based on the constructed knowledge graph, the graph neural network is used to learn and aggregate node features, and feature information is extracted for risk level prediction;
[0078] a23. Utilize forward propagation, loss calculation, parameter update and result evaluation to obtain the predicted results of risk level and implement multiple trainings;
[0079] A3. Introduce parameter data of the area to be tested into the risk level prediction model to predict the risk level.
[0080] By extracting entities and relationships in the data, a risk knowledge graph is constructed using node definition, edge definition, knowledge fusion and graph update operations. Based on the constructed knowledge graph, graph neural network is used to learn and aggregate node features, and feature information is extracted for risk level prediction. Therefore, the knowledge graph can represent entities, relationships and attributes in the form of a graph, and can effectively integrate and express complex knowledge and information. Graph neural network can effectively process graph structured data, capture the relationships and dependencies between nodes, and realize the modeling and analysis of complex systems. The knowledge graph and graph neural network are introduced into the prediction of flash flood disaster risk level, aiming to make full use of multi-source data, explore potential risk patterns and laws, improve the accuracy and reliability of prediction, and provide a scientific basis and decision-making support for disaster prevention and mitigation.
[0081] In the embodiment of the present invention, the data preprocessing operation in A1 is:
[0082] a11. Clean the data to remove missing values and outliers in the data, and compensate for missing values and outliers by filling or deleting them;
[0083] a12. Convert the categorical variables in the retained data into numerical features that can be used for subsequent model training through label encoding;
[0084] a13. Standardize the numerical features so that the processed numerical features conform to the standard normal distribution;
[0085] a14. Optimize the model performance by selecting numerical features with high contribution to model training through feature selection.
[0086] In the embodiment of the present invention, the formula for standardizing numerical features in a13 is:
[0087] ;
[0088] where X norm represents the standardized feature, X b represents the original feature, and α and λ respectively represent the mean and standard deviation values of the feature to determine the reduction of the difference of the numerical feature.
[0089] In the embodiment of the present invention, the operation of constructing a risk knowledge graph in a21 is:
[0090] B1. Identify each entity in the data and assign a unique identifier to define it as a node in the graph;
[0091] B2. Define the edges between nodes to implement the association operation between nodes, and form the basic structure of the risk knowledge graph through a construction formula;
[0092] B3. Through the knowledge fusion step, extract external influence data and integrate it into the knowledge graph to improve the risk knowledge graph structure;
[0093] B4. Add the preprocessed new data to the risk knowledge graph, and update the attributes of the nodes and the connection information of the edges. Delete the outdated data or the no-longer relevant relationships to implement the real-time update operation of the graph.
[0094] In the embodiment of the present invention, the construction formula in B2 is:
[0095] G = (V, E);
[0096] G represents the risk knowledge graph, V represents the set of nodes, and E represents the set of edges;
[0097] For the association operation between nodes, it is necessary to determine the relationship similarity between nodes to determine the edge weight and the association direction, and the specific determination operation is:
[0098] b21. Arbitrarily select two nodes. First, determine whether the two nodes are of the same type of regional location feature. If they are of the same type of regional location feature, it is necessary to determine whether the regional locations are adjacent, and then introduce geographical information to determine the risk flow direction to point to the flow node;
[0099] b22. If they are of different types of features and one node is a regional location feature, the association direction is transmitted from the current node to the other node. Summarize the nodes to be constructed related to the current node and label them as v n {v1 , v 2 , …, v n} Determine the contribution influence degree according to the economic loss value caused by each single node feature in the historical data, determine the influence order of each node feature according to the magnitude of the economic loss value, and the one with the largest value has the greatest influence. Then determine the edge weight according to the sorted node features;
[0100] b23. Thus, a complete risk knowledge graph architecture is constructed.
[0101] In the embodiment of the present invention, the calculation formula of the edge weight in b22 is:
[0102] W n(n-1) = sim(v n , v n-1 );
[0103] W n(n-1) represents the edge weight between node v n and node v n-1 , v n represents the nth node, and sim(v n , v n-1 ) represents the similarity between nodes;
[0104] And sim(v n , v n-1 ) = [(k 1 + k 2 + … + k n ) / n] / k n ;
[0105] k n represents the loss value of the nth place after sorting according to the magnitude of the economic loss value, and (k 1 + k 2 + … + k n ) / n is the average economic loss value of all influencing node features.
[0106] By defining the edges between nodes, the association operation between nodes is realized, and the basic architecture of the risk knowledge graph is formed by constructing a formula. Determine the contribution influence degree according to the economic loss value caused by each single node feature in the historical data, determine the influence order of each node feature according to the magnitude of the economic loss value, and the one with the largest value has the greatest influence. Then determine the edge weight according to the sorted node features, so that the knowledge graph can be quickly established and the relationship between the nodes in the knowledge graph can be associated. And determine the edge weight value according to the judgment training of the historical data, realize the continuous influx of new data, and dynamically update the graph structure through the graph update mechanism to ensure its timeliness and accuracy.
[0107] In the embodiment of the present invention, the operation of learning and aggregating node features in a22 is as follows:
[0108] C1. Use the node feature matrix and the adjacency matrix as the input layer of the risk level prediction model to ensure that the graph neural network can capture the relationships between nodes;
[0109] C2. Set the GCN layer to perform layer-by-layer aggregation and update of node features. Each layer of GCN uses the adjacency matrix and the degree matrix to aggregate the information of neighbor nodes;
[0110] C3. The graph neural network learns the local features of nodes, and maps the updated node feature matrix to a specific risk level in the output layer.
[0111] In the embodiment of the present invention, the expression for updating node features in C2 is:
[0112] H (l+1) =σ(D 1 / 2 AD 1 / 2 H (l) W (l) );
[0113] And H (l+1) represents the node feature matrix of the (l + 1)-th layer, H (l) represents the node feature of the adjacent l-th layer, matrix A represents the adjacency matrix, D represents the degree matrix, W (l) represents the weight matrix of the l-th layer, and σ represents the activation function;
[0114] And the node feature update formula is:
[0115] ;
[0116] where, h i represents the updated feature of node i, N(i) represents the set of neighbor nodes of node i, d i and d j represent the degrees of node i and node j respectively;
[0117] And the mapping formula for the risk level in C3 is:
[0118] Y = softmax(H (l) W (l+1) );
[0119] And W (l+1) is the weight matrix of the output layer, Y is the risk level prediction result of the node, each row represents the prediction probability distribution of a node, and softmax is a combination of a fully connected layer and a classifier to implement the mapping operation of the risk level.
[0120] By taking the node feature matrix and the adjacency matrix as the input layer of the risk level prediction model, a GCN layer is set to aggregate and update the node features layer by layer. Each layer of GCN uses the adjacency matrix and the degree matrix to aggregate the information of neighbor nodes. The graph neural network learns the local features of the nodes. In the output layer, the updated node feature matrix is mapped to specific risk levels, so that during the training process, it gradually adapts to the complexity of the data after multiple adjustments, the step-by-step convergence of the model parameter adjustment, and finally tends to be stable, achieving the optimal state of the model and better performing the risk level prediction operation.
[0121] In the present invention, two layers of GCN are used, with 64 hidden units in each layer, the activation function is ReLU, the Adam optimizer is used for model training, the learning rate is 0.01, and the number of training epochs is 500. During the training process, an early stopping strategy is adopted to prevent overfitting. The training environment is Python 3.8, and the PyTorch framework is used for implementation. It can be seen that as the number of training epochs increases, the model gradually converges. The loss curve during the specific experimental training process is as Figure 3 shown. In the later stage of model training (after about 1500 epochs), the loss curve begins to show slight oscillations, and as the number of training epochs increases, the amplitude of the oscillations gradually decreases, while the frequency of the oscillations gradually increases. Finally, the loss value tends to converge. This phenomenon indicates that the model gradually adapts to the complexity of the data after multiple adjustments during the training process. The gradual weakening of these oscillations indicates the stability of the model when finely learning complex features. The decrease in the amplitude of the oscillations reflects the step-by-step convergence of the model parameter adjustment and finally tends to be stable, indicating that the model has reached the optimal state and can better carry out risk level prediction.
[0122] In the embodiment of the present invention, the specific operation of the risk level prediction result in a23 is as follows:
[0123] D1. Output the risk prediction result based on the graph neural network and perform backpropagation;
[0124] D2. Use the cross-entropy loss function to calculate the error between the prediction result and the true label, and implement the compensation operation between the prediction result and the true label. The specific loss function formula is:
[0125] ;
[0126] N represents the number of samples, C represents the number of categories, y gc represents the true label of sample g, represents the predicted probability of sample g;
[0127] Then, update the model parameters through the gradient descent algorithm to minimize the loss function, and its formula is:
[0128] θnew = θ old − η(▽ θ )L;
[0129] θ new is the new model parameter, θ old is the old model parameter, η represents the learning rate, (▽ θ )L represents the gradient of the loss function with respect to the parameter;
[0130] D3. Use accuracy, precision, recall, and F1 score to evaluate the performance of the risk level prediction model.
[0131] In the embodiment of the present invention, the operation of predicting the risk level of the parameter data of the area to be measured in A3 to the risk level prediction model is as follows:
[0132] a31. After collecting the parameter data of the area to be measured, perform cleaning processing through the risk level prediction model, and then establish the current risk level prediction model of the area to be measured based on the combination of the knowledge graph and the graph neural network;
[0133] a32. Analyze and output the risk result of the area to be measured according to the risk level prediction model, and adjust the prevention strategy according to the predicted risk result.
[0134] Embodiment 2. The difference compared with Embodiment 1 is that the present invention also collects a region with a subtropical monsoon climate for detection. The current region has dense rivers and streams, and the average annual precipitation is 1638 mm. The unique climate characteristics and topographical and geomorphic conditions have led to frequent, severe, and repeated flash floods. According to the statistics of flash flood disaster investigation and evaluation data, the area of the current flash flood disaster prevention area reaches, accounting for 78% of the total area of the province. There are 9758 flash flood disaster risk areas, and the population in the risk areas reaches more than 800,000. The data set in this article comes from a wide range of sources, including flash flood disaster investigation and evaluation data, news reports, statistical yearbooks, etc. Among them, there are 9758 flash flood disaster investigation and evaluation data in total, and 9000 are obtained after removing noise; there are 1756 news reports with the keyword of flash flood disaster as the title in the past 15 years, and 857 are obtained after removing noise. The experimental data set contains 9000 records, with features such as geographical location and risk type, and is divided into a training set of about 7200 records, a validation set of 900 records, and a test set of 900 records according to a ratio of about 8:1:1, which are used for model training, model selection, hyperparameter tuning, and evaluation of the model generalization performance.
[0135] In the operation of specific model evaluation, in addition to using traditional classification metrics such as accuracy, precision, recall, and F1-score, the area under the receiver operating characteristic curve (AUC) is also used to further measure the performance of the model. AUC is an important indicator for measuring the ability of a classification model to rank positive and negative samples. Its value ranges from 0.5 to 1, where 1 indicates that the model has perfect discrimination ability, while 0.5 indicates that the model has no discrimination ability, equivalent to random guessing. Specifically, the closer the AUC is to 1, the stronger the classification ability of the model, and the change curve during the model training process is as Figure 4 shown.
[0136] In the present invention, 5-fold cross-validation is performed on the model to more robustly evaluate its performance. The AUC values for each fold are shown in Table 1, gradually increasing from 0.79 in the first fold to 0.84 in the fifth fold. At the same time, other performance indicators such as accuracy, precision, recall, and F1-score also show corresponding improvements. These results indicate that with the training of the model, the overall performance of the model has been significantly improved, verifying the effectiveness of the proposed method.
[0137] The model is evaluated using the 5-fold cross-validation method. The dataset is randomly divided into 5 subsets. Each time, 4 subsets are used for training and 1 subset is used for validation. This is repeated 5 times, and the final result is the average of the 5 experiments. The model evaluation results are shown in Table 1:
[0138] Table 1 Model Evaluation Results
[0139] AUC Accuracy Precision Recall F1 - Score Fold 1 0.79 0.85 0.83 0.84 0.83 Fold 2 0.81 0.86 0.84 0.85 0.84 Fold 3 0.82 0.87 0.85 0.86 0.85 Fold 4 0.83 0.88 0.86 0.87 0.86 Fold 5 0.84 0.89 0.87 0.88 0.87
[0140] In summary, cross-validation effectively evaluates the stability and generalization ability of the model.
[0141] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0142] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0143] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting flash flood disaster risk level based on knowledge graph and graph neural network, characterized by: The specific steps include: A1. Collect data related to flash flood disaster risks and perform data preprocessing operations; A2. Build and train a risk level prediction model combining knowledge graph and graph neural network. The specific operations are as follows: a21. Extract entities and relationships from the data, and construct a risk knowledge graph using node definition, edge definition, knowledge fusion, and graph update operations; a22. Based on the constructed knowledge graph, the graph neural network is used to learn and aggregate node features, and feature information is extracted for risk level prediction; a23. Utilize forward propagation, loss calculation, parameter update and result evaluation to obtain the predicted results of risk level and implement multiple trainings; A3. Introduce parameter data of the area to be tested into the risk level prediction model to predict the risk level; The operation of learning and aggregating node features using the graph neural network in a22 is: C1. Use the node feature matrix and adjacency matrix as the input layer of the risk level prediction model to ensure that the graph neural network can capture the relationship between nodes; C2. Set the GCN layer to aggregate and update the node features layer by layer. Each layer of GCN uses the adjacency matrix and degree matrix to aggregate the information of neighboring nodes. C3. The graph neural network learns the local features of the nodes and maps the updated node feature matrix to specific risk levels in the output layer.
2. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 1 is characterized in that: The data preprocessing operation in A1 is: a11. Clean the data to remove missing values and outliers in the data, and compensate for missing values and outliers by filling or deleting them; a12. Convert the categorical variables in the retained data into numerical features that can be used for subsequent model training through label encoding; a13. Standardize the numerical features so that the processed numerical features conform to the standard normal distribution; a14. Through feature selection, numerical features that contribute most to model training are selected to optimize model performance.
3. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 2 is characterized in that: The formula for normalizing the numerical features in a13 is: ; Among them, X norm represents the standardized features, X b represents the original feature, α and λ represent the mean and standard deviation values of the feature, respectively, to determine the reduction in variance of the numerical feature.
4. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 1 is characterized in that: The operation of constructing the risk knowledge graph in a21 is: B1. Identify each entity in the data and assign a unique identifier to define it as a node in the graph; B2. Define the edges between nodes, implement the association operations between nodes, and form the infrastructure of the risk knowledge graph by constructing formulas; B3. Through the knowledge fusion step, extract external impact data and integrate it into the knowledge graph to improve the risk knowledge graph architecture; B4. Add the preprocessed new data to the risk knowledge graph, update the attributes of the nodes and the connection information of the edges, delete outdated data or no longer relevant relationships, and realize real-time update of the graph.
5. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 4 is characterized in that: The formula constructed in B2 is: G = (V, E); G represents the risk knowledge graph, V represents the node set, and E represents the edge set; For the association operation between nodes, it is necessary to determine the relationship similarity between the nodes to determine the edge weight and the association direction, and the specific determination operation is: b21. Select any two nodes and first determine whether the two nodes are of the same type of regional location features. If they are of the same type, it is necessary to determine whether the regional locations are adjacent, and then introduce geographic information to determine the risk flow direction to point to the flow node; b22. If the features are not of the same type and one node is a regional location feature, the association direction is transmitted from the current node to another node, and the nodes that need to be constructed related to the current node are summarized and marked as v n {v1, v2, …, v n }, determine the contribution influence according to the economic loss value caused by each single node feature in the historical data, determine the influence order of each node feature according to the size of the economic loss value, and the one with the largest value has the greatest influence, and then determine the edge weight according to the sorted node features; b23. In this way, a complete risk knowledge graph architecture is constructed.
6. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 5 is characterized in that: The calculation formula of the edge weight in b22 is: W n(n-1) =sim(v n ,v n-1 ); W n(n-1) Represents node v n and node v n-1 The edge weight between n Represented as the nth node, sim(v n , v n-1 ) represents the similarity between nodes; And sim (v n , v n-1 ) = [(j1+j2+…+j n ) / n] / j n ; j n It is expressed as the nth loss value after sorting the economic loss value, and (j1+j2+…+j n ) / n is the average value of economic losses of all the influencing node characteristics.
7. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 1 is characterized in that: The expression for updating node features in C2 is: H (l+1) =σ(D 1 / 2 AD 1 / 2 H (l) W (l) ); And H (l+1) is the node feature matrix representing the (l+1)th layer, H (l) To represent the adjacent node features of the lth layer, matrix A represents the adjacency matrix, D represents the degree matrix, and W (l) represents the weight matrix of the lth layer, σ represents the activation function; The node feature update formula is: ; Among them, h i represents the updated features of node i, N(i) represents the set of neighbor nodes of node i, d i and d j Represent the degrees of nodes i and j respectively; The mapping formula for risk level in C3 is: Y=softmax(H (l) W (l+1) )) And W (l+1) is the weight matrix of the output layer, Y is the risk level prediction result of the node, each row represents the predicted probability distribution of a node, and softmax is a combination of a fully connected layer and a classifier to realize the mapping operation of risk level.
8. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 1 is characterized in that: The specific operation of the risk level prediction result in a23 is: D1. Output risk prediction results based on graph neural network and perform back propagation; D2. Use the cross entropy loss function to calculate the error between the predicted result and the true label to achieve the compensation operation between the predicted result and the true label. The specific loss function formula is: ; N represents the number of samples, C represents the number of categories, and y gc represents the true label of sample g, represents the predicted probability of sample g; Then, the model parameters are updated through the gradient descent algorithm to minimize the loss function, which is formulated as: i new =θ old -η(▽ θ )L; θ new is the new model parameter, θ old is the old model parameter, η represents the learning rate, (▽ θ )L represents the gradient of the loss function with respect to the parameters; D3. Use accuracy, precision, recall and F1 score to evaluate the performance of the risk level prediction model.
9. The method for predicting flash flood disaster risk level based on knowledge graph and graph neural network according to claim 1 is characterized in that: The operation of using the parameter data of the area to be tested in A3 to predict the risk level in the risk level prediction model is as follows: a31. After collecting parameter data of the area to be tested, the data is cleaned and processed by the risk level prediction model, and then a current risk level prediction model for the area to be tested is established based on the knowledge graph and graph neural network; a32. Analyze and output the risk results of the area to be tested based on the risk level prediction model, and adjust the prevention strategy based on the predicted risk results.
Citation Information
Patent Citations
Mountain torrent disaster risk division and prediction method based on GIS (geographic information system)-neural network integration
CN108280553A
Landslide space prediction method based on knowledge graph and representation learning
CN119398148A
Power grid natural disaster early warning method based on knowledge graph and computer equipment
CN119398225A