A geological disaster knowledge graph construction method, system and storage medium

By combining expert knowledge and machine knowledge correction, and utilizing relational dependency mining and multi-linked list optimization, the problem of information reliability in the construction of geological disaster knowledge graphs was solved, and the accurate construction and optimization of geological disaster knowledge graphs were achieved.

CN119721207BActive Publication Date: 2025-10-24山东省煤田地质局第四勘探队 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411783681.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-10-24
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

In existing technologies, the reliability of information cannot be guaranteed during the construction of geological disaster knowledge graphs, leading to errors and making it impossible to construct accurate geological disaster knowledge graphs.

Method used

Multimodal information is corrected using expert knowledge correction and machine knowledge correction methods. A geological hazard knowledge graph is constructed through a relation dependency mining layer and a multi-condition constrained random forest model. The information content is then optimized using a geological hazard multi-linked list.

Benefits of technology

It improves the accuracy and usability of the geological hazard knowledge graph, and enhances the reliability of geological hazard information retrieval, prediction and assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721207B_ABST
    Figure CN119721207B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of geological disaster knowledge graph construction, in particular to a geological disaster knowledge graph construction method, system and storage medium. First, a kind of information correction method is used to correct multi-modal information;The method is matched by establishing expert knowledge base to correct information;And by statistical verification to data information, the information with problems is corrected;Combining expert knowledge correction and machine knowledge correction realizes the accurate correction of original information;Second, a geological disaster knowledge graph generation model is proposed;The model constructs geological disaster knowledge graph by analyzing the dependency relationship between corrected information and disaster type;Finally, a geological disaster knowledge graph optimization method is proposed;The method uses corrected auxiliary information to perform correlation calculation to obtain auxiliary elements to optimize the knowledge graph. The above methods comprehensively realize the accurate construction of geological disaster knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of geological disaster knowledge graph construction, in particular to a geological disaster knowledge graph construction method, system and storage medium. BACKGROUND

[0002] Geological disasters refer to geological processes or phenomena that cause environmental damage under the action of natural or human factors, including landslides, debris flows, earthquakes and collapses, etc. These disasters have a very serious impact on humans and the environment. Therefore, a deep understanding of the occurrence mechanism, influencing factors and their mutual relationships of geological disasters is crucial for disaster prevention, risk management and emergency response.

[0003] With the continuous development of computer technology, geologists have gradually applied knowledge graph technology to the field of geological disasters. Knowledge graph is a technology that organizes and represents knowledge in a graphical way, and establishes the relevance of knowledge through nodes (entities) and edges (relationships). The advantage of knowledge graph is that it can integrate a large amount of heterogeneous data into a unified platform, thereby forming a knowledge network from a global perspective.

[0004] In the field of geological disasters, the construction of knowledge graph can effectively integrate various information related to geological disasters, including: disaster types, such as the definition, characteristics and occurrence conditions of various geological disasters such as landslides, debris flows and earthquakes; cause analysis, how natural factors (such as rainfall, earthquakes) and human factors (such as land use change, construction activities) lead to geological disasters; influencing factors, the influence of different geological environments, climate conditions and human activities on disaster occurrence; for the above mentioned content, professional personnel collects, organizes and analyzes data in a systematic way according to different information sources to construct a comprehensive geological disaster knowledge graph; however, the reliability of the information cannot be guaranteed in this process; the existing technology lacks an effective information correction mechanism, which will lead to errors in the construction of geological disaster knowledge graph, and an accurate geological disaster knowledge graph cannot be obtained.

[0005] Therefore, the present application proposes a geological disaster knowledge graph construction method, system and storage medium. SUMMARY

[0006] The purpose of the present application is to provide a geological disaster knowledge graph construction method, system and storage medium for constructing a geological disaster knowledge graph, which can be used for geological disaster information query, prediction and disaster assessment functions; The specific functions include: first, different channels are used to obtain multi-modal information related to geological disasters, and the multi-modal information is corrected; For information correction, the present application proposes a method of combining expert knowledge correction and machine knowledge correction; In the expert knowledge correction process, an expert knowledge base is established to generate expert knowledge matching text for different types of geological disasters, which is used to generate a candidate expert correction information set for the collected information; Machine knowledge correction is obtained by classifying and counting the collected multi-modal information, and the description information of different types of geological disasters is obtained, and the information co-occurrence probability is calculated; The information with high co-occurrence rate is cross-validated and corrected accordingly to obtain a selected machine correction information set; The candidate expert correction information set and the selected machine correction information set are combined to generate a comprehensive correction information set to correct the original information data; Secondly, the present application proposes a geological disaster knowledge graph generation model; The model takes the corrected disaster description information and the initial geological disaster graph as input data; The relationship dependency mining layer obtains the dependency relationship between information features, and a multi-condition constraint random forest model is used as the dependency relationship mining model in the relationship dependency mining layer; The geological disaster knowledge graph generation layer constructs the geological disaster knowledge graph according to the structured geological disaster dependency matrix and the geological disaster graph; Finally, the present application proposes an optimization method for the geological disaster knowledge graph; The method converts the geological disaster knowledge graph into a geological disaster multi-linked list, and explains the internal relationship of the geological disaster knowledge graph in the form of linked list; By means of auxiliary information elements in the corrected disaster auxiliary description information, the content of the geological disaster multi-linked list is increased and the data is updated; And through the reverse rule, the optimized geological disaster knowledge graph is obtained.

[0007] To achieve the above purpose, the present application provides the following technical scheme:

[0008] In the first aspect, the present application provides a geological disaster knowledge graph construction method, comprising:

[0009] Collecting multi-modal information related to geological disasters and performing deduplication processing to obtain deduplicated multi-modal information; wherein the multi-modal information includes: text information, video information and audio information;

[0010] Further, the deduplicated multi-modal information is classified according to the disaster type to obtain disaster description information and disaster auxiliary description information;

[0011] Wherein, the classification process of the deduplicated multi-modal information includes: obtaining description information corresponding to different types of geological disasters;

[0012] Further, keyword extraction is performed on the description information to obtain a disaster type keyword;

[0013] Further, the disaster type keyword is vectorized to obtain a disaster type matching word vector;

[0014] Further, the deduplicated multi-modal information is divided into a text data set, a video data set, and an audio data set according to information types;

[0015] Further, all text descriptions are extracted from the text data set to obtain text description information, all video descriptions are extracted from the video data set to obtain video description information, and all audio descriptions are extracted from the audio data set to obtain audio description information;

[0016] Further, feature extraction is performed on the text description information, the video description information, and the audio data respectively to obtain text description features, video description features, and audio description features;

[0017] Further, the similarity of the text description features, the video features, and the audio description features to the disaster type matching word vector is calculated respectively to obtain a text similarity, a video similarity, and an audio similarity;

[0018] Further, when the text similarity, the video similarity, and the audio similarity exceed a similarity matching threshold, the deduplicated multi-modal information is taken as the disaster explanation information of the corresponding disaster type; when the text similarity, the video similarity, and the audio similarity do not exceed the similarity matching threshold, the deduplicated multi-modal information is taken as the disaster auxiliary explanation information.

[0019] Further, the disaster explanation information and the disaster auxiliary explanation information are corrected to obtain corrected disaster explanation information and corrected disaster auxiliary explanation information;

[0020] The correction of the disaster explanation information and the disaster auxiliary explanation information includes an expert knowledge correction mechanism and a machine knowledge correction mechanism.

[0021] The expert knowledge correction mechanism corrects information through geologist knowledge, including: obtaining related knowledge of N geologists on geological disasters to construct an expert knowledge base.

[0022] Further, expert knowledge matching text is generated from the information in the expert knowledge base;

[0023] Further, a multi-layer knowledge matching index is established for the expert knowledge matching text;

[0024] Further, the disaster description information and the disaster auxiliary description information are matched with the expert knowledge matching text by using the multi-layer knowledge matching index.

[0025] Further, information consistency is calculated according to the matching result to obtain an information consistency value.

[0026] Further, when the information consistency value does not exceed a consistency threshold, an information correction algorithm is used to obtain a candidate expert correction information set.

[0027] The machine knowledge correction mechanism corrects information by using a machine algorithm, including: using a named entity recognition algorithm to identify entity fields of the disaster description information and the disaster auxiliary description information.

[0028] Further, data of the entity fields is counted, and a co-occurrence probability of the entity fields is calculated to obtain an entity field co-occurrence matrix.

[0029] Further, the disaster description information and the disaster auxiliary description information are filtered according to the entity field co-occurrence matrix to obtain a geological disaster filtered information set.

[0030] Further, the geological disaster filtered information set is cross-validated.

[0031] Further, when a cross-validation value of any information in the geological disaster filtered information set is lower than a cross-validation threshold, information that does not meet a condition is corrected according to similar information to obtain a candidate machine correction information set.

[0032] Further, the candidate expert correction information set and the candidate machine correction information set are merged to obtain a comprehensive correction information set; when the candidate expert correction information set and the candidate machine correction information set have the same correction item, if correction information of the correction item is the same, the correction information is merged, and if the correction information of the correction item is different, the correction information is regularized and then merged; the comprehensive correction information set is used to correct the disaster description information and the disaster auxiliary description information.

[0033] Further, an initial geological disaster graph is constructed, and the initial geological disaster graph is represented as: {Node(GHT, GHAT), Edg(NULL)}; Node() represents an initial node; GHT represents the disaster type; GHAT represents a geological disaster area type; and Edg() represents an initial edge with an initial value of NULL, i.e., an empty value.

[0034] Further, the corrected disaster description information and the initial geological disaster graph are input into a geological disaster knowledge graph generation model; the geological disaster knowledge graph generation model includes:

[0035] The receiving layer comprises: a disaster information receiving module, configured to receive the correction disaster description information; and a graph receiving module, configured to receive the initial geological disaster graph.

[0036] The processing layer is configured to process the correction disaster description information to obtain structured geological disaster description features.

[0037] The relationship dependence mining layer is configured to mine dependence relationships of the structured geological disaster description features, and comprises: establishing a structured geological disaster description feature matrix according to the structured geological disaster description features; wherein the structured geological disaster description feature matrix is expressed as: wherein GHTC i represents a disaster type feature of the ith structured geological disaster description feature; GHATC i represents a geological disaster area type feature of the ith structured geological disaster description feature; and GHTC represents the jth related information feature vector of the ith structured geological disaster description feature; M represents a total number of the structured geological disaster description features; and N represents a total number of related information feature vectors in the structured geological disaster description features.

[0038] The structured geological disaster description feature matrix is input into a multi-condition constraint random forest model to obtain a structured geological disaster dependence matrix.

[0039] The multi-condition constraint random forest model comprises: an input layer, configured to receive the structured geological disaster description feature matrix; a tree generation layer, configured to randomly generate a plurality of trees; wherein each tree independently selects feature data corresponding to a different data disaster type; a condition constraint setting layer, configured to set a limitation condition for each tree to constrain branches of the tree; wherein a constraint rule of the condition constraint setting layer is set as: an occurrence reliability constraint between the disaster type and the geological disaster area type; an association constraint between the disaster type and a disaster occurrence influencing factor; an influence constraint between the geological disaster area type and a disaster influence; a voting layer, configured to aggregate analysis results of the trees; a dependence relationship calculation layer, configured to calculate a dependence relationship according to the analysis results of the voting layer; and an output layer, configured to output the structured geological disaster dependence matrix.

[0040] The geological disaster knowledge graph generation layer is configured to generate a geological disaster knowledge graph.

[0041] The visual output layer is configured to visually output the geological disaster knowledge graph.

[0042] Further, the geological disaster knowledge graph is optimized by using the correction disaster auxiliary description information to obtain an optimized geological disaster knowledge graph.

[0043] wherein, the optimization process of the geological disaster knowledge graph comprises:

[0044] converting the geological disaster knowledge graph into a geological disaster multi-linked list; the geological disaster multi-linked list comprises: linked list nodes, linked list multi-edges and linked list connection directions; wherein, the linked list nodes are nodes in the geological disaster knowledge graph; the linked list multi-edges are connection relationships between nodes in the geological disaster knowledge graph; the linked list multi-edges record a plurality of information elements related to geological disasters; and the linked list connection directions are connection strengths between nodes in the geological disaster knowledge graph, and the connection directions of the linked list multi-edges are set according to the connection strengths;

[0045] Further, the relevance of the correction disaster auxiliary explanation information and the data in the geological disaster multi-linked list is calculated;

[0046] Further, the auxiliary information elements of the correction disaster auxiliary explanation information are extracted according to the relevance calculation results;

[0047] Further, the auxiliary information elements are integrated with the linked list nodes and the linked list multi-edges of the geological disaster multi-linked list, and the connection strengths are updated, and the connection directions are modified;

[0048] Further, the updated geological disaster multi-linked list is reversely converted into the optimized geological disaster knowledge graph.

[0049] In a second aspect, the present application provides a geological disaster knowledge graph construction system, which comprises:

[0050] an information collection unit for collecting information related to geological disasters from different channels;

[0051] an information classification unit for classifying geological disaster information;

[0052] an information correction unit for correcting the collected related information; wherein, the information correction unit comprises: an expert knowledge correction module and a machine knowledge correction module;

[0053] the expert knowledge correction module comprises:

[0054] obtaining the related knowledge of N geologists on geological disasters to construct an expert knowledge base;

[0055] Further, the information in the expert knowledge base is generated into an expert knowledge matching text;

[0056] Further, a multi-layer knowledge matching index is established for the expert knowledge matching text;

[0057] Further, the disaster description information and the disaster auxiliary description information are matched with the expert knowledge matching text according to the multi-layer knowledge matching index.

[0058] Further, information consistency is calculated according to the matching result to obtain an information consistency value; when the information consistency value does not exceed a consistency threshold value, an information correction algorithm is used to obtain a candidate expert correction information set.

[0059] The machine knowledge correction module comprises:

[0060] The named entity recognition algorithm is used to identify entity fields of the disaster description information and the disaster auxiliary description information.

[0061] Further, data statistics are performed on the entity fields, and co-occurrence probability of the entity fields is calculated to obtain an entity field co-occurrence matrix.

[0062] Further, the disaster description information and the disaster auxiliary description information are screened according to the entity field co-occurrence matrix to obtain a geological disaster screening information set; the geological disaster screening information set is cross-validated; when cross-validation values of any information in the geological disaster screening information set are lower than a cross-validation threshold value, information that does not meet a condition is corrected according to similar information to obtain a candidate machine correction information set.

[0063] Further, the candidate expert correction information set and the candidate machine correction information set are merged to obtain a comprehensive correction information set; when the candidate expert correction information set and the candidate machine correction information set have the same correction item; if correction information of the correction item is the same, the correction information is merged, and if the correction information of the correction item is different, the correction information is merged after regularization; the comprehensive correction information set is used to correct the disaster description information and the disaster auxiliary description information.

[0064] A knowledge graph construction unit is configured to construct a geological disaster knowledge graph; the knowledge graph construction unit comprises:

[0065] An initial geological disaster graph is constructed.

[0066] The corrected disaster description information and the initial geological disaster graph are input into a geological disaster knowledge graph generation model; the geological disaster knowledge graph generation model comprises:

[0067] The receiving layer comprises a disaster information receiving module configured to receive the corrected disaster description information, and a graph receiving module configured to receive the initial geological disaster graph.

[0068] The processing layer is configured to process the corrected disaster description information to obtain structured geological disaster description features.

[0069] The relationship-dependent mining layer is configured to mine the dependency relationship of the structured geological disaster description feature, and includes: establishing a structured geological disaster description feature matrix according to the structured geological disaster description feature; inputting the structured geological disaster description feature matrix into a multi-condition constraint random forest model to obtain a structured geological disaster dependency matrix;

[0070] The geological disaster knowledge graph generation layer is configured to generate a geological disaster knowledge graph.

[0071] The visual output layer is configured to visually output the geological disaster knowledge graph.

[0072] The knowledge graph optimization unit is configured to optimize the geological disaster knowledge graph, and includes:

[0073] The geological disaster knowledge graph is converted into a geological disaster multi-linked list, which includes: a linked list node, a linked list multi-edge, and a linked list connection direction; the linked list node is a node in the geological disaster knowledge graph; the linked list multi-edge is a connection relationship between nodes in the geological disaster knowledge graph; the linked list multi-edge records a plurality of information elements related to geological disasters; and the linked list connection direction is a connection strength between nodes in the geological disaster knowledge graph, and the connection direction of the linked list multi-edge is set according to the connection strength.

[0074] Further, the correlation between the correction disaster auxiliary description information and the geological disaster multi-linked list is calculated.

[0075] Further, the auxiliary information elements of the correction disaster auxiliary description information are extracted according to the correlation calculation result.

[0076] Further, the auxiliary information elements are integrated with the linked list node and the linked list multi-edge of the geological disaster multi-linked list, and the connection strength is updated, and the connection direction is modified.

[0077] Further, the updated geological disaster multi-linked list is reversely converted into an optimized geological disaster knowledge graph.

[0078] The visualization unit is configured to visually display the generated geological disaster knowledge graph.

[0079] The maintenance unit is configured to periodically update the geological disaster knowledge graph.

[0080] In a third aspect, the present application provides a geological disaster knowledge graph construction storage medium, wherein the storage medium stores a geological disaster knowledge graph construction program; and the geological disaster knowledge graph construction program, when executed by a processor, implements the geological disaster knowledge graph construction of any one of the first aspect or the second aspect.

[0081] Compared with the prior art, the present application has the following beneficial effects:

[0082] 1. The present application provides a geological disaster information correction method; the method comprehensively utilizes expert knowledge correction and machine knowledge correction; wherein, in the expert knowledge correction process, an expert knowledge base is established to generate expert knowledge matching text for different types of geological disasters, which is used to generate a candidate expert correction information set for the collected information; the machine knowledge correction obtains description information of different types of geological disasters by classifying and counting the collected multi-modal information, and calculates the information co-occurrence probability; the information with high co-occurrence rate is cross-validated for corresponding correction to obtain a selected machine correction information set; and the candidate expert correction information set and the selected machine correction information set are comprehensively combined to generate a comprehensive correction information set to correct the original information data; the accurate construction of the geological disaster knowledge graph is realized, thereby improving the usability of the geological disaster knowledge graph.

[0083] 2. The present application provides a geological disaster knowledge graph generation model; the model takes the corrected disaster description information and the initial geological disaster graph as input data; the dependency relationship between information features is obtained through a relationship dependency mining layer, and a multi-condition constraint random forest model is used as a dependency relationship mining model in the relationship dependency mining layer; a structured geological disaster dependency matrix and a geological disaster graph are used by a geological disaster knowledge graph generation layer to construct a geological disaster knowledge graph; this is conducive to the accurate construction of the geological disaster knowledge graph.

[0084] 3. The present application provides an optimization method for a geological disaster knowledge graph; the method converts the geological disaster knowledge graph into a geological disaster multi-linked list, and explains the internal relationship of the geological disaster knowledge graph in the form of a linked list; the addition of content and the update of data in the geological disaster multi-linked list are realized by means of auxiliary information elements in the corrected disaster auxiliary description information; and the optimized geological disaster knowledge graph is obtained through reverse rules; the method optimizes the accuracy of the information of the geological disaster knowledge graph and enriches the information content of the geological disaster knowledge graph. BRIEF DESCRIPTION OF DRAWINGS

[0085] Figure 1 A flowchart of a geological disaster knowledge graph construction method provided by an embodiment of the present application;

[0086] Figure 2 A structure diagram of a geological disaster knowledge graph construction system provided by an embodiment of the present application;

[0087] Figure 3 A structural diagram of the information correction unit provided for the embodiment of the present application is shown in FIG. 1;

[0088] Figure 4 A structural diagram of the multi-layer knowledge matching index provided for the embodiment of the present application is shown in FIG. 2;

[0089] Figure 5 A schematic diagram of the initial geological disaster map provided for the embodiment of the present application is shown in FIG. 3;

[0090] Figure 6 A structural diagram of the geological disaster knowledge graph generation model provided for the embodiment of the present application is shown in FIG. 4;

[0091] Figure 7 A schematic diagram of the geological disaster knowledge graph provided for the embodiment of the present application is shown in FIG. 5;

[0092] Figure 8 A schematic diagram of the geological disaster multi-linked list provided for the embodiment of the present application is shown in FIG. 6;

[0093] Figure 9 A schematic diagram of the optimized geological disaster knowledge graph provided for the embodiment of the present application is shown in FIG. 7. DETAILED DESCRIPTION

[0094] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without any creative work fall within the protection scope of the present application.

[0095] Geological disasters are natural disasters caused by geological actions or changes in geological environment, which cause damage to human beings and natural environment. They include various types, such as earthquakes, volcanic eruptions, landslides, debris flows, collapses, and ground subsidence. Geologists are committed to revealing the causes, processes, and potential impacts of these disasters in the process of studying these disasters.

[0096] In the research of geological disasters, due to the complexity of the background factors and the diversity of geological disasters, geologists cannot comprehensively understand and analyze all types of geological disasters by relying on a single method or model alone, making the research in the field of geological disasters very challenging. In recent years, with the continuous development of computer technology, geologists have applied computer technology to the field of geological disasters, and through research, it has been found that combining knowledge graph technology with geological disasters can deeply mine the occurrence mechanism, influencing factors and their mutual relationship of geological disasters; the knowledge graph is a technology for organizing and representing knowledge in a graphical way, and the correlation of knowledge is established through nodes (entities) and edges (relationships). The advantage of the knowledge graph is that it can integrate a large amount of heterogeneous data into a unified platform, thereby forming a knowledge network from a global perspective.

[0097] In the prior art, researchers achieve multiple tasks such as prediction, evaluation and management of geological disasters by establishing a geological disaster knowledge graph; whether these can achieve good results mainly depends on the construction of the geological disaster knowledge graph; the construction of the geological disaster knowledge graph collects a large amount of data, including relevant literature, research reports and related data, to form a huge knowledge base, and then analyzes the relationship of knowledge from the knowledge base to construct a knowledge graph; however, in this process, the reliability of the data cannot be verified, so there is a certain risk in the construction of the geological disaster knowledge graph; in order to solve the problem of errors in the data in the construction process of the geological disaster knowledge graph, the present application provides a geological disaster knowledge graph construction method, system and storage medium to improve the accuracy of the geological disaster knowledge graph; the following will be described in detail from two embodiments.

[0098] Embodiment one:

[0099] In the embodiments of the present application, the construction of the geological disaster knowledge graph is realized by the geological disaster knowledge graph construction method and system provided by the present application; in order to illustrate the utility of the present application, in embodiment one, taking three common types of geological disasters, i.e., earthquakes, landslides and debris flows, as examples, a geological disaster knowledge graph targeting earthquakes, landslides and debris flows is established. The method and system provided by the present application are described in the process of establishing the geological disaster knowledge graph; in Figure 1 and Figure 2 respectively show the flowchart of the method and the structure diagram of the system of the present application; as shown in Figure 1 and Figure 2 , the steps of the method of the present application include: S10. collecting geological disaster related multi-modal information; S20. performing deduplication processing on the multi-modal information; S30. classifying the deduplicated multi-modal information; S40. information correction; S50. constructing a geological disaster knowledge graph; S60. optimizing the geological disaster knowledge graph; Figure 2The construction system of the geological disaster knowledge graph given in the specification comprises: an information collection unit, an information classification unit, an information correction unit, a knowledge graph construction unit, a knowledge graph optimization unit, a visualization unit and a maintenance unit; and the specific geological disaster knowledge graph construction process comprises:

[0100] The information collection unit of the system is used to collect multi-modal data about earthquakes, landslides and debris flows, corresponding to Figure 1 S10; wherein the information collection unit mainly obtains text information, video information and audio information related to earthquakes, landslides and debris flows through the network; the text information is mainly obtained from literature, forums, social media and official electronic reports; the video information is mainly obtained from news reports, social media videos and geological official report analysis videos; and the audio information is mainly obtained from podcasts, news broadcasts and geologist interview audio;

[0101] Further, the multi-modal data is subjected to deduplication processing through S20, to obtain deduplicated multi-modal data;

[0102] Further, the information classification unit of the system is used to classify the deduplicated multi-modal data into disaster description information and disaster auxiliary description information, corresponding to S30; wherein the disaster description information refers to the relevant description information about a specific disaster type in the collected multi-modal information; and the disaster auxiliary description information refers to the relevant information in the case where the disaster type is not directly described; for example, the disaster description information about the epicenter location, magnitude and impact of an earthquake; and the disaster auxiliary description information about the relevant geological structure, soil type and climate conditions;

[0103] The classification process of the deduplicated multi-modal information comprises:

[0104] Obtaining description information of geological disasters (earthquakes, landslides and debris flows);

[0105] Further, the description information is subjected to keyword extraction, to obtain disaster type keywords;

[0106] Further, the disaster type keywords are subjected to vectorization, to obtain disaster type matching word vectors; refer to Table 1, which shows the conversion process of the description information of earthquakes, landslides and debris flows to disaster type matching word vectors;

[0107] Table 1: Disaster type description vector conversion

[0108]

[0109] The conversion process from "description information" to "disaster type matching word vector" is given in Table 1; in this process, the Word2Vec technology, a neural network-based model, is used to convert words into fixed-dimensional vectors, capturing the semantic relationship between words. Through the Skip-gram or CBOW (Continuous Bag of Words) model, word vectors are generated.

[0110] Further, the de-duplication multi-modal information is divided into a text data set, a video data set and an audio data set according to the information type;

[0111] Further, all text descriptions are extracted from the text data set using NLP (Natural Language Processing) technology to obtain text description information;

[0112] All video descriptions are extracted from the video data set to obtain video description information; the acquisition process of the video description information includes: obtaining the video metadata of the videos in the video data set: each video usually contains some metadata (such as title, description, duration and release time, etc.); using video processing libraries (such as OpenCV, FFmpeg and Moviepy, etc.) to analyze the video metadata and extract the text information in the frame; organizing the extracted description information to create a structured data format (such as list, dictionary or data frame);

[0113] All audio descriptions are extracted from the audio data set to obtain audio description information; including: extracting audio metadata: each audio file usually contains some metadata (such as title, duration and sampling rate, etc.), which can be extracted using some audio processing libraries (such as pydub, mutagen, etc.); using audio processing libraries: the basic information of the audio file can be read through these libraries; using speech recognition technology to transcribe the audio into text to obtain the audio description information.

[0114] Further, feature extraction is performed on the text description information, the video description information and the audio data respectively to obtain text description features, video description features and audio description features;

[0115] Further, the similarity of the text description features, the video description features and the audio description features with the disaster type matching word vector is calculated to obtain text similarity, video similarity and audio similarity;

[0116] The calculation formula of the text similarity, the video similarity and the audio similarity is represented as:

[0117]

[0118] CosSim() represents a similarity calculation function, i.e. cosine similarity; represents the text description feature; represents the video description feature; represents the audio description feature; represents the disaster type matching word vector; represents a norm operator;

[0119] Further, when the text similarity, the video similarity and the audio similarity exceed a similarity matching threshold; the deduplicated multi-modal information is taken as the disaster explanation information of the corresponding disaster type; when the text similarity, the video similarity and the audio similarity do not exceed the similarity matching threshold, the deduplicated multi-modal information is taken as the disaster auxiliary explanation information; in the embodiment of the present application, in order to obtain more disaster explanation information, the similarity matching threshold is set to 0.6.

[0120] In the embodiment of the present application, the deduplicated multi-modal information is classified into disaster explanation information and disaster auxiliary explanation information; in the classification process, the feature similarity between the text information, the video information and the audio information and the feature of the disaster type matching word vector is used to achieve the classification effect. Not only can the accuracy and usability of the geological disaster knowledge graph be improved, but also a solid data foundation can be provided for future disaster management and emergency response.

[0121] Further, the information correction unit of the system corrects the disaster explanation information and the disaster auxiliary explanation information to obtain corrected disaster explanation information and corrected disaster auxiliary explanation information, corresponding to S40; wherein, as shown in Figure 3 The information correction unit includes an expert knowledge correction module and a machine knowledge correction module;

[0122] The expert knowledge correction module includes: obtaining the related knowledge of N geologists on earthquakes, landslides and debris flows, and constructing an expert knowledge base; in the embodiment of the present application, the related knowledge of 10 geologists on earthquakes, landslides and debris flows is obtained; the structure of the expert knowledge base includes: expert number attribute description, disaster type attribute description, disaster area type attribute description, occurrence mechanism attribute description, influence factor attribute description and factor action attribute description; refer to Table 2, which illustrates the knowledge provided by the geologist numbered P001 in the expert knowledge base;

[0123] Table 2 Part of the expert knowledge in the expert knowledge base The content of each attribute under different disaster types is shown in Table 2 above. The expert knowledge base contains more detailed knowledge according to the knowledge provided by geology experts, and the technical solutions contained therein do not deviate from the present application.

[0124] Further, the information in the expert knowledge base is generated into an expert knowledge matching text;

[0125] Further, a multi-layer knowledge matching index is established for the expert knowledge matching text; wherein the multi-layer knowledge matching index includes: attribute key values established according to disaster types, disaster area types, occurrence mechanisms, influencing factors and factor effects; each key value is associated with a corresponding expert knowledge matching text; wherein the expert knowledge matching text stores a content index value; the specific information in the expert knowledge base can be queried according to the content index value; refer to Figure 4 , which gives a visual explanation of the multi-layer knowledge matching index; from Figure 4 , it can be known that the multi-layer knowledge matching index is divided into four layers, the first layer is the attribute key value layer corresponding to the disaster type, wherein the attribute key value corresponding relationship is: "attribute key value 1<->earthquake"; "attribute key value 2<->landslide"; "attribute key value 3<->debris flow"; the second layer is the attribute key value layer corresponding to the disaster area type, the attribute key value corresponding relationship is, for example, "attribute key value A1<->mountainous area" "attribute key value A2<->coastal area" and the like; the third layer is the related attribute of disaster occurrence, including: occurrence mechanism, influencing factor and factor effect, corresponding attribute key value; the fourth layer is a link layer, which is linked to the expert knowledge base through a database structure to obtain the corresponding expert index; in Figure 4 , the first layer and the second layer are bidirectionally linked, and the corresponding disaster area type can be queried through the disaster type, and the corresponding disaster type can be queried through the disaster area type;

[0126] Further, the disaster explanation information and the disaster auxiliary explanation information are matched with the expert knowledge matching text according to the multi-layer knowledge matching index;

[0127] Further, information consistency calculation is performed according to the matching result to obtain an information consistency value;

[0128] Further, information consistency calculation is performed according to the matching result to obtain an information consistency value; when the information consistency value does not exceed a consistency threshold, an information correction algorithm is used to obtain a candidate expert correction information set;

[0129] Among them, regarding the above information matching and consistency calculation process, the following examples are explained; for example, the description of one information in the disaster description information is: "In mountainous area A, a strong earthquake occurred, with epicenter at longitude 114.3, latitude 30.5, and magnitude 6.8. The earthquake was strongly felt, and residents of several surrounding villages felt obvious shaking. The geological structure of the area is complex, and the earthquake caused severe movement of the earth's crust and triggered a chain reaction. Due to the occurrence of the earthquake, landslides occurred in multiple places in the region, especially in sunny weather, with saturated soil moisture, greatly increasing the risk of landslides. The earthquake also caused the collapse of some buildings, causing road closures and severe traffic disruptions. Influencing factors: (1) Geological structure: the geological structure of the area is loose and easily affected by earthquakes; (2) Epicenter distance: the epicenter is close to nearby villages, exacerbating the degree of impact; (3) Focal depth: the focal depth is shallow, causing significant ground shaking and increasing the occurrence of secondary disasters. Secondary disaster description: Landslide: Landslides caused by the earthquake blocked dozens of roads, and some villages lost contact with the outside world. Building damage: multiple buildings were severely damaged by the shaking, and some houses collapsed, trapping people inside." According to the above information, a keyword list can be obtained;

[0130] Table 3 shows an example keyword list

[0131]

[0132] According to the disaster type in Table 3 ("earthquake, landslide"), attribute key value 1 and attribute key value 2 are obtained according to the first layer of the multi-layer knowledge matching index;

[0133] Further, the consistency of the region type in Table 3 is calculated using the consistency calculation formula, and the consistency of the second layer of disaster region type is calculated according to the attribute key value 1 and attribute key value 2 index, and when the consistency exceeds the consistency threshold, it is matched to the third layer; if it does not exceed, record the corresponding keyword of the region type in Table 3 as modification information;

[0134] According to the keyword corresponding to the region type in Table 3, the "region type" is consistent with the second layer index;

[0135] Further, it is matched to the third layer; through the attribute key value of the occurrence mechanism, influencing factors and factor action of the third layer, the corresponding knowledge of the expert knowledge base is obtained; the obtained knowledge is consistent with the corresponding keywords of the epicenter information, impact description, influencing factors, secondary disasters and weather conditions in Table 3;

[0136] In the consistency calculation process, the keyword corresponding to the weather condition in Table 3 is "sunny", and the consistency calculation formula of the expert knowledge is: Wherein, C represents a consistency value (range is usually between 0 and 1, 1 represents complete consistency, 0 represents complete inconsistency); P represents "sunny weather"; Q represents a disaster type; W(P, Q) represents a consistency weight of "sunny weather" and the disaster type; P' represents all possible weather conditions under the disaster type;

[0137] Referring to Table 4, consistency weights in the case of "geology and landslide" are given in Table 4;

[0138] Table 4: Meteorological consistency weight distribution in the case of geology and landslide

[0139]

[0140] Calculate the consistency value:

[0141] Further, according to the calculated consistency value, the consistency threshold is not exceeded; therefore, the keyword corresponding to the meteorological condition in Table 3 is error information; wherein, the consistency threshold for the meteorological condition is set to 0.7 in the embodiment of the application; the consistency value calculation for other item information is the same as the above principle, and the consistency threshold matched according to different information is set to a high threshold to ensure the accuracy of information verification;

[0142] For error information, an information correction algorithm is used to correct the information; wherein, the information correction algorithm mainly includes the following steps:

[0143] Step 1: Query the knowledge base: retrieve information similar to the error from the expert knowledge base, and compare;

[0144] Step 2: Calculate the matching degree: calculate the matching degree of the marked information and the related information in the knowledge base, and quantify;

[0145] Step 3: Find the matching high consistency information from the knowledge base, then replace the original error information;

[0146] Step 4: According to the correction strategy, update the original information to form a candidate expert correction information set;

[0147] The machine knowledge correction module includes:

[0148] The named entity recognition algorithm is used to recognize the entity field of the disaster description information and the disaster auxiliary description information; wherein, the named entity recognition algorithm uses a dictionary matching method as follows:

[0149] Match through a pre-defined entity dictionary, which is suitable for application in the scenarios of earthquake, landslide and debris flow; wherein, the pre-defined content of the entity dictionary is shown in Table 5;

[0150] Table 5 predefined content of entity dictionary

[0151]

[0152] Further, the disaster description information and the disaster auxiliary description information are subjected to word segmentation processing by using Jieba. Jieba is a very popular Chinese word segmentation tool, mainly used for word segmentation processing of Chinese text. It combines rule-based dictionary and statistical-based model, and can efficiently and accurately complete the word segmentation task.

[0153] Further, by traversing the word segmentation result, a predefined entity dictionary is used for matching to identify the entity field in the text.

[0154] Further, data statistics are performed on the entity field, and the co-occurrence probability of the entity field is calculated to obtain an entity field co-occurrence matrix. The obtaining process of the entity field co-occurrence matrix includes: creating a co-occurrence matrix, and the rows and columns are all the identified entities.

[0155] Further, the co-occurrence times between each group of identified entities are counted.

[0156] Further, each value in the co-occurrence matrix is divided by the total number of entities to obtain the entity field co-occurrence matrix.

[0157] Further, the disaster description information and the disaster auxiliary description information are filtered according to the entity field co-occurrence matrix to obtain a set of geological disaster screening information. When the co-occurrence rate in the entity field co-occurrence matrix is greater than a preset threshold, the entity fields have relevance. For example, if the co-occurrence rate between “earthquake” and “building collapse” is higher than the threshold, it can be considered that the two entities have significant relevance. In the embodiment of the present application, the preset threshold is set to 0.5.

[0158] Further, the set of geological disaster screening information is subjected to cross-validation. In the embodiment of the present application, the set of geological disaster screening information is divided into K parts by using K-fold cross-validation method.

[0159] Further, when the cross-validation value of any information in the set of geological disaster screening information is lower than a cross-validation threshold, the information that does not meet the condition is corrected according to the same type of information to obtain a set of candidate machine correction information.

[0160] For example, according to the following entities: entity A: earthquake; entity B: building collapse; entity C: landslide; and entity D: flood.

[0161] The entity field co-occurrence matrix of the above entities is calculated. As shown in Table 6:

[0162] Table 6 example entity field co-occurrence matrix

[0163] Entity Earthquake Building collapse Landslide Flood Earthquake 1 0.6 0.4 0.2 Building collapse 0.6 1 0.3 0.1 Landslide 0.4 0.3 1 0.5 Flood 0.2 0.1 0.5 1

[0164] According to the preset threshold value 0.5, the entities with significant relevance will be screened according to the co-occurrence rate; including: earthquake and building collapse: co-occurrence rate is 0.6 (greater than 0.5); earthquake and landslide: co-occurrence rate is 0.4 (less than 0.5, not retained); building collapse and landslide: co-occurrence rate is 0.3 (less than 0.5, not retained); landslide and flood: co-occurrence rate is 0.5 (equal to 0.5, retained); other combinations are less than the threshold value.

[0165] According to the above screening, the geological disaster screening information set is obtained, including: earthquake and building collapse; landslide and flood.

[0166] K-fold cross-validation is performed; set K=2, divide the screening information set into 2 parts; part 1: earthquake and building collapse; part 2: landslide and flood.

[0167] Model training and verification are performed on the two parts respectively, and different training data are used to verify the effect of the model;

[0168] The cross-validation value is evaluated: the cross-validation value of part 1 is 0.4 (lower than the threshold value, the cross-validation threshold value is set to 0.5); the cross-validation value of part 2 is 0.7 (higher than the threshold value)

[0169] Since the cross-validation value of part 1 is lower than the threshold value 0.5, the information of this part needs to be corrected, which can be adjusted according to the same type of information (for example, other information related to earthquake and building collapse);

[0170] Further, the candidate expert correction information set and the candidate machine correction information set are merged to obtain a comprehensive correction information set;

[0171] Further, when the candidate expert correction information set and the candidate machine correction information set have the same correction item; if the correction information of the correction item is the same, merge, if the correction information of the correction item is not the same, then regularize and merge;

[0172] In the embodiments of the present application, the formula of regularization is set as: Wherein, C r represents the information after regularization; C c represents the candidate expert correction information; C m represents the machine correction information; w c and w m represent different regularization weights, which can be set according to historical accuracy, reliability or other indicators;

[0173] Further, the disaster explanation information and the disaster auxiliary explanation information are corrected by using the comprehensive correction information set.

[0174] The information correction process is shown in the following example:

[0175] The candidate expert correction information set includes item 1: "Landslide may cause road interruption, immediate traffic diversion is needed"; item 2: "Residents in affected areas should be prepared for evacuation";

[0176] The candidate machine correction information set includes item 1: "Landslide will cause road closure, immediate traffic control is recommended"; item 2: "Residents in disaster areas need to evacuate in time to ensure safety";

[0177] The merging process: merging the candidate expert correction information set and the candidate machine correction information set to obtain the comprehensive correction information set.

[0178] (1) Check the same correction item

[0179] For item 1: expert information: "Landslide may cause road interruption, immediate traffic diversion is needed"; machine information: "Landslide will cause road closure, immediate traffic control is recommended"; the information is not the same, and needs to be normalized.

[0180] The normalization result is: C r = "Landslide may cause road interruption, immediate traffic diversion is recommended";

[0181] (2) For item 2, since the two pieces of information are similar, they can be directly merged:

[0182] The merging result is: "Residents in affected areas should be prepared for evacuation, and need to evacuate in time to ensure safety."

[0183] The final comprehensive correction information set is:

[0184] Item 1: "Landslide may cause road interruption, immediate traffic control and traffic diversion are recommended";

[0185] Item 2: "Residents in affected areas should be prepared for evacuation, and need to evacuate in time to ensure safety."

[0186] The comprehensive correction information set is used for correction; next, the comprehensive correction information set is applied to the correction of disaster explanation information and auxiliary explanation information. The assumed disaster explanation information is: "The recent landslide event has caused multiple roads to be closed, and residents need to evacuate";

[0187] Correction process: use the information in the comprehensive correction information set to update the original description, for: "the recent landslide event caused multiple road interruptions, it is recommended to immediately implement traffic control and traffic diversion. Residents in the affected area should be prepared for evacuation and need to evacuate in time to ensure safety."

[0188] The final corrected disaster description information is: "the recent landslide event caused multiple road interruptions, it is recommended to immediately implement traffic control and traffic diversion. Residents in the affected area should be prepared for evacuation and need to evacuate in time to ensure safety."

[0189] In the embodiments of the present application, the geological disaster information correction method is used to correct the information; the method is comprehensive information correction through expert knowledge correction and machine knowledge correction; wherein, the expert knowledge correction establishes an expert knowledge base according to multiple geological experts; and the acquired information is matched with the expert knowledge through an index method, and a candidate expert correction information set is generated according to the matching result; the machine knowledge correction is to collect entity information of the information, filter the information according to the co-occurrence relationship of the entity fields, and cross-verify the filtered information to obtain a machine correction information set; the original information is corrected according to the candidate expert correction information set and the machine correction information set; in this way, not only the accuracy and reliability of the information can be improved, but also the richness and adaptability of the knowledge graph can be enhanced, which provides sufficient data basis for the construction of the geological disaster knowledge graph.

[0190] Further, the knowledge graph construction unit of the system constructs a geological disaster knowledge graph related to landslides, debris flows and earthquakes, corresponding to step S50; the specific process includes:

[0191] Construct an initial geological disaster graph;

[0192] Wherein, the initial geological disaster graph is represented as: {Node(GHT, GHAT), Edg(NULL)}; wherein, Node() represents an initial node; GHT represents the disaster type; GHAT represents the geological disaster area type; Edg() represents an initial edge, and the initial value is NULL, i.e. empty; see Figure 5 , the initial geological disaster graph of the embodiments of the present application is given;

[0193] Input the corrected disaster description information and the initial geological disaster graph into the geological disaster knowledge graph generation model; wherein, see Figure 6 , the geological disaster knowledge graph generation model includes:

[0194] The receiving layer includes: a disaster information receiving module for receiving the corrected disaster description information; a graph receiving module for receiving the initial geological disaster graph;

[0195] a processing layer configured to process the disaster explanation information to obtain structured geological disaster explanation features;

[0196] a relationship dependency mining layer configured to mine dependency relationships of the structured geological disaster explanation features, including: establishing a structured geological disaster explanation feature matrix according to the structured geological disaster explanation features; inputting the structured geological disaster explanation feature matrix into a multi-condition constraint random forest model to obtain a structured geological disaster dependency matrix;

[0197] The multi-condition constraint random forest model includes: an input layer configured to receive the structured geological disaster explanation feature matrix; a tree generation layer configured to randomly generate multiple trees; each tree independently selects feature data corresponding to different disaster types; a condition constraint setting layer configured to set a constraint condition for each tree to constrain branches of the tree; the constraint rule of the condition constraint setting layer is set as: occurrence reliability constraint between the disaster type and the geological disaster region type; association constraint between the disaster type and a disaster occurrence influencing factor; influence constraint between the geological disaster region type and a disaster influence;

[0198] The constraint rule of the condition constraint setting layer is as follows:

[0199] (1) occurrence reliability constraint between the disaster type and the geological disaster region type: P(DR) ≥ θ1; wherein P(DR) represents a probability of occurrence of the disaster type D given the region type R; θ1 represents a set reliability threshold value, representing a minimum probability requirement for occurrence of the disaster;

[0200] (2) association constraint between the disaster type and the disaster occurrence influencing factor: I(D, F) ≥ θ2; wherein I(D, F) represents a correlation coefficient between the disaster type D and the influencing factor F; θ2 represents a set association threshold value, representing a minimum requirement for association;

[0201] (3) influence constraint between the geological disaster region type and the disaster influence: C(R, I D ) ≥ θ3; wherein C(R, I D ) represents an influence degree of the region type D on the disaster influence I D ; θ3 represents a set influence threshold value, representing a minimum requirement for influence degree;

[0202] In an embodiment of the present application, the correlation between geological disaster attributes is obtained through the relationship dependency mining layer of the geological disaster knowledge graph generation model; wherein, a multi-condition constraint random forest model is established in the relationship dependency mining layer to perform relationship mining; the multi-condition constraint random forest model is based on the random forest model and sets a condition constraint setting layer; the actual attribute relationship in the geological disaster is set as a rule to filter the connection between unrelated attributes; the multi-condition constraint random forest model can significantly enhance the construction quality and practicality of the geological disaster knowledge graph.

[0203] The voting layer is used to summarize the analysis results of each tree; the dependency calculation layer is used to calculate the dependency according to the analysis results of the voting layer; and the output layer is used to output the structured geological hazard dependency matrix.

[0204] The geological disaster knowledge graph generation layer is used to generate the geological disaster knowledge graph;

[0205] The visualization output layer is used to visualize the geological disaster knowledge graph; see Figure 7 ;

[0206] Furthermore, the geological disaster knowledge graph is optimized using the knowledge graph optimization unit of the system to obtain an optimized geological disaster knowledge graph, corresponding to the above-mentioned step S60; the geological disaster knowledge graph optimization process includes:

[0207] The geological hazard knowledge graph is converted into a geological hazard multiple linked list; wherein the geological hazard multiple linked list includes: linked list nodes, linked list multiple edges and linked list connection directions; wherein the linked list nodes are nodes in the geological hazard knowledge graph; the linked list multiple edges are connection relationships between nodes in the geological hazard knowledge graph; the linked list multiple edges record various information elements related to geological hazards; the linked list connection direction is the connection strength between nodes in the geological hazard knowledge graph, and the connection direction of the linked list multiple edges is set according to the connection strength; refer to Figure 8 , Figure 8 Content is for Figure 7 The geological disaster knowledge graph is converted into a multiple linked list of geological disasters;

[0208] Furthermore, the correlation between the auxiliary explanation information for correcting the disaster and the data in the geological disaster multi-link list is calculated; the specific data correlation calculation process includes:

[0209] Processing the corrected disaster auxiliary description information, including removing irrelevant words, unifying terminology, etc.;

[0210] performing data standardization on the processed disaster correction auxiliary explanation information to obtain standard disaster correction auxiliary explanation information;

[0211] The standard disaster correction auxiliary instruction information is vectorized, and relevant features are extracted;

[0212] The acquired features are used to construct a feature matrix;

[0213] The features of the feature matrix are analyzed for relevance with the data of the geological disaster multi-linked list using a chi-square test; the steps are:

[0214] Step 1: According to the relationship between the features of the feature matrix and the multi-linked list nodes, select appropriate features for analysis;

[0215] Step 2: Construct a contingency table according to the frequency between the features and the nodes;

[0216] Referring to Table 7, an example of a contingency table for the embodiment of the present application is shown;

[0217] Table 7 Contingency Table Example

[0218] Landslide Mudflow Earthquake Total Sand 10 5 2 17 Clay 8 12 4 24 Loam 6 3 1 10 Total 24 20 7 51

[0219] In Table 7, the relationship between different soil types and different geological disasters is shown; the rows represent different soil types (sand, clay, and loam). The columns represent different geological disaster types (landslide, debris flow, and earthquake). The numbers in Table 7 represent the number of a particular geological disaster observed under a particular soil type. The "Total" column and row on the right and bottom, respectively, represent the total for each soil type and each disaster.

[0220] Step 3: Perform a chi-square test to evaluate the independence between the features and the multi-linked list, and calculate the correlation value to determine the significance of the correlation.

[0221] Further, according to the correlation calculation results, auxiliary information elements of the disaster correction auxiliary instruction information are extracted; including:

[0222] (1) Determine the features that are significantly related to the type of geological disaster. For example, features such as soil type, rainfall, and topographic features, etc.

[0223] (2) According to the strength of the correlation, assign weights to different features to identify the most influential features;

[0224] (3) Extract information elements, classified information extraction: soil type: extract the physical and chemical properties of various types of soil, such as permeability, viscosity, etc. Geological structure: collect information about related geological structures, such as the existence and distribution of faults and folds.

[0225] Numerical information extraction: meteorological data: extract historical meteorological data related to geological disasters, such as precipitation, temperature, wind speed, etc. during a specific period.

[0226] Further, integrate the auxiliary information elements with the chain table nodes and the chain table multiple edges of the geological disaster multiple chain table, and update the connection strength and modify the connection direction; including:

[0227] (1) Identify target nodes: According to the extracted auxiliary information elements, determine the chain table nodes that need to be integrated. For example, the nodes that may need to be updated include "landslide", "soil type", etc.

[0228] (2) Update multiple edge information:

[0229] Enhance edge information: Add auxiliary information elements to existing multiple edges, such as including rainfall, historical disaster data, etc. into the information attributes of related edges.

[0230] New edge: If there is a new relationship between auxiliary information elements and existing nodes, a new multiple edge can be created. For example, a new edge from the "rainfall" node to the "landslide" node.

[0231] (3) Update connection strength:

[0232] Calculate new connection strength: Use statistical methods (such as linear regression, logistic regression, etc.) to calculate the new connection strength between auxiliary information elements and geological disasters. For example, analyze the relationship between rainfall and landslide occurrence rate.

[0233] (4) Modify connection direction: Determine the direction of the edge according to the updated connection strength.

[0234] (5) Update nodes: Update the integrated auxiliary information elements to the corresponding nodes to ensure the integrity of the information.

[0235] Further, the updated geological disaster multiple chain table is reversely converted into an optimized geological disaster knowledge graph. See Figure 9 , Figure 9 for the optimized geological disaster knowledge graph.

[0236] In the embodiments of the present application, a geological disaster knowledge graph optimization method is proposed to optimize the knowledge graph generated by the knowledge graph construction unit; the method converts the geological disaster knowledge graph into a multiple chain table form, clearly reflecting the relationship between nodes; and by correcting the association between disaster auxiliary explanation information and multiple chain table nodes, auxiliary information elements are obtained, thereby realizing the update of edges and nodes; the updated multiple chain table is reversely converted into an optimized geological disaster knowledge graph; the method not only further improves the accuracy of the geological disaster knowledge graph, but also enriches the content of the geological disaster knowledge graph.

[0237] Further, the visualization unit of the system displays the geological disaster knowledge graph;

[0238] Further, the maintenance unit is used to update the geological disaster knowledge graph by obtaining new information.

[0239] In the embodiments of the present application, the method and the system are combined to realize the construction of the geological disaster knowledge graph of earthquakes, landslides and debris flows. The following points are mainly considered. First, multi-modal information about geological disasters is obtained from multiple channels. In order to ensure the reliability of the information, an information correction method is proposed. The method includes expert knowledge correction and machine knowledge correction. The expert knowledge correction corrects the original information by establishing an expert knowledge base of geologists and by information matching. The machine knowledge correction finds and corrects the incorrect information by statistical analysis of the collected information. The information corrected by the expert knowledge correction and the machine knowledge correction is combined to ensure the correctness of the information. Second, a geological disaster knowledge graph generation model is proposed. The model generates the geological disaster knowledge graph according to the corrected disaster description information and the initial geological disaster graph. In the model, the dependency relationship between the information and the disaster is mined by relationship dependency mining layer to establish the corresponding knowledge graph. This is conducive to the construction of the geological disaster knowledge graph. Finally, a geological disaster knowledge graph optimization method is proposed. The method converts the geological disaster knowledge graph into a multi-linked list form to clearly reflect the relationship between the nodes. Auxiliary information elements are obtained by correcting the association between the disaster auxiliary description information and the multi-linked list nodes to update the edges and nodes. The updated multi-linked list is converted into an optimized geological disaster knowledge graph. The above methods realize the accurate construction of the geological disaster knowledge graph.

[0240] Embodiment two:

[0241] In the first embodiment, the method and the system of the present application are used to realize the construction of the geological disaster knowledge graph of earthquakes, landslides and debris flows. In the embodiments of the present application, the types of geological disasters are increased to show good flexibility and to realize the expansion of the geological disaster graph. The specific process includes:

[0242] Collect multi-modal information related to geological disasters and perform deduplication processing to obtain deduplicated multi-modal information.

[0243] Further, the deduplicated multi-modal information is classified according to the disaster types to obtain disaster description information and disaster auxiliary description information.

[0244] Further, the disaster description information and the disaster auxiliary description information are corrected to obtain corrected disaster description information and corrected disaster auxiliary description information.

[0245] Further, an initial geological disaster graph is constructed.

[0246] Further, the correction disaster description information and the initial geological disaster map are input into a geological disaster knowledge graph generation model; the geological disaster knowledge graph generation model comprises:

[0247] The receiving layer comprises: a disaster information receiving module, configured to receive the correction disaster description information; and a graph receiving module, configured to receive the initial geological disaster map.

[0248] The processing layer is configured to process the correction disaster description information to obtain structured geological disaster description features.

[0249] The relationship dependence mining layer is configured to mine dependence relationships of the structured geological disaster description features, and comprises: establishing a structured geological disaster description feature matrix according to the structured geological disaster description features; wherein the structured geological disaster description feature matrix is expressed as: wherein GHTC i represents a disaster type feature of the ith structured geological disaster description feature; GHATC i represents a geological disaster area type feature of the ith structured geological disaster description feature; and GHTC represents the jth related information feature vector of the ith structured geological disaster description feature; M represents a total number of the structured geological disaster description features; and N represents a total number of related information feature vectors in the structured geological disaster description features.

[0250] The structured geological disaster description feature matrix is input into a multi-condition constraint random forest model to obtain a structured geological disaster dependence matrix.

[0251] The multi-condition constraint random forest model comprises: an input layer, configured to receive the structured geological disaster description feature matrix; a tree generation layer, configured to randomly generate a plurality of trees; wherein each tree independently selects feature data corresponding to a different data disaster type; a condition constraint setting layer, configured to set a limitation condition for each tree to constrain branches of the tree; wherein a constraint rule of the condition constraint setting layer is set as: an occurrence reliability constraint between the disaster type and the geological disaster area type; an association constraint between the disaster type and a disaster occurrence influencing factor; an influence constraint between the geological disaster area type and a disaster influence; a voting layer, configured to aggregate analysis results of the trees; a dependence relationship calculation layer, configured to calculate a dependence relationship according to the analysis results of the voting layer; and an output layer, configured to output the structured geological disaster dependence matrix.

[0252] The geological disaster knowledge graph generation layer is configured to generate a geological disaster knowledge graph.

[0253] The visual output layer is configured to visually output the geological disaster knowledge graph.

[0254] Further, the geological disaster knowledge graph is optimized by using the correction disaster auxiliary description information, to obtain an optimized geological disaster knowledge graph.

[0255] The optimization process of the geological disaster knowledge graph includes:

[0256] The geological disaster knowledge graph is converted into a geological disaster multi-linked list, which includes linked list nodes, linked list multi-edges, and linked list connection directions. The linked list nodes are nodes in the geological disaster knowledge graph. The linked list multi-edges are connection relationships between nodes in the geological disaster knowledge graph. The linked list multi-edges record various information elements related to geological disasters. The linked list connection directions are connection strengths between nodes in the geological disaster knowledge graph, and the connection directions of the linked list multi-edges are set according to the connection strengths.

[0257] Further, the correlation between the correction disaster auxiliary description information and the data in the geological disaster multi-linked list is calculated.

[0258] Further, auxiliary information elements of the correction disaster auxiliary description information are extracted according to the correlation calculation results.

[0259] Further, the auxiliary information elements are integrated with the linked list nodes and the linked list multi-edges of the geological disaster multi-linked list, and the connection strengths are updated, and the connection directions are modified.

[0260] Further, the updated geological disaster multi-linked list is reversely converted into the optimized geological disaster knowledge graph.

[0261] Embodiment Three:

[0262] A storage medium for constructing a geological disaster knowledge graph, the storage medium storing a geological disaster knowledge graph construction program. When the geological disaster knowledge graph construction program is executed by a processor, the entire content of Embodiment One or Embodiment Two is realized.

[0263] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a geological disaster knowledge graph, characterized in that, The method comprises the following steps: Collecting multi-modal information related to geological disasters and performing deduplication processing to obtain deduplicated multi-modal information; Performing keyword extraction to obtain disaster type keywords; Vectorizing the disaster type keywords to obtain disaster type matching word vectors; According to the disaster type, classifying the deduplicated multi-modal information by the similarity between the features of the text information, video information and audio information and the features of the disaster type matching word vectors, and by a similarity threshold to obtain disaster explanation information and disaster auxiliary explanation information; Correcting the disaster explanation information and the disaster auxiliary explanation information to obtain corrected disaster explanation information and corrected disaster auxiliary explanation information; The correction of the disaster explanation information and the disaster auxiliary explanation information comprises an expert knowledge correction mechanism and a machine knowledge correction mechanism; The expert knowledge correction mechanism corrects information by geologist knowledge, comprising: acquiring the relevant knowledge of N geologists on geological disasters to construct an expert knowledge base; generating expert knowledge matching text from the information in the expert knowledge base; establishing a multi-layer knowledge matching index for the expert knowledge matching text; matching the disaster explanation information and the disaster auxiliary explanation information with the expert knowledge matching text by using the multi-layer knowledge matching index; calculating information consistency according to the matching result to obtain an information consistency value; when the information consistency value does not exceed a consistency threshold, using an information correction algorithm to obtain a candidate expert correction information set; The machine knowledge correction mechanism corrects information by machine algorithm, comprising: identifying the entity field of the disaster explanation information and the disaster auxiliary explanation information by using a named entity recognition algorithm; performing data statistics on the entity field and calculating the co-occurrence probability of the entity field to obtain an entity field co-occurrence matrix; screening the disaster explanation information and the disaster auxiliary explanation information according to the entity field co-occurrence matrix to obtain a geological disaster screening information set; cross-validating the geological disaster screening information set; when the cross-validation value of any information in the geological disaster screening information set is lower than a cross-validation threshold, correcting the information that does not meet the conditions according to similar information to obtain a candidate machine correction information set; Merging the candidate expert correction information set and the candidate machine correction information set to obtain a comprehensive correction information set; when the candidate expert correction information set and the candidate machine correction information set have the same correction item; if the correction information of the correction item is the same, merging; if the correction information of the correction item is not the same, then normalizing and merging; using the comprehensive correction information set to correct the disaster explanation information and the disaster auxiliary explanation information; Constructing an initial geological disaster graph; Inputting the corrected disaster explanation information and the initial geological disaster graph into a geological disaster knowledge graph generation model; the geological disaster knowledge graph generation model comprises: The receiving layer comprises a disaster information receiving module for receiving the corrected disaster explanation information and a graph receiving module for receiving the initial geological disaster graph; The disaster auxiliary explanation information is processed to obtain a structured geological disaster explanation feature; The relationship dependency mining processing layer is configured to mine dependency relationships of the structured geological disaster explanation feature, and includes: establishing a structured geological disaster explanation feature matrix according to the structured geological disaster explanation feature; and inputting the structured geological disaster explanation feature matrix into a multi-condition constraint random forest model to obtain a structured geological disaster dependency matrix; The geological disaster knowledge graph generation layer is configured to generate a geological disaster knowledge graph; The visual output layer is configured to visually output the geological disaster knowledge graph; The geological disaster knowledge graph is optimized by using the disaster auxiliary explanation information to obtain an optimized geological disaster knowledge graph. 2.The method of claim 1, wherein, The classification process of the deduplicated multi-modal information includes: obtaining description information corresponding to different types of geological disasters; extracting keywords from the description information to obtain disaster type keywords; vectorizing the disaster type keywords to obtain disaster type matching word vectors; dividing the deduplicated multi-modal information into a text data set, a video data set, and an audio data set according to information types; extracting all text descriptions from the text data set to obtain text description information; extracting all video descriptions from the video data set to obtain video description information; extracting all audio descriptions from the audio data set to obtain audio description information; extracting features from the text description information, the video description information, and the audio data to obtain text description features, video description features, and audio description features; calculating similarities of the text description features, the video description features, and the audio description features with the disaster type matching word vectors to obtain text similarity, video similarity, and audio similarity; when the text similarity, the video similarity, and the audio similarity exceed a similarity matching threshold; taking the deduplicated multi-modal information as the disaster explanation information of a corresponding disaster type; and when the text similarity, the video similarity, and the audio similarity do not exceed the similarity matching threshold, taking the deduplicated multi-modal information as the disaster auxiliary explanation information. 3.The method of claim 1, wherein, The multi-condition constraint random forest model includes: The input layer is configured to receive the structured geological disaster explanation feature matrix; The tree generation layer is configured to randomly generate multiple trees; each tree independently selects feature data corresponding to different data disaster types; The condition constraint setting layer is configured to set a limit condition for each tree to constrain branches of the tree; the constraint rule of the condition constraint setting layer is set as: occurrence reliability constraint between a disaster type and a geological disaster area type; association constraint between the disaster type and a disaster occurrence influencing factor; and influence constraint between the geological disaster area type and a disaster influence; The voting layer is configured to aggregate analysis results of the trees; The dependency relationship calculation layer is configured to calculate dependency relationships according to the analysis results of the voting layer; The output layer is configured to output the structured geological disaster dependency matrix. 4.The method of claim 1, wherein , the initial geological disaster map is represented as: ; wherein, is represented as an initial node; is represented as the disaster type; is represented as a geological disaster area type; is represented as an initial edge, with an initial value of , i.e., a null value.

5. The method of claim 3, wherein The rule constraint of the condition constraint setting layer is as follows: (1) Disaster type and geological disaster regional type between the occurrence of reliability constraints: ; wherein, represents the probability of the occurrence of a disaster type for a given regional type ; represents a set reliability threshold value, representing the minimum probability requirement for the occurrence of the disaster; (2) The relevance constraint between the disaster type and the disaster occurrence influencing factor: ; wherein, represents the correlation coefficient between the disaster type and the influencing factor ; represents the relevance threshold set, representing the minimum requirement for relevance. (3) Impact constraint between geological disaster area type and disaster impact: ; wherein, represents the area type of the disaster impact ; and represents the impact threshold set, representing the minimum requirement reached by the impact degree. 6.The method of claim 1, wherein, The optimization process of the geological disaster knowledge graph comprises: The geological disaster knowledge graph is converted into a geological disaster multilink list; the geological disaster multilink list comprises: a link list node, a link list multiedge and a link list connection direction; wherein the link list node is a node in the geological disaster knowledge graph; the link list multiedge is a connection relationship between nodes in the geological disaster knowledge graph; the link list multiedge records a plurality of information elements related to geological disasters; and the link list connection direction is a connection strength between nodes in the geological disaster knowledge graph, and the connection direction of the link list multiedge is set according to the connection strength; The correlation between the disaster auxiliary explanation information and the data in the geological disaster multilink list is calculated; According to the correlation calculation result, the auxiliary information elements of the disaster auxiliary explanation information are extracted; The auxiliary information elements are integrated with the link list node and the link list multiedge of the geological disaster multilink list, and the connection strength is updated, and the connection direction is modified; The updated geological disaster multilink list is reversely converted into the optimized geological disaster knowledge graph. 7.A system for constructing a geological disaster knowledge graph, the system being configured to perform the method for constructing a geological disaster knowledge graph according to any one of claims 1 to 6. Comprise: An information collection unit for collecting information related to geological disasters from different channels; An information classification unit for classifying geological disaster information; An information correction unit for correcting the collected related information; a knowledge graph construction unit for constructing a geological disaster knowledge graph; A knowledge graph optimization unit for optimizing the geological disaster knowledge graph; A visualization unit for visualizing the generated geological disaster knowledge graph; A maintenance unit for periodically updating the geological disaster knowledge graph. 8.The system of claim 7, wherein, The information correction unit comprises an expert knowledge correction module and a machine knowledge correction module; The expert knowledge correction module comprises: Obtain the related knowledge of N geologists on geological disasters, and construct an expert knowledge base; Generate expert knowledge matching text from the information in the expert knowledge base; Establish a multi-layer knowledge matching index for the expert knowledge matching text; Match the disaster explanation information and the disaster auxiliary explanation information with the expert knowledge matching text according to the multi-layer knowledge matching index; According to the matching result, the information consistency is calculated to obtain an information consistency value; when the information consistency value does not exceed the consistency threshold, an information correction algorithm is used to obtain a candidate expert correction information set; The machine knowledge correction module comprises: An named entity recognition algorithm is used to identify the entity fields of the disaster explanation information and the disaster auxiliary explanation information; The entity fields are statistically analyzed, and the co-occurrence probability of the entity fields is calculated to obtain an entity field co-occurrence matrix; According to the entity field co-occurrence matrix, the disaster explanation information and the disaster auxiliary explanation information are screened to obtain a geological disaster screening information set; the geological disaster screening information set is cross-validated; when the cross-validation value of any information in the geological disaster screening information set is lower than the cross-validation threshold, the information that does not meet the condition is corrected according to the same type of information to obtain a candidate machine correction information set; The candidate expert correction information set and the candidate machine correction information set are merged to obtain a comprehensive correction information set; when the candidate expert correction information set and the candidate machine correction information set have the same correction items; if the correction information of the correction items are the same, they are merged; if the correction information of the correction items are different, they are regularized and then merged; the disaster description information and the disaster auxiliary description information are corrected using the comprehensive correction information set. 9.The system of claim 7, wherein, The knowledge graph construction unit includes: Construct an initial geological hazard map; Inputting the corrected disaster description information and the initial geological disaster map into a geological disaster knowledge graph generation model; wherein the geological disaster knowledge graph generation model includes: The receiving layer includes: a disaster information receiving module for receiving the corrected disaster description information; a map receiving module for receiving the initial geological disaster map; a processing layer for processing the corrected disaster description information to obtain structured geological disaster description features; A relationship dependency mining layer is used to mine the dependency relationships of the structured geological hazard description features, including: establishing a structured geological hazard description feature matrix based on the structured geological hazard description features; inputting the structured geological hazard description feature matrix into a multi-condition constrained random forest model to obtain a structured geological hazard dependency matrix; The geological disaster knowledge graph generation layer is used to generate the geological disaster knowledge graph; The visualization output layer is used to visualize the geological disaster knowledge graph. 10.The geological disaster knowledge graph construction system of claim 7, wherein, The knowledge graph optimization unit includes: The geological hazard knowledge graph is converted into a geological hazard multiple linked list; the geological hazard multiple linked list includes: linked list nodes, linked list multiple edges and linked list connection directions; wherein the linked list nodes are nodes in the geological hazard knowledge graph; the linked list multiple edges are connection relationships between nodes in the geological hazard knowledge graph; the linked list multiple edges record various information elements related to geological hazards; the linked list connection direction is the connection strength between nodes in the geological hazard knowledge graph, and the connection direction of the linked list multiple edges is set according to the connection strength; Calculate the correlation between the auxiliary explanation information for correcting disasters and the data in the geological disaster multiple linked list; extract the auxiliary information elements of the auxiliary explanation information for correcting disasters according to the correlation calculation result; integrate the auxiliary information elements with the linked list nodes and the multiple edges of the linked list of the geological disaster multiple linked list, update the connection strength, and modify the connection direction; reversely transform the updated geological disaster multiple linked list into an optimized geological disaster knowledge graph. 11.A storage medium for constructing a geological disaster knowledge graph, comprising the following steps of: The storage medium stores a program for constructing a geological hazard knowledge graph. When the program for constructing a geological hazard knowledge graph is executed by a processor, a method for constructing a geological hazard knowledge graph as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Tunnel construction geological disaster early warning and prevention and control intelligent decision-making method and auxiliary platform

    CN116562656A

  • Risk prediction method based on multi-modal data fusion

    CN117708746A