Disaster information fusion and semantic reasoning algorithm based on cross-modal graph attention mechanism

Through the disaster information fusion and semantic reasoning algorithm of the cross-modal graph attention mechanism, the problem of multimodal data fusion is solved, the efficient fusion and accurate prediction of disaster information are achieved, and disaster emergency management is supported.

CN120633837APending Publication Date: 2025-09-12TIANJIN BAIZE TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510704438.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

Smart Images

  • Figure CN120633837A_ABST
    Figure CN120633837A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of disaster information early warning, and particularly relates to a disaster information fusion and semantic reasoning algorithm based on a cross-modal diagram attention mechanism, which comprises a data preprocessing module, a cross-modal diagram construction module, a cross-modal diagram attention mechanism module, a semantic reasoning module and a result output module. According to the method, data sources from multiple modals can be effectively fused through a cross-modal graph attention mechanism, so that more comprehensive disaster information is provided, and meanwhile, the graph attention mechanism is realized through the graph attention mechanism, so that the model can dynamically and adaptively adjust the weight according to the correlation of different modal data, and the accuracy of information fusion is improved. The graph neural network is adopted for semantic reasoning of disaster events, the types and influences of disasters can be accurately predicted, decision support can be provided, disaster emergency management personnel can be helped to respond quickly, various different types of disaster data can be processed, and the method has high adaptability to data sources and formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of disaster information early warning technology, and specifically to a disaster information fusion and semantic reasoning algorithm based on a cross-modal graph attention mechanism. Background Art

[0002] Natural variations on Earth, including those induced by human activities, breed natural disasters within the Earth's surface environment, which comprises the atmosphere, lithosphere, hydrosphere, and biosphere. They occur constantly and everywhere. When these variations harm human society, they constitute natural disasters. This is because they inflict varying degrees of damage on human production and daily life, including the relationship between humans and nature, mediated by labor, and the interpersonal relationships associated with these. Disasters are always negative or destructive. Therefore, natural disasters are a manifestation of the conflict between humans and nature, possessing both natural and social attributes. They are one of the most severe challenges facing humanity in the past, present, and future.

[0003] With the rapid development of modern information technology, the means of obtaining disaster information are becoming increasingly diverse, covering multiple modalities such as remote sensing images, meteorological data, social media text, and sensor data. The data sources and formats of different modalities vary greatly, making it difficult for traditional disaster information processing methods to effectively integrate multi-source information, which in turn affects the timely prediction and decision support of disaster events.

[0004] The advantages of graph neural networks (GNNs) in processing graph-structured data are widely recognized, but effectively integrating multimodal data into GNNs and performing semantic reasoning remains a research challenge. In particular, how to dynamically fuse multi-source data through cross-modal graph attention mechanisms to achieve efficient semantic reasoning for disaster information analysis and emergency response remains an urgent issue.

[0005] The existing technology has the following defects or problems:

[0006] Existing disaster information fusion and semantic reasoning algorithms are unable to effectively integrate multimodal data sources. They lack unified modeling for data from different modalities, making efficient cross-modal data fusion impossible. Consequently, they are unable to provide more comprehensive disaster information. Furthermore, existing multimodal fusion methods mostly rely on shallow feature concatenation or simple weighted summation, which fails to fully exploit the deep semantic connections between different modalities. While graph neural networks (GNNs) can process graph-structured data, they lack effective cross-modal data fusion capabilities and are unable to organically combine multimodal information such as images, text, and sensor data.

[0007] It should be noted that the above content falls within the technical knowledge of the inventor and does not necessarily constitute prior art. Summary of the Invention

[0008] In response to the shortcomings of the existing technology, the present invention provides a disaster information fusion and semantic reasoning algorithm based on a cross-modal graph attention mechanism to solve the current problems.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a disaster information fusion and semantic reasoning algorithm based on a cross-modal graph attention mechanism, comprising:

[0010] The data preprocessing module is used to receive various disaster information from different data sources, including but not limited to image data, text data, and sensor data, and perform necessary preprocessing on such data to ensure the quality and consistency of the input data;

[0011] A cross-modal graph construction module, which converts preprocessed data into a graph structure where each modality is represented as a node and the relationships between modalities are represented by edges. Furthermore, the edges of the graph can be weighted based on the similarity between modalities.

[0012] A cross-modal graph attention mechanism module, which applies the graph attention mechanism to the aforementioned graph structure. It dynamically adjusts the weight of modal data based on the correlation between nodes. By calculating the similarity between nodes, the attention mechanism can identify the key parts of different modal information and strengthen the focus on important data.

[0013] The semantic reasoning module uses graph neural networks to perform semantic reasoning based on the fused graph data to deduce the type, severity, and affected area of ​​the disaster event. This module is based on the content of the disaster-related database and can obtain specific disaster prediction results through reasoning;

[0014] The output result module converts the reasoning results into visual reports and decision support information, enabling disaster emergency responders to make timely decisions.

[0015] In some embodiments, the specific process of the data preprocessing module is as follows:

[0016] Data cleaning process, first deal with missing data, choose to fill or delete rows and columns containing missing values, check for duplicate records in the data set, delete redundant data, and then use the standard deviation method to detect and remove outliers in the data;

[0017] Data normalization was performed by using the Z-score normalization method to convert numerical data into the same scale;

[0018] Lowercase words, remove stop words and punctuation marks, and perform stemming operations on text data;

[0019] For categorical data, use one-hot encoding to convert it into numerical data;

[0020] Data format conversion processing, using TF-IDF method to convert text data into vector representation;

[0021] For graphic data, resizing, denoising, and enhancement can be performed to increase the robustness of the model;

[0022] For time series data, timestamp unification, periodicity adjustment and smoothing are performed;

[0023] Feature selection and dimensionality reduction: Pearson correlation coefficient is used to select features that are closely related to the target variable, and principal component analysis is used to reduce the dimensionality of high-dimensional data to retain the main features in the data;

[0024] Data set partitioning and data integration: divide the data into K subsets, use K-1 subsets for training each time, and use the remaining subset for validation. Repeat K times to ensure the stability of the model.

[0025] Data balancing and preservation, reducing the data of majority class samples to achieve class balance;

[0026] Depending on the data size and subsequent processing requirements, you can choose one of the formats including CSV, JSON, Parquet and HDF5 to store the data. At the same time, for large-scale data sets, use data version control tools to manage data versions for easy traceability and management.

[0027] In some embodiments, the specific process of the cross-modal graph construction module is as follows:

[0028] Data feature extraction: using deep learning models to extract high-level features of images;

[0029] Use the Word2Vec pre-trained language model to extract the semantic features of the text;

[0030] Extract the audio spectrum features and Mel-frequency cepstral coefficients and convert them into numerical vectors;

[0031] Use the combination of image features and audio features of video frames to extract the spatiotemporal features of the video;

[0032] Cross-modal alignment processing: For time series data, different modal data need to be aligned according to timestamps;

[0033] For image, text, and audio modalities, mapping is required to enable them to be connected together;

[0034] Graph structure construction requires connecting data nodes of different modalities and expressing the relationship between them, as follows:

[0035] The data of each modality is converted into nodes. Nodes can be feature vectors of images, text, and audio. Edges are constructed based on the similarity and semantic relationships between modalities. Cosine similarity is used to calculate the similarity between nodes. Similar nodes are connected through edges. A heterogeneous graph is constructed, integrating different types of nodes and their edge relationships into a single graph.

[0036] Cross-modal information fusion: Use GCN to learn the node representations in the graph structure constructed by the graph structure, and simultaneously perform cross-modal information fusion, and map the features of different modalities into a shared embedding space. Use the embedding layer to fuse the information of different modalities, select a deep learning model, and perform joint training based on data from different modalities to achieve information fusion;

[0037] Graph optimization, processing, and application. After the graph is constructed, it is usually necessary to further optimize the graph, including optimizing the quality of the graph by removing low-correlation edges and nodes;

[0038] Graph smoothing technology is used in graph convolutional networks to reduce the impact of noise. At the same time, the graph can be simplified according to task requirements, removing unnecessary edges and nodes to improve processing efficiency.

[0039] The graph structure constructed by the cross-modal graph can be applied to a variety of tasks. It can find relevant matching values ​​from data of different modalities based on image and text queries.

[0040] Graph neural networks can be used to jointly learn data from different modalities to improve the performance of classification and detection tasks;

[0041] Model evaluation and tuning: Depending on the task, model performance is evaluated using precision, recall, and F1 score indicators.

[0042] In some embodiments, the cross-modal graph attention mechanism module needs to pre-process and extract features from data of different modalities before implementation. The specific process of the cross-modal graph attention mechanism module is as follows:

[0043] A cross-modal graph attention mechanism is constructed. Each node uses a self-attention mechanism to focus on its most relevant neighbor nodes and updates the node representation based on the attention weight.

[0044] Self-attention can determine the weight by calculating the similarity between nodes, and then weight the influence of neighboring nodes;

[0045] Nodes of different modalities transfer information through a cross-modal attention mechanism. Image nodes interact with text nodes and audio nodes through cross-modal attention. The attention mechanism adjusts the transfer weight between nodes of different modalities by calculating the similarity between them.

[0046] After assigning weights to each edge through the attention mechanism, the features of neighboring nodes are weighted and aggregated to obtain the updated representation of the current node;

[0047] Establish early warning status analysis, select early warning thresholds to match the characteristics of each node, and simulate data anomalies through deep learning algorithms to achieve early warning effects;

[0048] Model application and reasoning: Based on the cross-modal graph attention mechanism, users can perform cross-modal retrieval through image, text, or audio queries to find the most relevant content. At the same time, through cross-modal graph generation tasks, relevant images or videos can be generated based on given text, and vice versa.

[0049] In addition, the relationship between different modalities can be deeply analyzed, such as exploring the similarity or causal relationship between images, text and audio;

[0050] Performance optimization and expansion: By reducing redundant calculations in the model, the efficiency of graph neural networks is optimized and the inference speed is improved. For large-scale data sets, a distributed training framework can be used to accelerate the training process.

[0051] In some embodiments, the specific process of the semantic reasoning module is as follows:

[0052] Data processing and preparation: further processing of various data to ensure that the quality and format of input data meet the requirements of subsequent models;

[0053] Build a semantic reasoning model and select a model based on the task requirements. If the task is common sense reasoning, use the RERT bidirectional Transformer model; if the task is natural language production, use the GPT generative model;

[0054] Design an appropriate neural network structure and use a multi-layer Transformer to model the context and semantic information of the text;

[0055] In the feature extraction and reasoning process, a deep learning model is used to encode the input text, extracting contextual information and semantic relationships between sentences. A dedicated reasoning module is also designed to process implicit information in the text. If the task requires multiple rounds of reasoning, the model must remember and understand the context from multiple rounds to ensure consistent reasoning.

[0056] Train the model and implement evaluation. For classification tasks, use cross-entropy loss, and for generation tasks, use maximum likelihood estimation loss. When algorithm optimization is required, use Adam optimization and update model parameters through backpropagation. Adjust the training strategy according to the complexity of the task and use transfer learning to accelerate training.

[0057] Evaluation criteria include accuracy, precision, recall, F1 value and PLEU indicators;

[0058] Model deployment and application: Deploy the model as an API service to support online reasoning, integrate the semantic reasoning module with actual application scenarios, and optimize the model's response speed and processing capabilities;

[0059] Continuous optimization and feedback loops are used to fine-tune the model using new data to maintain continuous improvement in its reasoning capabilities; when processing complex semantics, the model's error types are analyzed to improve the model's reasoning capabilities.

[0060] In some embodiments, the specific process of the output result module is as follows:

[0061] Result formatting and processing: Before output, the model output needs to be cleaned and post-processed to convert the model output into a language that is easy for users to understand and remove unnecessary information;

[0062] Select the output format for the results. For text tasks, select plain text and JSON format; for image tasks, select PNG image file;

[0063] Error handling and exception management: verify the output to ensure it meets expectations. If the output is empty or incorrectly formatted, it needs to be processed or the user is prompted. When an exception occurs, the error log is recorded and retained for subsequent debugging and improvement.

[0064] Performance optimization: when processing large amounts of data, the output speed can be increased through parallel processing; for results of repeated calculations, the cache mechanism can be used to reduce repeated calculations and improve system response speed;

[0065] The diversification and customization of results allow users to choose different output methods and formats. The output module can also support multi-language translation and format adjustment. Depending on the task, the presentation of the output results needs to be adjusted.

[0066] Results presentation and interaction: Output results in the form of graphs, charts, or other visual formats. Ensure that the output is easy to understand and interactive, such as options for zooming in and out, and clicking to view details. Also, provide user feedback channels, allowing users to evaluate the output results or provide suggestions for improvement, thereby providing a basis for subsequent optimization.

[0067] Integration interface and deployment: For scenarios where external systems need to access output results, provide either a RESTful API or a GraphQL interface to ensure that the results can be delivered to other systems over the network.

[0068] Compared with the existing technology, the present invention provides a disaster information fusion and semantic reasoning algorithm based on the cross-modal graph attention mechanism, which has the following beneficial effects:

[0069] This disaster information fusion and semantic reasoning algorithm, based on a cross-modal graph attention mechanism, uses a data preprocessing module to preprocess various disaster data before entering a cross-modal graph construction module. Within this module, satellite imagery, meteorological data, social media content, and sensor data are constructed as nodes, with relationships between them represented by graph edges. The cross-modal graph attention mechanism weights the nodes, dynamically adjusting the weights of data from different modalities.

[0070] This invention uses a cross-modal graph attention mechanism to effectively integrate multimodal data sources, thereby providing more comprehensive disaster information. At the same time, through the graph attention mechanism, the model can dynamically and adaptively adjust the weights according to the relevance of different modal data, thereby improving the accuracy of information fusion. Using graph neural networks for semantic reasoning of disaster events can not only accurately predict the type and impact of disasters, but also provide decision support, helping disaster emergency management personnel to respond quickly;

[0071] The present invention can process a variety of different types of disaster data, has strong adaptability to data sources and formats, and is suitable for a variety of disaster emergency management scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a schematic diagram of the disaster information fusion and semantic reasoning algorithm module of the present invention;

[0073] Figure 2 This is a schematic diagram of the data preprocessing module of the present invention;

[0074] Figure 3 This is a schematic diagram of the cross-modal graph construction module of the present invention;

[0075] Figure 4 This is a schematic diagram of the cross-modal graph attention mechanism module of the present invention;

[0076] Figure 5 This is a schematic diagram of the semantic reasoning module of the present invention;

[0077] Figure 6 This is a schematic diagram of the output result module of the present invention. DETAILED DESCRIPTION

[0078] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention and the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0079] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are executed.

[0080] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0081] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0082] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.

[0083] See also Figure 1-6 In this implementation plan: disaster information fusion and semantic reasoning algorithm based on cross-modal graph attention mechanism, including:

[0084] The data preprocessing module is used to receive various disaster information from different data sources, including but not limited to image data, text data, and sensor data, and perform necessary preprocessing on such data to ensure the quality and consistency of the input data;

[0085] The specific process of the data preprocessing module is as follows:

[0086] Data cleaning process, first deal with missing data, choose to fill or delete rows and columns containing missing values, check for duplicate records in the data set, delete redundant data, and then use the standard deviation method to detect and remove outliers in the data;

[0087] Data normalization was performed by using the Z-score normalization method to convert numerical data into the same scale;

[0088] Lowercase words, remove stop words and punctuation marks, and perform stemming operations on text data;

[0089] For categorical data, use one-hot encoding to convert it into numerical data;

[0090] Data format conversion processing, using TF-IDF method to convert text data into vector representation;

[0091] For graphic data, resizing, denoising, and enhancement can be performed to increase the robustness of the model;

[0092] For time series data, timestamp unification, periodicity adjustment and smoothing are performed;

[0093] Feature selection and dimensionality reduction: Pearson correlation coefficient is used to select features that are closely related to the target variable, and principal component analysis is used to reduce the dimensionality of high-dimensional data to retain the main features in the data;

[0094] Data set partitioning and data integration: divide the data into K subsets, use K-1 subsets for training each time, and use the remaining subset for validation. Repeat K times to ensure the stability of the model.

[0095] Data balancing and preservation, reducing the data of majority class samples to achieve class balance;

[0096] Depending on the data size and subsequent processing requirements, you can choose one of the formats CSV, JSON, Parquet, and HDF5 to store the data. For large-scale data sets, use data version control tools to manage data versions for easy traceability and management.

[0097] Through the above process, the data preprocessing module can ensure the high quality and consistency of input data, providing a good data foundation for subsequent graph structure construction and disaster information fusion;

[0098] A cross-modal graph construction module, which converts preprocessed data into a graph structure where each modality is represented as a node and the relationships between modalities are represented by edges. Furthermore, the edges of the graph can be weighted based on the similarity between modalities.

[0099] The specific process of building a cross-modal graph module is as follows:

[0100] Data feature extraction: using deep learning models to extract high-level features of images;

[0101] Use the Word2Vec pre-trained language model to extract the semantic features of the text;

[0102] Extract the audio spectrum features and Mel-frequency cepstral coefficients and convert them into numerical vectors;

[0103] Use the combination of image features and audio features of video frames to extract the spatiotemporal features of the video;

[0104] Cross-modal alignment processing: For time series data, different modal data need to be aligned according to timestamps;

[0105] For image, text, and audio modalities, mapping is required to enable them to be connected together;

[0106] Graph structure construction requires connecting data nodes of different modalities and expressing the relationship between them, as follows:

[0107] The data of each modality is converted into nodes. Nodes can be feature vectors of images, text, and audio. Edges are constructed based on the similarity and semantic relationships between modalities. Cosine similarity is used to calculate the similarity between nodes. Similar nodes are connected through edges. A heterogeneous graph is constructed, integrating different types of nodes and their edge relationships into a single graph.

[0108] Cross-modal information fusion: Use GCN to learn the node representations in the graph structure constructed by the graph structure, and simultaneously perform cross-modal information fusion, and map the features of different modalities into a shared embedding space. Use the embedding layer to fuse the information of different modalities, select a deep learning model, and perform joint training based on data from different modalities to achieve information fusion;

[0109] Graph optimization, processing, and application. After the graph is constructed, it is usually necessary to further optimize the graph, including optimizing the quality of the graph by removing low-correlation edges and nodes;

[0110] Graph smoothing technology is used in graph convolutional networks to reduce the impact of noise. At the same time, the graph can be simplified according to task requirements, removing unnecessary edges and nodes to improve processing efficiency.

[0111] The graph structure constructed by the cross-modal graph can be applied to a variety of tasks. It can find relevant matching values ​​from data of different modalities based on image and text queries.

[0112] Graph neural networks can be used to jointly learn data from different modalities to improve the performance of classification and detection tasks;

[0113] Model evaluation and tuning: Depending on the task, use precision, recall, and F1 score metrics to evaluate model performance;

[0114] Through these processes, effective cross-modal graph construction can be achieved, facilitating the fusion and analysis of different modal data;

[0115] A cross-modal graph attention mechanism module, which applies the graph attention mechanism to the aforementioned graph structure. It dynamically adjusts the weight of modal data based on the correlation between nodes. By calculating the similarity between nodes, the attention mechanism can identify the key parts of different modal information and strengthen the focus on important data.

[0116] Before implementing the cross-modal graph attention mechanism module, it is necessary to preprocess and extract features from data of different modalities. The specific process of the cross-modal graph attention mechanism module is as follows:

[0117] A cross-modal graph attention mechanism is constructed. Each node uses a self-attention mechanism to focus on its most relevant neighbor nodes and updates the node representation based on the attention weight.

[0118] Self-attention can determine the weight by calculating the similarity between nodes, and then weight the influence of neighboring nodes;

[0119] Nodes of different modalities transfer information through a cross-modal attention mechanism. Image nodes interact with text nodes and audio nodes through cross-modal attention. The attention mechanism adjusts the transfer weight between nodes of different modalities by calculating the similarity between them.

[0120] After assigning weights to each edge through the attention mechanism, the features of neighboring nodes are weighted and aggregated to obtain the updated representation of the current node;

[0121] Establish early warning status analysis, select early warning thresholds to match the characteristics of each node, and simulate data anomalies through deep learning algorithms to achieve early warning effects;

[0122] Model application and reasoning: Based on the cross-modal graph attention mechanism, users can perform cross-modal retrieval through image, text, or audio queries to find the most relevant content. At the same time, through cross-modal graph generation tasks, relevant images or videos can be generated based on given text, and vice versa.

[0123] In addition, the relationship between different modalities can be deeply analyzed, such as exploring the similarity or causal relationship between images, text and audio;

[0124] Performance optimization and expansion: By reducing redundant calculations in the model, the efficiency of graph neural networks is optimized and the inference speed is improved. For large-scale data sets, a distributed training framework can be used to accelerate the training process.

[0125] Through the above steps, we can successfully implement the cross-modal graph attention mechanism, which helps the model better integrate multimodal information and perform efficient cross-modal task processing.

[0126] The semantic reasoning module uses graph neural networks to perform semantic reasoning based on the fused graph data to deduce the type, severity, and affected area of ​​the disaster event. This module is based on the content of the disaster-related database and can obtain specific disaster prediction results through reasoning;

[0127] The specific process of the semantic reasoning module is as follows:

[0128] Data processing and preparation: further processing of various data to ensure that the quality and format of input data meet the requirements of subsequent models;

[0129] Build a semantic reasoning model and select a model based on the task requirements. If the task is common sense reasoning, use the RERT bidirectional Transformer model; if the task is natural language production, use the GPT generative model;

[0130] Design an appropriate neural network structure and use a multi-layer Transformer to model the context and semantic information of the text;

[0131] In the feature extraction and reasoning process, a deep learning model is used to encode the input text, extracting contextual information and semantic relationships between sentences. A dedicated reasoning module is also designed to process implicit information in the text. If the task requires multiple rounds of reasoning, the model must remember and understand the context from multiple rounds to ensure consistent reasoning.

[0132] Train the model and implement evaluation. For classification tasks, use cross-entropy loss, and for generation tasks, use maximum likelihood estimation loss. When algorithm optimization is required, use Adam optimization and update model parameters through backpropagation. Adjust the training strategy according to the complexity of the task and use transfer learning to accelerate training.

[0133] Evaluation criteria include accuracy, precision, recall, F1 value and PLEU indicators;

[0134] Model deployment and application: Deploy the model as an API service to support online reasoning, integrate the semantic reasoning module with actual application scenarios, and optimize the model's response speed and processing capabilities;

[0135] Continuous optimization and feedback loops, using new data to fine-tune the model and maintain continuous improvement of its reasoning ability; analyzing the types of errors the model makes when processing complex semantics to improve the model's reasoning ability;

[0136] Through these processes, the semantic reasoning module can effectively understand and reason about semantic information in natural language, and is applied to various tasks such as question-answering systems, text reasoning, and dialogue generation;

[0137] The output result module converts the reasoning results into visual reports and decision support information, enabling disaster emergency responders to make timely decisions;

[0138] The specific process of the output result module is as follows:

[0139] Result formatting and processing: Before output, the model output needs to be cleaned and post-processed to convert the model output into a language that is easy for users to understand and remove unnecessary information;

[0140] Select the output format for the results. For text tasks, select plain text and JSON format; for image tasks, select PNG image file;

[0141] Error handling and exception management: verify the output to ensure it meets expectations. If the output is empty or incorrectly formatted, it needs to be processed or the user is prompted. When an exception occurs, the error log is recorded and retained for subsequent debugging and improvement.

[0142] Performance optimization: when processing large amounts of data, the output speed can be increased through parallel processing; for results of repeated calculations, the cache mechanism can be used to reduce repeated calculations and improve system response speed;

[0143] The diversification and customization of results allow users to choose different output methods and formats. The output module can also support multi-language translation and format adjustment. Depending on the task, the presentation of the output results needs to be adjusted.

[0144] Results presentation and interaction: Output results in the form of graphs, charts, or other visual formats. Ensure that the output is easy to understand and interactive, such as options for zooming in and out, and clicking to view details. Also, provide user feedback channels, allowing users to evaluate the output results or provide suggestions for improvement, thereby providing a basis for subsequent optimization.

[0145] Integration interface and deployment: For scenarios where external systems need to access output results, we provide either a RESTful API or a GraphQL interface to ensure that the results can be delivered to other systems over the network.

[0146] Through these processes, the output result module can provide output accurately and efficiently according to user needs, ensuring the format, content, visibility and interactivity of the results.

[0147] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are described briefly because they are generally similar to the method embodiments. For relevant parts, refer to the description of the method embodiments.

[0148] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. Disaster information fusion and semantic reasoning algorithm based on cross-modal graph attention mechanism, characterized by: include: The data preprocessing module is used to receive various disaster information from different data sources, including but not limited to image data, text data, and sensor data, and perform necessary preprocessing on such data to ensure the quality and consistency of the input data; A cross-modal graph construction module, which converts preprocessed data into a graph structure where each modality is represented as a node and the relationships between modalities are represented by edges. Furthermore, the edges of the graph can be weighted based on the similarity between modalities. A cross-modal graph attention mechanism module, which applies the graph attention mechanism to the aforementioned graph structure. It dynamically adjusts the weight of modal data based on the correlation between nodes. By calculating the similarity between nodes, the attention mechanism can identify the key parts of different modal information and strengthen the focus on important data. The semantic reasoning module uses graph neural networks to perform semantic reasoning based on the fused graph data to deduce the type, severity, and affected area of ​​the disaster event. This module is based on the content of the disaster-related database and can obtain specific disaster prediction results through reasoning; The output result module converts the reasoning results into visual reports and decision support information, enabling disaster emergency responders to make timely decisions.

2. The disaster information fusion and semantic reasoning algorithm based on the cross-modal graph attention mechanism according to claim 1 is characterized in that: The specific process of the data preprocessing module is as follows: Data cleaning process, first deal with missing data, choose to fill or delete rows and columns containing missing values, check for duplicate records in the data set, delete redundant data, and then use the standard deviation method to detect and remove outliers in the data; Data normalization was performed by using the Z-score normalization method to convert numerical data into the same scale; Lowercase words, remove stop words and punctuation marks, and perform stemming operations on text data; For categorical data, use one-hot encoding to convert it into numerical data; Data format conversion processing, using TF-IDF method to convert text data into vector representation; For graphic data, resizing, denoising, and enhancement can be performed to increase the robustness of the model; For time series data, timestamp unification, periodicity adjustment and smoothing are performed; Feature selection and dimensionality reduction: Pearson correlation coefficient is used to select features that are closely related to the target variable, and principal component analysis is used to reduce the dimensionality of high-dimensional data to retain the main features in the data; Data set partitioning and data integration: divide the data into K subsets, use K-1 subsets for training each time, and use the remaining subset for validation. Repeat K times to ensure the stability of the model. Data balancing and preservation, reducing the data of majority class samples to achieve class balance; Depending on the data size and subsequent processing requirements, you can choose one of the formats including CSV, JSON, Parquet and HDF5 to store the data. At the same time, for large-scale data sets, use data version control tools to manage data versions for easy traceability and management.

3. The disaster information fusion and semantic reasoning algorithm based on the cross-modal graph attention mechanism according to claim 1 is characterized in that: The specific process of the cross-modal graph construction module is as follows: Data feature extraction: using deep learning models to extract high-level features of images; Use the Word2Vec pre-trained language model to extract the semantic features of the text; Extract the audio spectrum features and Mel-frequency cepstral coefficients and convert them into numerical vectors; Use the combination of image features and audio features of video frames to extract the spatiotemporal features of the video; Cross-modal alignment processing: For time series data, different modal data need to be aligned according to timestamps; For image, text, and audio modalities, mapping is required to enable them to be connected together; Graph structure construction requires connecting data nodes of different modalities and expressing the relationship between them, as follows: The data of each modality is converted into nodes. Nodes can be feature vectors of images, text, and audio. Edges are constructed based on the similarity and semantic relationships between modalities. Cosine similarity is used to calculate the similarity between nodes. Similar nodes are connected through edges. A heterogeneous graph is constructed, integrating different types of nodes and their edge relationships into a single graph. Cross-modal information fusion: Use GCN to learn the node representations in the graph structure constructed by the graph structure, and simultaneously perform cross-modal information fusion, and map the features of different modalities into a shared embedding space. Use the embedding layer to fuse the information of different modalities, select a deep learning model, and perform joint training based on data from different modalities to achieve information fusion; Graph optimization, processing, and application. After the graph is constructed, it is usually necessary to further optimize the graph, including optimizing the quality of the graph by removing low-correlation edges and nodes; Graph smoothing technology is used in graph convolutional networks to reduce the impact of noise. At the same time, the graph can be simplified according to task requirements, removing unnecessary edges and nodes to improve processing efficiency. The graph structure constructed by the cross-modal graph can be applied to a variety of tasks. It can find relevant matching values ​​from data of different modalities based on image and text queries. Graph neural networks can be used to jointly learn data from different modalities to improve the performance of classification and detection tasks; Model evaluation and tuning: Depending on the task, model performance is evaluated using precision, recall, and F1 score indicators.

4. The disaster information fusion and semantic reasoning algorithm based on the cross-modal graph attention mechanism according to claim 1 is characterized in that: Before implementing the cross-modal graph attention mechanism module, it is necessary to preprocess and extract features from data of different modalities. The specific process of the cross-modal graph attention mechanism module is as follows: A cross-modal graph attention mechanism is constructed. Each node uses a self-attention mechanism to focus on its most relevant neighbor nodes and updates the node representation based on the attention weight. Self-attention can determine the weight by calculating the similarity between nodes, and then weight the influence of neighboring nodes; Nodes of different modalities transfer information through a cross-modal attention mechanism. Image nodes interact with text nodes and audio nodes through cross-modal attention. The attention mechanism adjusts the transfer weight between nodes of different modalities by calculating the similarity between them. After assigning weights to each edge through the attention mechanism, the features of neighboring nodes are weighted and aggregated to obtain the updated representation of the current node; Establish early warning status analysis, select early warning thresholds to match the characteristics of each node, and simulate data anomalies through deep learning algorithms to achieve early warning effects; Model application and reasoning: Based on the cross-modal graph attention mechanism, users can perform cross-modal retrieval through image, text, or audio queries to find the most relevant content. At the same time, through cross-modal graph generation tasks, relevant images or videos can be generated based on given text, and vice versa. In addition, we can conduct in-depth analysis of the relationship between different modalities, such as exploring the similarity or causal relationship between images, text, and audio; Performance optimization and expansion: By reducing redundant calculations in the model, the efficiency of graph neural networks is optimized and the inference speed is improved. For large-scale data sets, a distributed training framework can be used to accelerate the training process.

5. The disaster information fusion and semantic reasoning algorithm based on cross-modal graph attention mechanism according to claim 1 is characterized in that: The specific process of the semantic reasoning module is as follows: Data processing and preparation: further processing of various data to ensure that the quality and format of input data meet the requirements of subsequent models; Build a semantic reasoning model and select a model based on the task requirements. If the task is common sense reasoning, use the RERT bidirectional Transformer model; if the task is natural language production, use the GPT generative model; Design an appropriate neural network structure and use a multi-layer Transformer to model the context and semantic information of the text; In the feature extraction and reasoning process, a deep learning model is used to encode the input text, extracting contextual information and semantic relationships between sentences. A dedicated reasoning module is also designed to process implicit information in the text. If the task requires multiple rounds of reasoning, the model must remember and understand the context from multiple rounds to ensure consistent reasoning. Train the model and implement evaluation. For classification tasks, use cross-entropy loss, and for generation tasks, use maximum likelihood estimation loss. When algorithm optimization is required, use Adam optimization and update model parameters through backpropagation. Adjust the training strategy according to the complexity of the task and use transfer learning to accelerate training. Evaluation criteria include accuracy, precision, recall, F1 value and PLEU indicators; Model deployment and application: Deploy the model as an API service to support online reasoning, integrate the semantic reasoning module with actual application scenarios, and optimize the model's response speed and processing capabilities; Continuous optimization and feedback loops, using new data to fine-tune the model and maintain continuous improvement in its reasoning capabilities; Analyze the types of errors the model makes when processing complex semantics to improve the model's reasoning ability.

6. The disaster information fusion and semantic reasoning algorithm based on the cross-modal graph attention mechanism according to claim 1 is characterized by: The specific process of the output result module is as follows: Result formatting and processing: Before output, the model output needs to be cleaned and post-processed to convert the model output into a language that is easy for users to understand and remove unnecessary information; Select the output format for the results. For text tasks, select plain text and JSON format; for image tasks, select PNG image file; Error handling and exception management: verify the output to ensure it meets expectations. If the output is empty or incorrectly formatted, it needs to be processed or the user is prompted. When an exception occurs, the error log is recorded and retained for subsequent debugging and improvement. Performance optimization: when processing large amounts of data, the output speed can be increased through parallel processing; for results of repeated calculations, the cache mechanism can be used to reduce repeated calculations and improve system response speed; The diversification and customization of results allow users to choose different output methods and formats. The output module can also support multi-language translation and format adjustment. Depending on the task, the presentation of the output results needs to be adjusted. Results presentation and interaction: Output results in the form of graphs, charts, or other visual formats. Ensure that the output is easy to understand and interactive, such as options for zooming in and out, and clicking to view details. Also, provide user feedback channels, allowing users to evaluate the output results or provide suggestions for improvement, thereby providing a basis for subsequent optimization. Integration interface and deployment: For scenarios where external systems need to access output results, one of the RESTful API and GraphQL interface is provided to ensure that the results can be delivered to other systems over the network.

Citation Information

Cited By

  • Landslide disaster multi-element correlation analysis method, device and equipment and storage medium

    CN121434665A

  • Landslide disaster multi-element correlation analysis method, device and equipment and storage medium

    CN121434665B

  • Geological disaster intelligent analysis method and system based on large-scale language model

    CN121456331A

  • Parking control method, storage medium and vehicle

    CN121536283A