Typhoon disaster air-ground collaborative dynamic monitoring method and system based on deep learning

By employing a cross-modal attention mechanism and a deep learning model, the advantages of satellite remote sensing data and UAV remote sensing data are complemented, enabling dynamic tracking of typhoon disaster processes. This solves the problems of multi-source data fusion and quantification of prediction reliability, thereby improving the real-time performance and accuracy of typhoon disaster monitoring.

CN121661525APending Publication Date: 2026-03-13CHINA TOWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies lack a collaborative mechanism for multi-source heterogeneous data in typhoon disaster monitoring, which makes it difficult to deeply integrate the macroscopic view of satellites with the microscopic details of drones, making it difficult to capture the continuous dynamic evolution of the typhoon disaster process, and the prediction results of deep learning models lack reliable quantification.

Method used

By employing a cross-modal attention mechanism and a deep learning model, the advantages of satellite remote sensing data and UAV remote sensing data are complemented and feature fusion is performed. Furthermore, a 3D convolutional neural network and a Transformer encoder are used for continuous temporal analysis to simultaneously assess the uncertainty of the prediction results.

Benefits of technology

It achieves the complementary advantages of UAV and satellite data, dynamically tracks the typhoon disaster process, provides reliable and quantifiable prediction results, improves the real-time performance and prediction accuracy of disaster monitoring, and supports scientific emergency decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661525A_ABST
    Figure CN121661525A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of typhoon monitoring, and discloses a typhoon disaster air-ground collaborative dynamic monitoring method based on deep learning, and the method comprises the steps: obtaining original multi-modal data of a target region, the original multi-modal data at least comprising satellite remote sensing data and unmanned aerial vehicle remote sensing data, and carrying out the preprocessing of the satellite remote sensing data and the unmanned aerial vehicle remote sensing data, and performing feature extraction on satellite remote sensing data and unmanned aerial vehicle remote sensing data after space registration, performing feature fusion on satellite image features and unmanned aerial vehicle image features based on a cross-modal attention mechanism, and analyzing continuous time sequence fusion features during a typhoon transit period to obtain a disaster evolution prediction result. According to the method, pixel-level space registration is realized through feature point matching and affine transformation, and adaptive weighting and deep fusion are carried out on data from different platforms on a feature level by using a cross-modal attention mechanism; therefore, the macroscopic meteorological environment information contained in the satellite image can effectively guide and correct the microscopic detail analysis of the unmanned aerial vehicle image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of typhoon monitoring, and specifically relates to a deep learning-based method and system for dynamic monitoring of typhoon disasters via air-ground collaboration. Background Technology

[0002] Typhoons, as one of the most destructive natural disasters globally, pose a serious threat to people's lives and property due to storm surges, heavy rainfall, floods, and secondary disasters they trigger. Traditional disaster monitoring mainly relies on meteorological satellites, radar, and ground observation station networks. These methods can provide large-scale dynamic meteorological information, providing crucial evidence for early warning. With the rapid development of remote sensing technology and artificial intelligence, deep learning-based computer vision methods have been introduced into the field of disaster monitoring, significantly improving the automation level of information extraction. For example, using satellite remote sensing imagery for automatic identification of large-scale inundation areas, or using high-resolution imagery acquired by drones for building damage detection, has become a current research and application hotspot. In addition, existing technologies have also seen systematic attempts to integrate multi-dimensional data, such as the three-dimensional simulation system for typhoon disaster monitoring and assessment disclosed in authorization announcement number CN8170714A. This system integrates multi-source data and displays and evaluates it on a three-dimensional geographic information platform, aiming to provide more intuitive decision support for disaster prevention command.

[0003] However, when dealing with the rapidly evolving and complex disaster of typhoons, most of the existing technologies still process data from single platforms or single types, such as satellites and drones, in isolation. They lack effective mechanisms for multi-source heterogeneous data collaboration, resulting in a lack of deep integration between the macroscopic view from satellites and the microscopic details from drones, and failing to fully leverage the complementary advantages of information. Secondly, most existing analyses focus on static images at specific time phases, making it difficult to capture the continuous dynamic evolution of typhoon disaster processes from occurrence and development to decline, thus limiting the ability to track disasters in real time and predict trends. Finally, mainstream deep learning models typically output a definitive classification or segmentation result, but cannot provide a measure of the uncertainty of the prediction itself, while clearly knowing the reliability of the judgment is crucial for risk assessment and emergency decision-making.

[0004] Therefore, how to achieve intelligent collaboration of multi-source data, dynamic analysis throughout the entire process, and credibility assessment of prediction results are key technical issues that urgently need to be addressed to improve the effectiveness of typhoon disaster monitoring and early warning systems. Summary of the Invention

[0005] To address the aforementioned issues, this application provides a deep learning-based method for joint air-ground dynamic monitoring of typhoon disasters, which has the advantage of leveraging the complementary strengths of UAV and satellite remote sensing data.

[0006] This application provides a deep learning-based method for air-ground coordinated dynamic monitoring of typhoon disasters, comprising the following steps: Acquire raw multimodal data of the target area, which includes at least satellite remote sensing data and UAV remote sensing data; Preprocessing of satellite remote sensing data and UAV remote sensing data, and spatial registration of the preprocessed satellite remote sensing data and UAV remote sensing data; Feature extraction is performed on spatially registered satellite remote sensing data and UAV remote sensing data to obtain satellite image features and UAV image features. Based on a cross-modal attention mechanism, the satellite image features and UAV image features are fused to obtain fused features. Based on the fusion characteristics, the continuous temporal fusion characteristics during the passage of typhoons are analyzed to obtain disaster evolution prediction results, and the uncertainty of the prediction results is evaluated simultaneously to obtain uncertainty information of the prediction results. Disaster risk assessment is conducted based on disaster evolution prediction results and uncertainty assessment results, combined with external geographic information; The output includes monitoring information that includes the assessment results of disaster risks and the corresponding uncertainty prediction results.

[0007] Furthermore, the preprocessing includes: Preprocessing includes radiometric and geometric correction to transform satellite and UAV remote sensing data into standardized image data. Radiometric correction includes radiometric calibration and atmospheric correction.

[0008] Furthermore, preprocessing also includes: Based on their correlation with the target disaster identification task, feature bands are selected from the acquired satellite remote sensing data and UAV remote sensing data to filter out the data subset with the largest amount of information. The selection of characteristic bands is based on the principle of maximizing mutual information.

[0009] Furthermore, it also includes: A first deep learning model is adopted, which is used to achieve feature fusion of satellite image features and UAV image features based on a cross-modal attention mechanism; A second deep learning model is adopted, which is used to analyze the continuous temporal fusion characteristics during the passage of typhoons based on the time series analysis architecture, obtain disaster evolution prediction results, and simultaneously assess the uncertainty of the prediction results.

[0010] Furthermore, the second deep learning model includes a 3D convolutional neural network for extracting local spatiotemporal features, and a self-attention encoder for modeling global dependencies.

[0011] Furthermore, the uncertainty of this prediction result is assessed simultaneously, including: During the inference process of the second deep learning model, stochastic forward inference is performed to generate a set of prediction results; Calculate the statistical mean of the set of prediction results as the final prediction value; Calculate the statistical variance of the predicted result set as the uncertainty information corresponding to the final predicted value.

[0012] Furthermore, feature fusion based on cross-modal attention mechanisms for satellite image features and UAV image features includes: Satellite image features are mapped to key vectors and value vectors, and UAV image features are mapped to query vectors; Calculate the attention weights between the query vector and the key vector; The value vector is weighted according to the attention weight, and the weighted result is fused with the UAV image features to obtain fused features.

[0013] Furthermore, feature fusion is performed on satellite image features and UAV image features through a feature pyramid network.

[0014] Furthermore, disaster risk assessment includes: The disaster evolution prediction results are overlaid with external digital elevation model data to calculate the actual inundation depth distribution; Based on building outline data, the average flood depth within each building unit is calculated, and its risk level is determined according to a preset threshold.

[0015] Furthermore, the method also includes: Collect user feedback data on monitoring information; Based on the feedback data and the corresponding original multimodal data, the models involved in feature fusion and time series analysis are optimized and updated.

[0016] This application also provides a deep learning-based air-ground collaborative dynamic monitoring system for typhoon disasters, the system comprising: The data acquisition module is used to simultaneously acquire satellite remote sensing data and UAV remote sensing data of the target area; The data preprocessing and registration module is used for preprocessing and spatial registration of satellite remote sensing data and UAV remote sensing data; The deep fusion analysis module is used to perform feature fusion on the registered data based on a cross-modal attention mechanism to obtain fused features. The time-series dynamic analysis module is used to process the continuous time-series fusion features during the passage of typhoons based on fusion features and using a time-series analysis model to obtain disaster evolution prediction results and simultaneously assess the uncertainty of the prediction results. The disaster risk assessment module is used to conduct disaster risk assessments based on disaster evolution prediction results and their uncertainties, combined with external geographic information. The information service and output module is used to generate and output monitoring information that includes disaster risk assessment results and their corresponding uncertainty information.

[0017] This application also provides an electronic device, which includes at least one processor and at least one memory, the memory being data-connected to the processor, wherein the memory stores instructions executable by at least one processor, the instructions being executed by at least one processor to enable at least one processor to perform the methods described above.

[0018] This application also provides a computer-storeable medium storing computer instructions, which, when executed by a processor, specifically perform the steps in the above-described method.

[0019] This application also provides a computer program product, including computer instructions, which, when executed by a processor, specifically perform the steps in the above-described method.

[0020] Compared with the prior art, this application has the following advantages: This invention achieves pixel-level spatial registration through feature point matching and affine transformation, and utilizes a cross-modal attention mechanism to adaptively weight and deeply fuse data from different platforms at the feature level. This enables the macro-meteorological environment information contained in satellite imagery to effectively guide and correct the micro-detail analysis of UAV imagery. At the same time, local high-precision observations can also provide feedback and refine macro-situation judgments, achieving complementary advantages between UAV and satellite remote sensing data, generating a comprehensive disaster situation map with both breadth and accuracy, and overcoming the inherent limitations of a single data source in terms of coverage or resolution.

[0021] To address the problem that existing static analysis methods struggle to capture the dynamic processes of disasters, this invention constructs a spatiotemporal cube from the fused features of continuous time phases and employs an architecture combining a 3D convolutional neural network and a Transformer encoder. This architecture can extract local evolution patterns in short time series and model global dependencies in long time series. By using causal masks to ensure autoregressive predictions over time series, the system can dynamically track processes such as the expansion of inundation zones and the evolution of damage to disaster-bearing bodies, and predict their future short-term trends. This achieves a leap from monitoring to predicting typhoon disasters, providing a crucial time window for forward-looking disaster prevention deployment.

[0022] To address the lack of reliability metrics in existing model predictions, this invention outputs a quantified uncertainty score along with the prediction conclusion, explicitly informing decision-makers of the reliability of the judgment. This design transforms AI predictions into a risk-assessable decision-making reference, enabling emergency command to distinguish between high-confidence areas and areas requiring verification. This allows for the rational allocation of limited rescue resources and more prudent and scientific decision-making, significantly enhancing the system's practical value in complex and uncertain disaster environments.

[0023] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart illustrating an embodiment of this application is shown; Figure 2 A system framework diagram according to an embodiment of this application is shown; Figure 3 A flowchart according to an embodiment of this application is shown. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] This embodiment details a deep learning-based method for coordinated air-ground dynamic monitoring of typhoon disasters. For example... Figure 1 and Figure 3 As shown, the method includes the following steps executed in an ordered manner: The first stage is data preparation and preprocessing.

[0028] S1, collaboratively acquires multi-source remote sensing data.

[0029] After a typhoon warning is issued, the system simultaneously plans satellite transit observation missions and UAV emergency patrol routes. Satellite observation focuses on acquiring large-scale image data, preferably including infrared bands sensitive to cloud top temperature, water vapor channel bands sensitive to atmospheric water vapor content, and shortwave infrared bands effective for identifying surface water bodies. UAV observation targets key areas that may be affected, such as urban low-lying areas and areas around important infrastructure, acquiring image data with significantly higher spatial resolution than satellites, including at least visible light (RGB) bands and near-infrared bands with obvious spectral response characteristics to water bodies / vegetation. Through this collaborative observation strategy, information on both macro-meteorological conditions and micro-surface details is obtained simultaneously.

[0030] Specifically, satellite data includes multi-band data such as infrared, water vapor channel, and visible light, while UAV data includes RGB, multispectral, and thermal infrared data, to achieve all-weather, multi-level disaster monitoring capabilities. Let the set of satellite bands be:

[0031] The UAV band set is as follows:

[0032] Based on typhoon monitoring requirements, select the optimal band combination:

[0033] Where I represents mutual information and Y represents disaster label.

[0034] S2, Data Preprocessing and Precise Registration.

[0035] Because satellite and UAV data differ in sensor characteristics, imaging geometry, and resolution, standardization is necessary for subsequent fusion. Therefore, radiometric calibration and atmospheric correction are performed on the raw satellite and UAV images separately, converting sensor values ​​into physical quantities such as surface reflectance or radiance to eliminate the effects of atmospheric scattering and absorption. Next, geometric correction is performed, using ground control points or sensor attitude parameters to eliminate image distortion caused by terrain undulations and sensor viewing angles, giving the images accurate geographic coordinates.

[0036] Subsequently, spatial registration is performed. On the geometrically corrected satellite reference image and UAV image, a large number of local feature points are extracted using the Scale Invariant Feature Transform (SIFT) algorithm. The feature descriptor matching algorithm is used to find corresponding feature point pairs in the two images. Then, based on these successfully matched point pairs, the Random Sample Consensus (RANSAC) algorithm is used to robustly estimate an affine transformation model that can describe the geometric relationships such as rotation, scaling, and translation between the two images.

[0037] Finally, this model is used to resample the UAV imagery so that each pixel is precisely aligned to the coordinate grid of the satellite imagery, achieving pixel-level spatial registration. This ensures that data from different platforms can be effectively fused under the same spatial reference.

[0038] The second stage is model building and computation.

[0039] S3, cross-modal deep fusion analysis.

[0040] This step is the core of solving the problem of multi-source data collaboration, and its specific implementation includes the following steps: S3-1. Input the registered satellite and UAV images into two independent convolutional neural networks (e.g., ResNet50) for feature extraction. The network will output feature maps at multiple levels, which form a feature pyramid. High-level feature maps have low resolution but contain rich semantic information (e.g., "water" and "building" areas); low-level feature maps have high resolution and retain more texture and detail information (e.g., building edges and wave patterns).

[0041] S3-2. For satellite feature maps (denoted as Fs) and UAV feature maps (denoted as Fu) from the same semantic level, they are transformed into query, key, and value vectors, denoted as Qs, Ks, Vs and Qu, Ku, Vu, respectively, through three different linear transformation layers. The cross-modal attention is as follows:

[0042] in, This is a feature map of satellite data, where R is the spatial resolution ratio of the image, a scaling factor, and H and W are the height and width of the image, respectively. The number of channels for satellite data depends on the type of satellite sensor. This is a feature map of the UAV data, where R, H, and W represent the image height and width. Let Q be the number of channels for the drone data, Q be the current location that needs attention calculation (to query information from other locations), K be the identifiers provided by all locations, used to match with Q to calculate attention weights, and V be the substantive information carried by all locations, which is finally aggregated based on the attention weights. , Generated from linear layers of projected satellite features. , , The generation method is the same as above, but a linear layer specific to the UAV mode is used.

[0043] Preferably, the fusion strategy for the feature pyramid network to fuse features of different scales from satellite and UAV data is as follows: ) Where: l represents the level of the pyramid; UP stands for upsampling operation; Concat is used for concatenating channels; Conv1 1 represents a 1x1 convolution used for dimensionality reduction.

[0044] The pyramid's structure corresponds to its hierarchical levels: The UAV feature Qu can query and weighted aggregate macroscopic environmental information from satellite features Vs, enabling satellite information to guide and correct the microscopic analysis of UAVs.

[0045] The fusion process is bidirectional, using UAV features as the query source (Qu) to query satellite features (Ks). By calculating the similarity between all locations in Qu and Ks (typically a dot product operation) and normalizing it using the Softmax function, a set of attention weights is obtained. These weights quantify which locations in the global context of the satellite imagery need to be extracted, and how much, when understanding each location in the current local scene of the UAV. Then, these weights are used to perform a weighted summation of the satellite value vector (Vs) to obtain a new feature modulated or enhanced by global satellite information. This new feature is then fused with the original UAV features.

[0046] Similarly, satellite features can be used as the query source, and drone features can be modulated. Through this cross-attention mechanism, the macroscopic cloud and rain information from the satellite can help understand the cause of water accumulation in the drone footage, while the clear details from the drone can help correct the satellite's coarse-grained judgment of the disaster area.

[0047] Finally, the features fused from each layer are integrated through a feature pyramid network to output a unified, highly complementary multi-scale fused feature map.

[0048] Phase 3: Post-processing and visualization of results.

[0049] S4, time series dynamic modeling and prediction.

[0050] This step aims to analyze the continuous temporal fusion characteristics during the passage of a typhoon based on the aforementioned fusion features, obtain disaster evolution prediction results, and thereby capture the dynamic evolution patterns of the disaster. S4-1. During the typhoon's passage, repeat steps S1 to S3 at fixed time intervals (e.g., every 3 hours) to obtain a series of fused feature maps at time points t1, t2, ..., tn. Stack these feature maps in chronological order to form a three-dimensional spatiotemporal cube (two spatial dimensions and one temporal dimension).

[0051] S4-2. Input the spatiotemporal cube into a three-dimensional convolutional neural network. The 3D convolutional kernel slides and convolves simultaneously in the spatial and temporal dimensions, which can effectively capture the local change patterns of disasters in space within a short time window. For example, it can identify the trend of floods spreading from east to west in two consecutive images.

[0052] S4-3, Modeling Long-Term Dependencies and Prediction: The feature sequence extracted by the 3DCNN is input into a Transformer encoder. The Transformer's self-attention mechanism can calculate the correlation between features at any two time points in the sequence, thereby learning the long-term patterns of disaster evolution throughout the typhoon process. To prevent the model from peeking at future answers during training, a causal mask is applied in the self-attention calculation, allowing only current-time features to focus on past-time features. This setting enables the model to be trained and predicted in an autoregressive manner: predicting the state of the next moment based on the historical sequence, then adding the predicted value back to the sequence to continue predicting the more distant future. In this way, the system can not only describe the historical evolution "trajectory" of the disaster but also predict its short-term development trend.

[0053] Specifically, the architecture combining a 3D convolutional neural network (the first deep learning model) with a Transformer encoder has the following processing flow: continuous time steps fused feature sequences Constructed as a spacetime cube; First, a 3D CNN is used to extract short-term local spatiotemporal features, which are then input into a Transformer encoder. Its self-attention mechanism is used to model long-term global dependencies. The self-attention calculation is as follows:

[0054] M is a lower triangular mask matrix that ensures that the current moment can only focus on historical moments, enabling autoregressive modeling of dynamic processes such as typhoon path and flood zone expansion.

[0055] 3DCNN adopts the C3D architecture, the number of Transformer encoder layers is configurable, it supports variable-length temporal inputs, and a course learning strategy is used during training to gradually increase the temporal length. Spatiotemporal feature extraction:

[0056]

[0057] Where M is the causal mask matrix.

[0058] Step S5: Quantify the uncertainty of the prediction.

[0059] This step involves simultaneously assessing the uncertainty of the prediction result and obtaining uncertainty information. The aim is to assign a confidence label to the prediction result. In the deep learning model constructed in steps S3 and S4, these Dropout layers are not turned off when the model is used for prediction; instead, they remain active. For the same input data to be analyzed (e.g., the most recently acquired image), the system performs T (e.g., T=50) forward propagation calculations. Because Dropout randomly "masks" a portion of neurons in the network each time, each forward propagation will yield slightly different prediction outputs (e.g., the probability value of a pixel being classified as "flooded").

[0060] Specifically, Monte Carlo Dropout is used as a Bayesian approximation in the key fully connected or convolutional layers of the first and second deep learning models. For an input x, perform N random forward propagations to obtain N predicted outputs. ; The final prediction result is the mean. ; The uncertainty of a forecast is quantified by the variance of the forecast values: , As not The determination score is output along with the prediction result.

[0061] The system collects these T prediction results and calculates their arithmetic mean. This mean is used as the final prediction value for the pixel. At the same time, the variance of these T results is calculated and used as the uncertainty score of this prediction. The larger the variance, the more inconsistent the model's internal judgment is, and the more uncertain the prediction result is. The smaller the variance, the more consistent the model's judgments are and the higher the reliability of the prediction results. This score will be output along with the final classification or regression results.

[0062] Step S6: Disaster simulation and refined assessment.

[0063] This step, based on the aforementioned data-driven analysis, uses the disaster evolution prediction results and their uncertainty information, combined with external geographic information, to conduct a disaster risk assessment.

[0064] Step S6-1: For tasks such as urban flooding prediction, construct a physical information neural network (second deep learning model).

[0065] Specifically, for urban flooding prediction, the loss function L is defined as: L= +λ

[0066] in: To predict water depth To observe water depth The mean square error is based on the shallow water equation residuals. The physical constraint terms are calculated using the following formula:

[0067] Where: h is the water depth, υ is the flow velocity vector, R is the precipitation, and λ is the tradeoff hyperparameter. By minimizing this loss function, the network's predictions are made to conform to physical laws.

[0068] During training, the network's loss function includes not only the error between the predicted and observed values, but also an additional physical constraint loss, which is calculated based on the fundamental equations of fluid mechanics (such as the Saint-Venant equations describing shallow water flow).

[0069] Specifically, variables such as water depth and flow velocity predicted by the network are substituted into the equation, and the residuals on both sides of the equation are calculated. By optimizing to minimize the residuals, the prediction results of the neural network are forced to not only match the training data, but also be reasonable in terms of physical laws, thereby enhancing the model's generalization and prediction capabilities in extreme or unseen scenarios.

[0070] Step S6-2: The system connects to an external geographic information system database.

[0071] First, the water depth distribution map predicted by the network is overlaid with the high-precision digital elevation model (DEM, i.e., terrain height data) Z(x,y) and the predicted water depth distribution h(x,y,t) to generate the inundation water depth map D(x,y,t)=h(x,y,t)-(Z(x,y)-Zbase), where Zbase is the baseline water level elevation; Then, a vector layer containing the building outline is loaded. For each building polygon, the average predicted flood depth of all raster cells within it is calculated. If the average depth exceeds a preset disaster threshold (e.g., 0.3 meters), the building is automatically identified as a "high-risk building".

[0072] This assessment method combines macro-level disaster prediction with micro-level vulnerability analysis of disaster-bearing bodies, achieving refined and spatially precise risk warning.

[0073] The formula for calculating the average submerged depth is as follows:

[0074] The criteria for determining the risk of flooding are as follows:

[0075] in, Average flood depth within building units Let be the inundation depth at the i-th discrete point within the building outline. The preset flood risk water depth threshold is used, and R is a variable indicating the building's flood risk status. R=1 indicates a building at risk of flooding, and R=0 indicates a building at no risk of flooding. Step S7: Information integration and dynamic visualization.

[0076] Specifically, all generated monitoring products, forecast results, risk assessment maps, and corresponding uncertainty scores are assigned geographic coordinates and stored in a spatial database in the cloud.

[0077] The system provides a web client access interface and adopts progressive detail rendering technology: when users view the macro situation at the provincial or municipal level, the client loads and displays an overview map generated primarily from satellite data with fast rendering speed; when users zoom in to the district or street level, the client automatically requests and loads the corresponding high-resolution detail map and refined prediction results generated from UAV data for that area from the server.

[0078] Users can click on any location on the map to check the latest disaster status, forecast trends, risk level, and credibility indicators provided by the system for that location.

[0079] Step S8, model optimization based on feedback.

[0080] To enable the system to learn continuously, a feedback optimization loop was established. The visualization client provides a convenient feedback entry point, allowing frontline emergency responders or experts to annotate the system's prediction results, such as noting that the prediction was accurate here or that the event did not actually occur here.

[0081] The system backend automatically associates and stores this feedback information, the original multi-source data that triggered the prediction, the model's initial prediction value, and its uncertainty score as a new sample, adding it to a continuously growing validation dataset. The system periodically starts offline training tasks, using this validation dataset to fine-tune the core deep learning model from steps S3, S4, and S6. This allows the system's analytical and predictive capabilities to continuously iterate and optimize as practical applications deepen and feedback accumulates, thereby maintaining and improving its monitoring performance over the long term.

[0082] This embodiment provides a monitoring system for implementing the method described above. The system logically and physically includes functional modules corresponding to the steps of the method in Embodiment 1. Each module can be implemented in a server cluster via software programming and communicates and collaborates through a predefined data interface protocol.

[0083] Specifically, the system includes: Multimodal data acquisition module: It is configured to perform step S1 of embodiment one, has an interface for communicating with satellite ground station and UAV control station, and is responsible for receiving and caching raw remote sensing image data.

[0084] Data preprocessing and registration module: It is configured to execute step S2, which integrates radiometric correction, geometric correction algorithms and feature matching registration algorithms to produce spatially aligned standard data products.

[0085] Deep Fusion Analysis Module: It is configured to execute step S3. This module encapsulates the neural network model structure and forward computation logic required to implement the cross-modal attention mechanism and feature pyramid network.

[0086] Temporal dynamic analysis module: It is configured to execute step S4. This module has a built-in hybrid model of 3DCNN and Transformer encoder, which is responsible for loading spatiotemporal sequence data and outputting dynamic analysis and prediction results.

[0087] Uncertainty Quantification Module: It is configured to execute step S5. This module controls the deep learning model to enable Monte Carlo Dropout mode during inference and is responsible for statistically analyzing the results of multiple inferences, calculating the mean and variance.

[0088] Disaster Prediction and Simulation Module: It is configured to execute step S6. This module integrates a physical information neural network and provides a standard interface with external GIS databases (such as OGCWFS) to acquire data such as DEM and building outlines for detailed assessment.

[0089] Information Service and Visualization Module: It is configured to execute step S7. This module includes a cloud server (responsible for business logic and API), a spatial database (such as PostGIS, responsible for data storage), and a web front-end application (responsible for interaction and rendering), forming a complete information publishing system.

[0090] Feedback optimization loop module: As the data pipeline and control logic connecting the information service module and the aforementioned analysis modules, it is configured to execute step S8, which is responsible for collecting and storing feedback data, and scheduling and triggering the periodic fine-tuning tasks of the model.

[0091] During system operation, the data flow and method description steps are completely consistent: raw data is input from the data acquisition module, flows sequentially through modules such as preprocessing, fusion, time series analysis, uncertainty quantification, and simulation evaluation for processing, and the final results are published in the information service module. Simultaneously, user feedback flows back to the analysis model through the feedback optimization loop module, forming a closed loop. Each module performs its specific function, collectively achieving intelligent, dynamic, and interpretable monitoring of typhoon disasters.

[0092] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A deep learning-based method for dynamic monitoring of typhoon disasters via air-ground coordination, characterized in that, Includes the following steps: Acquire raw multimodal data of the target area, wherein the raw multimodal data includes at least satellite remote sensing data and UAV remote sensing data; The satellite remote sensing data and UAV remote sensing data are preprocessed, and the preprocessed satellite remote sensing data and UAV remote sensing data are spatially registered. Feature extraction is performed on the spatially registered satellite remote sensing data and the UAV remote sensing data to obtain satellite image features and UAV image features. Based on the cross-modal attention mechanism, the satellite image features and UAV image features are fused to obtain fused features. Based on the aforementioned fusion characteristics, the continuous temporal fusion characteristics during the typhoon's passage are analyzed to obtain disaster evolution prediction results, and the uncertainty of the prediction results is assessed simultaneously to obtain uncertainty information of the prediction results. Based on the disaster evolution prediction results and their uncertainty information, and in conjunction with external geographic information, a disaster risk assessment is conducted; The output includes monitoring information containing the assessment results of the disaster risk and its corresponding uncertainty prediction results.

2. The method according to claim 1, characterized in that, The preprocessing includes radiometric correction and geometric correction to convert satellite remote sensing data and UAV remote sensing data into standardized image data. The radiometric correction includes radiometric calibration and atmospheric correction.

3. The method according to claim 1, characterized in that, The preprocessing also includes: Based on their correlation with the target disaster identification task, feature bands are selected from the acquired satellite remote sensing data and UAV remote sensing data to filter out the data subset with the largest amount of information. The selection of the characteristic bands is performed based on the mutual information maximization criterion.

4. The method according to claim 1, characterized in that, Also includes: The first deep learning model is adopted, which can realize feature fusion of satellite image features and UAV image features based on cross-modal attention mechanism; A second deep learning model is adopted, which can analyze the continuous temporal fusion characteristics during the passage of typhoons based on the temporal analysis architecture, obtain disaster evolution prediction results, and simultaneously assess the uncertainty of the prediction results.

5. The method according to claim 4, characterized in that, The second deep learning model includes a three-dimensional convolutional neural network for extracting local spatiotemporal features and a self-attention encoder for modeling global dependencies.

6. The method according to claim 5, characterized in that, The uncertainty of the simultaneous assessment of the prediction result includes: During the inference process of the second deep learning model, stochastic forward inference is performed to generate a set of prediction results; Calculate the statistical mean of the predicted result set as the final predicted value; Calculate the statistical variance of the predicted result set as the uncertainty information corresponding to the final predicted value.

7. The method according to claim 1, characterized in that, The feature fusion of satellite image features and UAV image features based on the cross-modal attention mechanism includes: The satellite image features are mapped into key vectors and value vectors, and the UAV image features are mapped into query vectors; Calculate the attention weight between the query vector and the key vector; The value vector is weighted according to the attention weight, and the weighted result is fused with the UAV image features to obtain the fused features.

8. The method according to claim 1, characterized in that, The satellite image features and UAV image features are fused using a feature pyramid network.

9. The method according to claim 1, characterized in that, The disaster risk assessment includes: The disaster evolution prediction results are overlaid with external digital elevation model data to calculate the actual inundation depth distribution; Based on building outline data, the average flood depth within each building unit is calculated, and its risk level is determined according to a preset threshold.

10. The method according to claim 1, characterized in that, The method further includes: Collect user feedback data on the monitoring information; Based on the feedback data and the corresponding original multimodal data, the model involved in the feature fusion and time series analysis is optimized and updated.

11. A deep learning-based air-ground collaborative dynamic monitoring system for typhoon disasters, characterized in that, The system includes: The data acquisition module is used to simultaneously acquire satellite remote sensing data and UAV remote sensing data of the target area; The data preprocessing and registration module is used to preprocess and spatially register the satellite remote sensing data and the UAV remote sensing data; The deep fusion analysis module is used to perform feature fusion on the registered data based on a cross-modal attention mechanism to obtain fused features. The time-series dynamic analysis module is used to process the continuous time-series fusion features during the typhoon's passage using a time-series analysis model based on the fusion features, obtain disaster evolution prediction results, and simultaneously assess the uncertainty of the prediction results. The disaster risk assessment module is used to conduct disaster risk assessment based on the disaster evolution prediction results and their uncertainties, and in conjunction with external geographic information; The information service and output module is used to generate and output monitoring information containing the disaster risk assessment results and their corresponding uncertainty information.

12. An electronic device, characterized in that, The electronic device includes at least one processor and at least one memory, the memory being data-connected to the processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A computer-storable medium, characterized in that, The storable medium stores computer instructions, which, when executed by a processor, specifically perform the steps of the method as described in any one of claims 1-10.

14. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they specifically perform the steps in the method as described in any one of claims 1-10.