Space-ground-terrestrial integrated geological disaster identification method and system based on multi-source fusion perception

By using generative diffusion models and generative adversarial networks for satellite image processing, combined with multi-source feature fusion of ViT and U-net models and dynamic path planning of UAVs, the data fusion and collaborative response problems of the integrated air-space-ground geological disaster monitoring system were solved, achieving efficient and accurate disaster identification and real-time response.

CN121280928BActive Publication Date: 2026-04-17安徽明生恒卓科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
安徽明生恒卓科技有限公司
Filing Date
2025-09-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing integrated air-space-ground geological disaster monitoring systems are inadequate in terms of high-precision spatiotemporal fusion of cross-domain multi-source heterogeneous data, fine-grained feature modeling, and dynamic collaborative response capabilities, making it difficult to achieve efficient and accurate disaster identification and response.

Method used

Generative diffusion models and generative adversarial networks are used for temporal interpolation and super-resolution reconstruction of satellite images. By combining multimodal feature fusion of ViT networks and fine-grained recognition of U-net models, along with UAV dynamic path planning and data supplementation, a hierarchical closed-loop collaborative perception system is constructed.

Benefits of technology

It achieves spatiotemporal alignment and resolution consistency of multi-source satellite data, improves the accuracy of disaster area identification and monitoring efficiency, meets the real-time and accuracy requirements of disaster monitoring, and enhances the defense capability under extreme disasters, especially in the monitoring of power facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280928B_ABST
    Figure CN121280928B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of artificial intelligence, and particularly relates to a space-air-ground integrated geological disaster identification method and system based on multi-source fusion perception. The scheme first acquires satellite images of a target area collected by stationary orbit satellites, synthetic aperture radar satellites and polar orbit satellites in real time; then performs time interpolation and spatial super-resolution reconstruction on multi-source heterogeneous satellite image data, and further obtains three types of satellite images with consistent time and space and unified resolution; the three types of satellite images are synchronously input into a pre-trained coarse-grained disaster identification model to obtain a macro probability graph reflecting geological disaster risk; after identifying the local area with high risk, ground monitoring equipment and unmanned aerial vehicles are used to obtain complete high-resolution images; finally, the high-resolution images are input into a pre-trained fine-grained disaster identification model to obtain the identification result of the geological disaster. The present application overcomes the problems of existing geological disaster monitoring schemes in real-time and precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence, specifically relating to a multi-source fusion perception integrated air-space-ground geological disaster identification method, system, and corresponding computer program product. Background Technology

[0002] Currently, natural disaster monitoring technology has gradually moved from static observation based on a single data source to an intelligent monitoring stage characterized by multi-platform collaboration and multi-modal fusion. Current intelligent monitoring based on multi-modal fusion mainly includes two categories: monitoring schemes based on multi-source satellite remote sensing data fusion and integrated air-space-ground collaborative monitoring schemes. Monitoring schemes based on multi-source satellite remote sensing data fusion utilize space-based satellite remote sensing imagery (optical, radar, etc.) to achieve large-scale, all-weather surface observation, and are widely used for macroscopic identification and dynamic tracking of disasters such as floods, wildfires, and debris flows. Its core technologies include multi-source remote sensing data fusion and processing, as well as disaster monitoring based on machine learning.

[0003] In the field of multi-source remote sensing data fusion and processing, image fusion technology is a core component for achieving collaborative analysis of cross-modal information, playing a crucial role, especially in large-scale, time-sensitive applications such as disaster monitoring. With the diversification of remote sensing platforms and sensor types, effectively integrating heterogeneous data has become a prerequisite for improving sensing accuracy. To enhance robustness, some research has shifted towards feature fusion methods based on deep learning, achieving cross-modal matching by extracting key points and structural features from images. In machine learning-based disaster monitoring tasks, multimodal fusion methods often achieve more accurate monitoring results by fusing optical and SAR imagery.

[0004] Despite the strong potential of these methods, existing approaches remain insufficient in modal complementarity mining and feature modeling. Furthermore, limitations imposed by the spatial resolution of data sources hinder more refined disaster boundary identification and local anomaly detection. Most satellite remote sensing-based monitoring methods rely on low- to medium-resolution imagery, making it difficult to capture small-scale disaster details such as localized landslide cracks, micro-deformation around power transmission lines, or the initial spread of fire points. Simultaneously, single-platform observations suffer from long revisit periods, fixed perspectives, and susceptibility to weather interference, leading to monitoring blind spots or delayed responses in complex terrain or continuously changing disaster scenarios. Moreover, existing fusion strategies often focus on simple coupling at the data or feature layers, lacking the ability to model deep semantic relationships among multimodal information, thus failing to fully unleash the synergistic potential of multi-source data such as optical, SAR, and infrared data.

[0005] To overcome the limitations of single-scale data sources, integrated space-air-ground collaborative monitoring solutions have emerged and become the mainstream development direction in the field of disaster monitoring. The space-based platform, with unmanned aerial vehicles (UAVs) as its core carrier, undertakes the task of disaster perception at a mesoscale, with high spatiotemporal resolution and high mobility. UAVs typically carry multispectral cameras, thermal infrared sensors, lidar, and high-definition visible light imaging systems, enabling rapid deployment immediately after a disaster to conduct low-altitude patrols and multi-angle imaging of key areas. Compared to space-based satellites, UAVs fly at lower altitudes, have shorter revisit cycles, and achieve centimeter-level spatial resolution, clearly capturing crucial details such as minute surface deformations, power line tilting, initial fire spread, and localized water accumulation. For example, in wildfire monitoring, UAVs use infrared thermal imaging to identify high-temperature areas in real time and combine this with visible light imagery to determine the fire line's trajectory; in landslide or flood disasters, oblique photography generates 3D reality models to accurately assess terrain changes and their impact range. The ground-based system consists of a sensor network deployed around critical infrastructure, including video surveillance, weather stations, displacement monitors, and vibration sensors, responsible for continuous, high-precision monitoring of localized areas. These devices can collect parameters such as temperature, humidity, rainfall, and surface displacement in real time, and combine them with AI algorithms to automatically identify abnormal events (such as tower tilting and mountain cracks). Ground-based data is mainly used to verify space-based observation results and support on-site response decisions. Its advantage lies in high accuracy, but its coverage is limited and it relies on fixed deployment points.

[0006] The integrated air-space-ground collaborative monitoring solution can fuse multi-source data from air-based (satellite remote sensing), space-based (drones), and ground-based (ground sensors, monitoring equipment) sources to achieve collaborative perception and analysis of multi-dimensional, multi-scale, and multi-modal data, thereby improving the response capability to sudden disasters such as floods, wildfires, and mudslides. However, the current integrated air-space-ground system still faces three key technical challenges that restrict its development in terms of intelligence, real-time capabilities, and precision:

[0007] First, the problem of high-precision spatiotemporal fusion of cross-domain, multi-source, heterogeneous data has not yet been solved. Space-based, space-based, and ground-based platforms differ significantly in data modality, spatial resolution, temporal resolution, and coordinate systems, leading to registration bias and semantic inconsistencies in joint analysis of multi-source data. Achieving efficient alignment, unified calibration, and semantic-level fusion of cross-modal and cross-scale remote sensing and sensor data, and constructing a unified spatiotemporal perception benchmark framework, is the core bottleneck for achieving high-quality collaborative sensing.

[0008] Second, there is a lack of fine-grained feature modeling and accurate identification capabilities in complex disaster scenarios. Faced with nonlinear disaster processes such as the spread of floodwaters, the diffusion of heat anomalies from wildfires, and the deformation of terrain by debris flows, existing methods remain weak in the deep representation of multimodal features. Multimodal information fusion often remains at the data layer or shallow feature splicing, lacking effective modeling of the correlation between the physical mechanisms of disasters and remote sensing responses. This results in high false alarm rates, blurred boundaries, and inaccurate judgments of disaster development stages, making it difficult to support refined risk assessment and decision support.

[0009] Third, the dynamic coordination and real-time response mechanism of the air-space-ground system is still incomplete. Current systems generally adopt a centralized architecture of "sensing-transmission-centralized processing," resulting in high response latency and difficulty in meeting the timeliness requirements of disaster emergency response. At the same time, the lack of autonomous task scheduling and closed-loop coordination mechanisms between platforms makes it difficult to achieve dynamic linkage from "satellite initial screening—UAV re-inspection—ground verification—model feedback." How to construct a task-driven air-space-ground collaborative networking mechanism to achieve cross-platform intelligent collaboration has become a key challenge in improving the overall system response capability. Summary of the Invention

[0010] In order to overcome the problems of real-time performance and accuracy of existing integrated air-space-ground geological disaster monitoring schemes, this invention provides a multi-source fusion sensing integrated air-space-ground geological disaster identification method, system, and corresponding computer program product.

[0011] The technical solution provided by this invention is as follows:

[0012] A multi-source fusion sensing integrated air-space-ground geological hazard identification method includes the following process:

[0013] Real-time acquisition of satellite images of the target area from geostationary satellites, synthetic aperture radar satellites, and polar-orbiting satellites.

[0014] First, using the sampling frequency of geostationary satellites as a benchmark, a pre-trained image interpolator based on a generative diffusion model is used to perform temporal interpolation on the three types of satellite images. Then, using the resolution of polar orbit satellites as a benchmark, a pre-trained pixel interpolator based on a generative adversarial model is used to perform super-resolution reconstruction on the three types of satellite images. This results in a sequence of visible light image P1, infrared image P2, and SAR image P3 with spatiotemporal consistency and uniform resolution.

[0015] P1, P2, and P3 are synchronously input into a pre-trained coarse-grained disaster identification model in a time sequence to obtain macroscopic probability maps representing the probability of geological disasters occurring in each region during the corresponding time period. The coarse-grained disaster identification model employs a ViT network incorporating multimodal feature fusion. The network model includes a feature extraction module, a feature fusion module, and a prediction head. The feature extraction module uses a ViT-based encoder as the backbone network and extracts image features F from P1, P2, and P3. optical F IR and F SAR The feature fusion module employs a cross-modal attention mechanism and a spatial adaptive fusion strategy based on F... optical F IR and F SAR Generate fusion feature F final The prediction head is used to predict based on F. final Generate a macro probability map for the corresponding time period.

[0016] The system identifies high-risk local areas in the macro-probability map output by the aforementioned model, and acquires high-resolution images of the corresponding areas from ground monitoring equipment. When the spatial coverage of the high-pole monitoring equipment is insufficient, it utilizes drones to conduct on-site image acquisition. Then, the high-resolution images acquired by the ground monitoring equipment and / or drones are stitched together to obtain a high-resolution image of the local area.

[0017] High-resolution images of local areas are input into a pre-trained U-net-based fine-grained disaster identification model, and the model outputs the final geological disaster identification results.

[0018] As a further improvement of the present invention, the image interpolator generates the interpolated image at the missing time point through a forward process of progressively adding noise and a reverse process of denoising and reconstruction. The forward process generates a series of intermediate images x by adding noise. t The reverse process estimates the denoised distribution through a neural network and gradually reconstructs the image x0 at the missing time points.

[0019] Among them, the loss function L of the image interpolator in the pre-training stage difusion for:

[0020]

[0021] Where E represents the desired optimization objective of the image interpolator; ∈ represents Gaussian noise, ∈ θ The noise is the model's prediction, and t is the intermediate time.

[0022] As a further improvement of the present invention, the pixel interpolator includes a generator and a discriminator; the generator uses a high-resolution optical image I ref For reference, the low-resolution target modal image I lowGenerate a corresponding high-resolution image; the discriminator determines the category of the image generated by the generator.

[0023] Among them, the loss function L of the pixel interpolator in the pre-training stage total for:

[0024]

[0025] In the above formula, L perceptual Indicates perceived loss; L reconstruction L represents pixel-level reconstruction loss. GAN λ1 and λ2 represent the generation adversarial loss of the pixel interpolator; L represents the generation adversarial loss of the pixel interpolator. perceptual and L reconstruction The weights in the final loss; D(·) represents the discriminator; G(·) represents the generator.

[0026] As a further improvement of this invention, in the feature extraction module of the coarse-grained disaster recognition model, the multimodal images input from different channels are first segmented into fixed-size patches; each patch is linearly embedded to generate a high-dimensional feature representation incorporating positional encoding; then, it is input into a Transformer encoder containing a multi-head self-attention mechanism and a feedforward network to extract the image features F of the corresponding channels. optical F IR and F SAR .

[0027] As a further improvement of this invention, in the feature fusion module, the cross-modal attention mechanism first calculates the cross-modal attention weights Attention(Q,K,V) through the interaction of the query vector Q, key vector K, and value vector V:

[0028]

[0029] In the above formula, d k It is represented as a scaling factor for the feature dimension.

[0030] Then, utilizing the bidirectional gating mechanism in the spatial adaptive fusion strategy, the overall representation capability of the optical image and the detail reconstruction capability of the SAR image are dynamically adjusted according to the spatial distribution characteristics of the disaster area; thus generating the final fusion feature F. final :

[0031] F final =Gate(F optical ,F IR ,F SAR )⊙(F optical ,F IR ,F SAR );

[0032] In the above formula, Gate(·) represents the gate function of the bidirectional gating mechanism; ⊙ represents the element-wise multiplication operation.

[0033] As a further improvement of the present invention, the UAV performs adaptive cruise and image acquisition in a local area according to a preset path planning and image acquisition strategy.

[0034] As a further improvement of this invention, in the path planning stage, a dynamic path planning algorithm based on heuristic search combined with a real-time feedback mechanism is used to dynamically adjust the path; and the target area is rasterized into multiple sub-regions, with the following path planning optimization objective set to prioritize coverage of high-risk areas while minimizing the total flight distance Dist:

[0035]

[0036] In the above formula, d i,i+1 N represents the flight distance between subregions i and i+1; N represents the number of subregions; w i Let w represent the risk weight of the i-th sub-region. i The fusion weights of each pixel in the macroscopic probability map generated by the coarse-grained disaster identification model can be determined.

[0037] And / or, during the image acquisition phase, the UAV employs a serpentine or helical scanning mode within each sub-region; let the flight path of the UAV in sub-region i be P. i Path coverage C i for:

[0038]

[0039] In the above formula, This represents the area actually covered by the drone within sub-region i. Let i be the total area of ​​subregion i;

[0040] When acquiring images, the drone dynamically adjusts its flight altitude and scanning mode to achieve a coverage rate of C. i The preset threshold has been reached.

[0041] As a further improvement of this invention, the fine-grained disaster identification model includes an encoder and a decoder. The encoder extracts multi-scale features of the disaster area through multi-layer convolution and downsampling. The decoder restores spatial resolution through progressive upsampling, uses a multi-level feature fusion module to fuse global and local features, and combines skip connections to fuse the high-resolution features of the encoder.

[0042] As a further improvement of this invention, the convolution module of the encoder section adopts a dilated convolution module, and the expression of the dilated convolution module is:

[0043]

[0044] In the above formula, r is the void ratio, and w k x represents the weight of the k-th sampling point in the convolution kernel; K represents the number of sampling points in the convolution kernel; x i,j Let F be the pixel value at (i, j) in the input feature map; dilated (x i,j The output of the dilated convolution module; x i+r·k,j+r·k The input feature map contains the pixel value at (i+rk, j+rk);

[0045] As a further improvement of this invention, the decoder uses depthwise separable convolution instead of traditional convolution. Depthwise separable convolution includes two parts: depthwise convolution and pointwise convolution. The expression for depthwise separable convolution is:

[0046]

[0047] In the above formula, x i,j,c The original input to the depthwise separable convolution is represented by w. c and w c '' represents the weight of the c-th channel in depthwise convolution and pointwise convolution, respectively; C is the number of channels; F depthwise F represents the output of a depthwise convolution; pointwise This represents the final output of a depthwise separable convolution.

[0048] As a further improvement of this invention, the multi-level feature fusion module is implemented using a Spatial Pyramid Pooling (SPP) module. The SPP module extracts global and local features through pooling operations at different scales, and then concatenates them before inputting them into the corresponding layer of the decoder. The expression for the SPP module is:

[0049] F SPP =Concat(Pooling 1×1 (F), Pooling 3×3 (F), Pooling 5×5 (F));

[0050] In the above formula, F represents the original input of the multi-level feature fusion module; F SPP This represents the output of the multi-level feature fusion module; Pooling 1×1 and Pooling 3×3 These represent pooling operations with pooling window sizes of 1×1 and 3×3, respectively; Concat represents the feature concatenation operation.

[0051] As a further improvement of this invention, both the coarse-grained disaster identification model and the fine-grained disaster identification model employ cross-entropy loss L during the training phase. CE As a loss function, its expression is:

[0052]

[0053] In the above formula, y i,j Indicates the true label; H represents the predicted probability; H and W represent the height and width of the image, respectively.

[0054] The present invention also includes an integrated air-space-ground geological disaster identification system, which includes satellite receiving equipment, ground monitoring equipment, unmanned aerial vehicles and a command and control center.

[0055] The satellite receiving equipment is used to acquire monitoring data transmitted by geostationary satellites, synthetic aperture radar (SAR) satellites, and polar-orbiting satellites in orbit. The geostationary and polar-orbiting satellites acquire satellite images in different wavelengths within their respective orbits, including visible and infrared bands. The SAR satellite acquires SAR images. Ground-based monitoring equipment includes various multimodal sensors mounted on high towers to acquire environmental images of the surrounding area. The unmanned aerial vehicle (UAV) carries multimodal imaging equipment.

[0056] The command and control center is connected to satellite receiving equipment, ground monitoring equipment, and drones. It employs a multi-source fusion sensing integrated air-space-ground geological hazard identification method, as described above, to identify early geological hazard risks in target areas based on monitoring data from satellite receiving equipment and ground monitoring equipment, and sends dispatch instructions to drones.

[0057] As a further improvement of the present invention, the command and control center is equipped with a data preprocessing module, a coarse-grained disaster identification model, an unmanned aerial vehicle (UAV) scheduling component, and a fine-grained disaster identification model.

[0058] The data and processing module uses the sampling frequency of geostationary satellites as the time reference and a pre-trained image interpolator based on a generative diffusion model to perform temporal interpolation on three types of satellite images. Then, using the resolution of polar orbit satellites as the spatial reference, a pre-trained pixel interpolator based on a generative adversarial model is used to perform super-resolution reconstruction on the three types of satellite images. This results in a sequence of visible light image P1, infrared image P2, and SAR image P3 with spatiotemporal consistency and uniform resolution.

[0059] The coarse-grained disaster identification model adopts the ViT network with multimodal feature fusion. The network model includes a feature extraction module, a feature fusion module, and a prediction head. Based on the synchronously input P1, P2, and P3 images, the coarse-grained disaster identification model outputs a macroscopic probability map for the corresponding time period.

[0060] The drone scheduling component is used to identify the geographic information of high-risk local areas in the macro probability map, acquire high-resolution images of the corresponding areas from ground monitoring equipment, and call upon drones to conduct on-site image supplementation when the spatial coverage of high-pole monitoring equipment is insufficient. The data supplemented by drones and the data acquired by ground monitoring equipment can be stitched together to produce a high-resolution image of the high-risk local area in the macro probability map.

[0061] The fine-grained disaster identification model employs a pre-trained image processing model based on U-net for disaster identification and segmentation. This model is used to output the final geological disaster identification result based on a stitched image of the input local region.

[0062] The present invention also includes a computer program product comprising a computer program that, when executed by a processor, implements the aforementioned multi-source fusion sensing integrated air-space-ground geological disaster identification method.

[0063] The technical solution provided by this invention has the following beneficial effects:

[0064] This invention utilizes a generative diffusion model and a generative adversarial network to construct an image interpolator and a pixel interpolator capable of temporal interpolation and super-resolution reconstruction of satellite images, thereby solving the problem of inconsistency in temporal and spatial resolution of multi-source satellite data. The generated spatiotemporally aligned and resolution-consistent multi-source heterogeneous satellite data significantly improves the input data quality of disaster monitoring models, providing a reliable foundation for the accurate identification of disaster areas.

[0065] This invention, on the one hand, constructs a coarse-grained disaster identification model based on a ViT-based deep learning framework, combining cross-modal attention mechanisms and spatial adaptive fusion strategies. This model fully leverages the complementary characteristics of various satellite imagery data, including visible light, SAR, and infrared, to generate a macroscopic probability map of disaster areas. This network model significantly enhances the intelligent identification capability of disaster areas, providing efficient and intelligent technical support for disaster monitoring. On the other hand, based on the U-net framework, the encoder and decoder of the network model are optimized by combining dilated convolution, depthwise separable convolution, and spatial pyramid pooling modules, resulting in a fine-grained disaster identification model with multi-scale spatial information enhancement. This ensures adaptability to complex disaster scenarios and improves the real-time performance and accuracy of the network model in disaster monitoring.

[0066] In UAV scheduling and data acquisition tasks, this invention designs a dynamic path planning algorithm based on heuristic search, combined with a risk-reward-priority real-time feedback mechanism, to ensure that UAVs can prioritize coverage of high-risk areas and dynamically adjust inspection strategies. Through real-time acquisition of high-resolution imagery and multimodal sensor data, disaster monitoring information is further supplemented, improving the accuracy and efficiency of disaster monitoring.

[0067] This invention constructs a hierarchical, closed-loop collaborative sensing system, ranging from large-scale macroscopic initial screening to small-scale refined monitoring. Through dynamic task scheduling and a space-air-ground collaborative interaction mechanism, it achieves seamless integration of wide-area disaster monitoring coverage and localized refined analysis, improving the timeliness and accuracy of disaster monitoring. When applied to power facility monitoring, this solution can enhance the proactive defense capabilities of power facilities under extreme disaster scenarios, providing a scientific and systematic technical guarantee for the safe operation of power infrastructure. Attached Figure Description

[0068] Figure 1 This is a flowchart of the multi-source fusion sensing integrated air-space-ground geological disaster identification method provided in Embodiment 1 of the present invention.

[0069] Figure 2 This is a schematic diagram illustrating the interaction between different components in a multi-source fusion sensing integrated air-space-ground geological disaster identification method.

[0070] Figure 3 This is a schematic diagram illustrating the principle of temporal interpolation and spatial super-resolution reconstruction of multi-source heterogeneous satellite images in Embodiment 1 of the present invention.

[0071] Figure 4 This is a system architecture diagram of the integrated air-space-ground geological disaster identification system provided in Embodiment 2 of the present invention. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0073] Example 1

[0074] While existing technologies have established a multi-dimensional geological disaster monitoring system encompassing airborne, space-based, and ground-based equipment, the level of collaboration among these devices is low, and the ability for multimodal comprehensive analysis is weak. Practical applications remain limited by several key technological shortcomings, exposing deep-seated deficiencies in the intelligence and collaboration capabilities of the current technological system.

[0075] First, the insufficient fusion capability of multi-source satellite observation data has become a fundamental bottleneck restricting the accuracy of wide-area disaster perception. Currently, remote sensing satellites are diverse, covering multiple modalities such as optical, SAR, infrared, hyperspectral, and meteorological, but their imaging mechanisms, spatial resolution, time revisit periods, and radiometric characteristics vary greatly. For example, optical images are susceptible to cloud and fog interference, SAR images suffer from severe speckle noise and geometric distortion, and while meteorological satellites offer wide coverage, they lack spatial detail. Existing fusion methods mostly rely on empirical weighting, simple stitching, or coarse alignment based on traditional statistical models, lacking the ability to deeply model the physical mechanisms of multi-modal data correlation. More seriously, data acquired from different satellites generally suffer from spatiotemporal mismatch and registration errors, leading to spatial distortion, spectral distortion, or information redundancy in the fusion results, making it difficult to support highly reliable initial disaster screening and dynamic assessment. Essentially, current "fusion" still mainly remains at the "data stacking" stage, falling short of achieving semantically consistent intelligent integration.

[0076] Secondly, disaster feature extraction methods are outdated and lack intelligence. Faced with complex nonlinear processes such as flooding, wildfire anomalies, and debris flow deformation, mainstream methods still rely on manually designed features or shallow classification models, offering extremely limited deep semantic mining of multi-temporal and multi-modal satellite data. Especially in typical interference scenarios such as cloud cover and nighttime imaging, single-modal information becomes ineffective, and existing fusion strategies often only perform simple splicing at the data or feature layers, failing to effectively model the complementary relationships between visible light, SAR, and thermal infrared image data. This leads to high false alarm rates in anomaly identification, blurred boundaries, and frequent misjudgments in the disaster development stage, making it difficult to meet the high-precision and robust identification requirements of emergency response. Existing technologies exhibit significant "vulnerability" in complex environments, heavily relying on ideal observation conditions and lacking adaptive discrimination capabilities.

[0077] More significantly, the inter-platform coordination mechanism is weak, resulting in a sluggish overall system response. Currently, air-space-ground systems generally employ a linear, passive "sensing-transmission-centralized processing" workflow, lacking task-driven dynamic interaction and closed-loop feedback between platforms. While satellites possess wide-area coverage capabilities, their long revisit cycles and delayed responses prevent them from proactively triggering subsequent observations. UAVs mostly fly along preset routes, lacking the ability to autonomously adjust their trajectories based on real-time disaster conditions. Data collected by ground sensors is typically used only for post-event verification, making it difficult to drive airborne platforms to refocus and re-fly. In typical operational processes, UAV dispatch instructions are only issued after manual assessment by the ground center, with the entire process taking several hours, far from meeting the rigid requirement of "minute-level response" for sudden disasters. Furthermore, due to the lack of a unified mission planning and resource coordination mechanism, air-based, space-based, and ground-based systems operate independently, often resulting in overlapping coverage, resource waste, and monitoring blind spots, leading to overall system inefficiency.

[0078] To address the aforementioned problems, this embodiment provides a multi-source fusion sensing integrated air-space-ground geological disaster identification method, such as... Figure 1 As shown, this method, based on multi-source satellite remote sensing data and combined with ground sensors and UAV platforms, constructs a hierarchical, closed-loop collaborative sensing system from large-scale macroscopic initial screening to small-scale refined monitoring. Through key technologies such as spatiotemporal joint modeling, deep learning multimodal feature fusion, dynamic task scheduling, and refined semantic segmentation, it achieves efficient acquisition, intelligent identification, and rapid response to disaster information. Applying this solution to the field of pre-power infrastructure safety inspection can significantly improve the proactive defense capabilities of power facilities under extreme disaster scenarios.

[0079] Specifically, such as Figure 2 As shown, the solution provided in this embodiment includes the following steps:

[0080] First, in the large-scale macroscopic preliminary screening stage, multi-source satellite data is acquired and preprocessed. Temporal interpolation and spatial super-resolution reconstruction of multi-source satellite remote sensing data address issues such as cloud obstruction in optical images, SAR image noise, and low resolution of meteorological data. Specifically, in practical applications, a generative diffusion model can be used to temporally interpolate images at missing moments, and a generative adversarial network (GAN) can be combined to achieve spatial super-resolution reconstruction of low-resolution images, thereby generating spatiotemporally consistent multi-source heterogeneous satellite data. This high-quality input data provides a reliable foundation for subsequent disaster monitoring.

[0081] Then, multi-modal feature fusion deep learning techniques are used to uniformly model and analyze multi-source satellite data. For example, Vision Transformer (ViT) can be used as the core backbone network, combined with cross-modal attention mechanisms and spatial adaptive fusion strategies, to fully exploit the complementary characteristics of optical, SAR, infrared, and hyperspectral imagery, generating macroscopic probability maps characterizing the disaster risk of various local areas within the target region. Each pixel value in the macroscopic probability map represents the probability of a geological disaster occurring in that area, providing guidance for subsequent dynamic task scheduling and refined monitoring. The core backbone network here can be replaced with technologies such as Swing Transformer or ResNet.

[0082] Next, in the small-scale, refined monitoring phase, a dynamic interactive mechanism combining air, space, and ground-based monitoring is used to dynamically determine the availability of ground-based observation data for the target area, in conjunction with a macroscopic probability map. If the target area is already covered by ground-based monitoring equipment, real-time data from the ground-based monitoring equipment is directly used for disaster analysis; if the target area is a blind spot for ground-based observation, UAV path planning and inspection tasks are automatically triggered. The UAV prioritizes coverage of high-risk areas through dynamic path planning algorithms and acquires high-resolution images in real time to further supplement disaster monitoring information.

[0083] Finally, a refined semantic segmentation model is used to perform high-precision disaster area boundary extraction and local anomaly detection on data collected by ground-based monitoring equipment and space-based UAVs. For example, a disaster area identification and segmentation model can be constructed based on the U-net convolutional neural network architecture, combining dilated convolution and depthwise separable convolution to achieve multi-level feature fusion, ensuring accurate modeling of the global structure and local details of the disaster area. The spatial pyramid pooling module further enhances the modeling capability for multi-scale features, generating semantic segmentation results consistent with the input image size, providing technical support for accurate identification and dynamic monitoring of disaster areas.

[0084] In detail, the multi-source fusion sensing integrated air-space-ground geological disaster identification method provided in this embodiment includes the following process:

[0085] S1: Real-time acquisition of satellite images of the target area collected by geostationary satellites, synthetic aperture radar satellites, and polar-orbiting satellites.

[0086] Synthetic Aperture Radar (SAR) satellites are Earth observation remote sensing satellites carrying Synthetic Aperture Radar (SAR). They can achieve large-area surface imaging, thereby acquiring high-resolution SAR images. Geostationary orbit satellites and polar orbit satellites are Earth monitoring satellites located in different orbits. They can use their onboard multi-channel optical observation equipment to acquire multi-band images of different areas of the Earth's surface below, according to preset resolutions and sampling frequencies. In this embodiment, visible light and infrared images acquired by relevant satellites in the visible and infrared bands are mainly used for geological disaster monitoring. Due to the different orbits and altitudes of geostationary orbit satellites and polar orbit satellites, the time, resolution, and coverage area of ​​the satellite images acquired by the three types of satellites differ. Among them, geostationary orbit satellites have the highest data sampling frequency for a specified area, approximately once every 15 minutes. Polar orbit satellites have the highest resolution, enabling higher-precision imaging.

[0087] It should be noted that the solution provided in this embodiment is mainly used for monitoring the geological disaster risk of a specified target area. Therefore, it is only necessary to acquire a portion of the satellite images covering the target area from various types of satellite images. Technicians can directly request historical images of the target area within a specified time window, or crop the image portion of the target area contained in the full-view satellite image based on geographic information.

[0088] S2: Perform spatiotemporal interpolation and super-resolution reconstruction on multi-source satellites to generate spatiotemporally consistent and resolution-uniform multi-source heterogeneous satellite data.

[0089] The raw multi-source satellite data acquired in the previous step of this implementation exhibits significant differences in temporal and spatial resolution. For example, optical imagery possesses high spatial resolution and rich spectral information, but is susceptible to cloud and fog obstruction, leading to data loss at certain times. SAR imagery offers all-weather, all-time observation capabilities, but suffers from speckle noise and geometric distortion. Meteorological satellite data covers a wide area, but has lower spatial resolution, making it difficult to capture local details. These differences result in inconsistencies in the spatiotemporal dimensions of multimodal data, directly impacting the accuracy of disaster monitoring models in identifying disaster areas.

[0090] To address this issue, this embodiment fully utilizes the temporal continuity of multi-temporal remote sensing sequences and the complementary characteristics of multimodal data, combining a generative diffusion model and a generative adversarial model (GAN) to convert the acquired multi-source heterogeneous satellite data into spatiotemporally consistent satellite data with uniform resolution. For example... Figure 3 As shown, the satellite data conversion process includes: first, using the sampling frequency of geostationary satellites as a benchmark, a pre-trained image interpolator based on a generative diffusion model is used to perform temporal interpolation on the three types of satellite images. Then, using the resolution of polar orbit satellites as a benchmark, a pre-trained pixel interpolator based on a generative adversarial model is used to perform super-resolution reconstruction on the three types of satellite images.

[0091] In this embodiment, the temporal interpolation and super-resolution reconstruction operations are performed separately in three dimensions: visible light image, infrared image, and SAR image; ultimately resulting in a sequence of images—visible light image P1, infrared image P2, and SAR image P3—that are spatiotemporally and spatially consistent and have uniform resolution. It is important to emphasize that in the temporal interpolation stage of the generative diffusion model and the super-resolution reconstruction stage of the generative adversarial model, the original samples of the visible light and infrared images include both satellite images of the corresponding bands acquired by geostationary satellites and satellite images of the corresponding bands acquired by polar-orbiting satellites.

[0092] In this embodiment, the generative diffusion model used by the image interpolator, after pre-training, can generate interpolated images for any missing moments in the original sequence of satellite images. The image interpolation process of the generative diffusion model includes a forward process of progressively adding noise and a reverse process of denoising and reconstruction; ultimately resulting in a high-quality temporal interpolated image.

[0093] For example, suppose the satellite image at the missing time point is x0; the forward process of a generative diffusion model can generate a series of intermediate images x by adding noise. t The forward process can be represented as:

[0094]

[0095] In the above formula, xt-1 and x t These are intermediate images of state t-1 and state t, respectively; α t Let q(x) be the noise adjustment coefficient for state t, and I be the identity matrix; t |x t-1 The transition probability (t) represents the probability of transitioning from state t-1 to state t, which describes how the next state is generated based on existing conditions. In this embodiment, it can be assumed that this transition process follows a Gaussian distribution. In This represents a Gaussian distribution with a mean of x. t The variance is determined by and (1-α) t I control means that the magnitude of the variance changes over time, controlling the degree of influence of the previous state on the current state and the amount of noise introduced.

[0096] In the reverse process, the model parameters of the generative diffusion model can be combined to estimate the denoising distribution, and then the image x0 at the missing time point can be reconstructed step by step; the process is represented as:

[0097]

[0098] In the above formula, μ θ and Σ θ These are the prediction functions for the mean and variance, respectively, both learned by the network model during the training phase; p θ (x t-1 |x t p represents the conditional probability of obtaining the previous state given the current state; θ A time-series generation model to be learned is defined, which is to generate image data for missing time points based on image data at existing time points.

[0099] To obtain temporal interpolated images that meet accuracy requirements, the image interpolator based on a generative diffusion model provided in this embodiment needs to be pre-trained using sequence images of real satellite data before practical application. The loss function L of the image interpolator during the pre-training phase... difusion for:

[0100]

[0101] Where E represents the desired optimization objective of the image interpolator; ∈ represents Gaussian noise, ∈ θ The noise is the model's prediction, and t is the intermediate time.

[0102] In this embodiment, a pixel interpolator based on generative adversarial networks (GANs) can enhance the resolution of various satellite images. Addressing the low spatial resolution issue of meteorological satellites and infrared imagery, the pixel interpolator in this embodiment, after pre-training, can generate high-resolution target modal images using high-resolution optical or SAR imagery as references. The pixel interpolator essentially reconstructs entirely new high-resolution sample images from low-resolution sample images, and it comprises two parts: a generator and a discriminator. The generator uses high-resolution optical imagery I... ref For reference, the low-resolution target modal image I low A corresponding high-resolution image is generated. The discriminator classifies the image generated by the generator, determining whether the input image originates from a real high-resolution image or is a high-resolution image generated by the generator. When the image generated by the trained generator on the corresponding satellite sample dataset can "fool" the discriminator, preventing it from effectively determining the image's origin, then the generator can be used for high-resolution satellite image reconstruction tasks.

[0103] Among them, the loss function L of the pixel interpolator in the pre-training stage total for:

[0104]

[0105] In the above formula, L perceptual Indicates perceived loss; L reconstruction L represents pixel-level reconstruction loss. GAN λ1 and λ2 represent the generation adversarial loss of the pixel interpolator; L represents the generation adversarial loss of the pixel interpolator. perceptual and L reconstruction The weights in the final loss; D(·) represents the discriminator; G(·) represents the generator. In practical applications of this embodiment, the perceived loss L perceptual This can be achieved using the cross-entropy loss function; while the pixel-level reconstruction loss L... reconstruction This can be achieved using a structural similarity loss function.

[0106] Ultimately, the multi-source satellite data, after generative diffusion model interpolation and spatial super-resolution reconstruction, were aligned to a unified spatiotemporal scale. This preprocessed, high-quality multi-source heterogeneous satellite data can effectively assist subsequent disaster monitoring models in mining recurring disaster-related semantic representations, improving the accuracy of disaster area identification and dynamic monitoring, and providing reliable data support for disaster emergency response.

[0107] S3: Input P1, P2, and P3 synchronously in time sequence into a pre-trained coarse-grained disaster identification model to obtain a macroscopic probability map representing the probability of geological disasters occurring in each region during the corresponding time period.

[0108] To address the heterogeneity of multi-source satellite data and the complexity of disaster monitoring tasks, a deep learning-based multimodal fusion framework, namely the coarse-grained disaster identification model, is proposed. This model uses the Vision Transformer (ViT) as its core backbone network and combines cross-modal attention mechanisms and spatial adaptive fusion strategies to fully leverage the complementary characteristics of visible light satellite images, SAR satellite images, and infrared satellite images, achieving end-to-end optimization from multimodal feature extraction to deep semantic fusion. Through pixel-level classification and probabilistic map generation of disaster areas, the model can accurately identify the spatial distribution of disaster areas such as floods, wildfires, and mudslides, providing efficient and intelligent technical support for disaster monitoring.

[0109] Specifically, the coarse-grained disaster identification model constructed in this embodiment is a ViT network incorporating multimodal feature fusion. This network model includes a feature extraction module, a feature fusion module, and a prediction head. The feature extraction module uses a ViT-based encoder as the backbone network and extracts visible light features F from the input visible light image P1, infrared image P2, and SAR image P3 in three different branches. optical Infrared feature F IR and radar signature F SAR .

[0110] Specifically, in each branch, the input multimodal image is first segmented into fixed-size patches; linear embedding of each patch generates a high-dimensional feature representation. To preserve spatial information, positional encoding is introduced during feature embedding. The embedded feature sequence is then fed into multiple Transformer encoders, each containing a multi-head self-attention mechanism (MSA) and a feedforward network (FFN). Residual connections and layer normalization ensure network stability. This embodiment's ViT network, by stacking multiple encoders, captures the global dependencies and deep semantic features of disaster areas, ultimately obtaining feature maps F for the corresponding three types of images. optical F IR and F SAR .

[0111] The expression for the data processing procedure of ViT at this stage is:

[0112]

[0113] In the above formula, Z (l-1) Z (l) and Z (l+1) represents the feature representations of layers l-1, l, and l+1, respectively; MSA is a multi-head self-attention mechanism, FFN is a feedforward network, and LN is the representation layer normalization operation.

[0114] In this embodiment, the feature fusion module for constructing the coarse-grained disaster identification model is a deep semantic-level fusion framework based on cross-modal attention mechanisms and spatial adaptive fusion. It first captures the deep semantic relationships between feature maps of different types of satellite images through a cross-modal attention mechanism; to further improve the spatial consistency of the fused features, a spatial adaptive fusion strategy is introduced into the feature fusion module. Therefore, the feature fusion module adopts a cross-modal attention mechanism and a spatial adaptive fusion strategy based on F... optical F IR and F SAR Generate fusion feature F final .

[0115] Specifically, the cross-modal attention mechanism in the feature fusion module calculates the cross-modal attention weights Attention(Q,K,V) through the interaction of the query vector Q, key vector K, and value vector V:

[0116]

[0117] In the above formula, d k It is represented as a scaling factor for the feature dimension.

[0118] Then, utilizing the bidirectional gating mechanism in the spatial adaptive fusion strategy, the overall representation capability of the optical image and the detail reconstruction capability of the SAR image are dynamically adjusted according to the spatial distribution characteristics of the disaster area; thus generating the final fusion feature F. final :

[0119] F final =Gate(F optical ,F IR ,F SAR )⊙(F optical ,F IR ,F SAR );

[0120] In the above formula, Gate(·) represents the gate function of the bidirectional gating mechanism; ⊙ represents the element-wise multiplication operation.

[0121] Finally, the prediction head is used based on F final A macroscopic probability map for the corresponding time period is generated. In this embodiment, each pixel value of the probability map represents the probability that it belongs to an area where a geological disaster may occur, which can intuitively reflect the spatial distribution and severity of the disaster area.

[0122] In practical applications, coarse-grained disaster identification models also need to be trained and tested using sample data from three types of satellite images that have been collected in real time and preprocessed to achieve spatiotemporal uniformity and resolution alignment. The model parameters of the network model that meets the performance requirements after training should be retained. In this embodiment, cross-entropy loss L is used during the training phase of the coarse-grained disaster identification model. CE As a loss function, its expression is:

[0123]

[0124] In the above formula, y i,j Indicates the true label; H represents the predicted probability; H and W represent the height and width of the image, respectively.

[0125] After training the coarse-grained disaster identification model using this loss function, the model can achieve accurate classification and boundary extraction of disaster areas. The generated probability map provides efficient and intuitive decision support for disaster monitoring, significantly improving the intelligence level of disaster monitoring.

[0126] S4: Combine the macroscopic probability map output by the coarse-grained disaster identification model to locate risk areas and acquire full-coverage high-resolution images in real time.

[0127] In the solution provided in this embodiment, guided by a large-scale macroscopic probability map related to disasters, technicians can dynamically determine the potential disaster risk level of a target area and assess the availability of ground observation data in the relevant area. For example, after receiving the probability map, the ground command and control center first dynamically determines the availability of ground observation data within the target area. If the target area is already covered by ground sensors, the real-time data from the ground sensors is directly used for disaster analysis; if the target area is a blind spot for ground observation, a drone inspection mission is automatically triggered to conduct high-density inspections and data supplementation in the area. That is, the geographical information of high-risk local areas in the macroscopic probability map output by the aforementioned model is identified, and high-resolution images collected by ground monitoring equipment in the corresponding area are obtained. When the spatial coverage of the high-pole monitoring equipment is insufficient, drones are used to collect images on-site; then, the high-resolution images collected by the ground monitoring equipment and / or drones are stitched together to obtain a high-resolution image of the local area.

[0128] In practical applications, UAV inspections can acquire high-resolution images in real time. This data, along with real-time monitoring data collected by ground-based equipment, can be transmitted back to the ground command and control center to stitch together high-resolution images of identified high-risk areas. If no ground-based monitoring equipment is installed in the corresponding high-risk area, UAVs can be dispatched to traverse the area to complete image acquisition. In addition to ground-based monitoring equipment, UAV-acquired image data can also be jointly analyzed with satellite remote sensing data. The flexibility of UAVs allows them to quickly cover areas with a high probability of disasters and obtain crucial detailed information, such as local terrain changes, power line damage, or wildfire spread trends. Simultaneously, for various image data of the inspected area, super-resolution processing technology can be further employed to enhance image details and generate high-precision disaster monitoring results. Through a dynamic interactive mechanism combining air, space, and ground, the ground command and control center can integrate ground sensor data, UAV inspection data, and even remote sensing satellite data to form a multi-layered, multi-source collaborative disaster monitoring system.

[0129] To address the needs of UAV path planning and image acquisition in disaster monitoring, this embodiment also designs an efficient and dynamic path planning and data acquisition scheme in practical application. This scheme comprehensively considers the risk distribution, terrain complexity, and UAV flight performance of the disaster area; thereby ensuring that the UAV can fully cover the target area within a limited time, and dynamically adjust the inspection strategy in combination with real-time feedback to improve the accuracy and efficiency of disaster monitoring.

[0130] For example, during the path planning phase, this implementation rasterizes the target area into multiple sub-regions and sets the following path planning optimization objective to prioritize coverage of high-risk areas while minimizing the total flight distance Dist:

[0131]

[0132] In the above formula, d i,i+1 N represents the flight distance between subregions i and i+1; N represents the number of subregions; w i Let w represent the risk weight of the i-th sub-region. i The fusion weights of each pixel in the macroscopic probability map generated by the coarse-grained disaster identification model can be determined. In practical applications, w is dynamically adjusted. i This allows drones to prioritize inspecting high-risk areas while skipping low-risk areas, thus optimizing resource allocation.

[0133] To ensure the flexibility and real-time performance of path planning, this embodiment employs a dynamic path planning algorithm based on heuristic search (such as the A* algorithm), combined with a real-time feedback mechanism to dynamically adjust the path. When the UAV discovers that the actual risk in certain areas is lower than expected during inspection, the system lowers the priority of those areas and replans the remaining paths, concentrating more resources on high-risk areas. The dynamic update rule for path planning is: when the risk is below a threshold, for w... i Update:

[0134] w i →w i ·α risk

[0135] Where, α risk α is the pre-defined attenuation factor for risk weights. risk ∈(0,1).

[0136] During the image acquisition phase, each UAV can be equipped with multimodal sensors (such as visible light cameras, infrared cameras, and LiDAR) to perform high-resolution imaging and data acquisition of the target area. To ensure the coverage and quality of the acquired data, the UAV can employ a serpentine or helical scanning mode within each sub-region. Assume that the UAV's flight path in sub-region i is P. i Path coverage C i for:

[0137]

[0138] In the above formula, This represents the area actually covered by the drone within sub-region i. Let i be the total area of ​​subregion i;

[0139] Therefore, when a drone is acquiring images, its flight altitude and scanning mode should be dynamically adjusted to ensure coverage C. i The preset threshold is reached. Combining the aforementioned path planning and image acquisition scheme, this embodiment can utilize drones to achieve comprehensive coverage and refined monitoring of high-risk disaster areas, significantly improving the timeliness and accuracy of disaster monitoring and providing reliable data support for disaster emergency response. Of course, in practical applications, the location information provided by positioning satellites can also guide the drone's path planning.

[0140] S4: High-resolution images of local areas collected jointly by UAVs and ground monitoring equipment are input into a pre-trained U-net-based fine-grained disaster identification model, and the model outputs the final geological disaster identification results.

[0141] The fine-grained disaster identification model constructed in this embodiment is a more refined semantic segmentation model. It can utilize multi-source data collected by UAVs and ground monitoring equipment to perform high-precision disaster area boundary extraction and local anomaly detection. This network model, through multi-level feature fusion design of convolutional neural networks (CNN), balances generalization ability, robustness, and efficiency, meets the consistency requirements of multi-source data, and provides accurate disaster monitoring results in real-time response scenarios.

[0142] Specifically, the fine-grained disaster identification model is based on the U-net framework and employs a classic encoder-decoder structure. It optimizes both the encoder and decoder by incorporating dilated convolution and depthwise separable convolution. The encoder extracts multi-scale features of the disaster area through multi-layer convolution and downsampling. The data processing procedure for each layer of the encoder is as follows:

[0143]

[0144] In the above formula, and represents the feature maps of the encoder and decoder at layer l, respectively. Upsample represents the upsampling operation, and Conv represents the convolution operation.

[0145] To enhance the model's generalization ability and robustness, the encoder's convolutional module employs dilated convolutional modules and expands the receptive field to capture global contextual information of the disaster area, while maintaining computational efficiency. The expression for the dilated convolutional module is:

[0146]

[0147] In the above formula, r is the void ratio, and w k x represents the weight of the k-th sampling point in the convolution kernel; K represents the number of sampling points in the convolution kernel; x i,j Let F be the pixel value at (i, j) in the input feature map; dilated (x i,j The output of the dilated convolution module; x i+r·k,j+r·k The input feature map contains the pixel values ​​at (i+rk, j+rk).

[0148] The decoder recovers spatial resolution through progressive upsampling, employs a multi-level feature fusion module to fuse global and local features, and combines skip connections to fuse high-resolution features from the encoder. Through skip connections, the model can directly utilize high-resolution features from the encoding stage during decoding, avoiding information loss. In the decoder, to further optimize computational efficiency and reduce the number of parameters, depthwise separable convolution is used instead of traditional convolution operations. Depthwise separable convolution decomposes standard convolution into two steps: depthwise convolution and pointwise convolution. The expression for the data processing of depthwise separable convolution is as follows:

[0149]

[0150] In the above formula, x i,j,c The original input to the depthwise separable convolution is represented by w. c and w c '' represents the weight of the c-th channel in depthwise convolution and pointwise convolution, respectively; C is the number of channels; F depthwise F represents the output of a depthwise convolution; pointwise This represents the final output of the depthwise separable convolution. Employing depthwise separable convolution effectively maintains the feature extraction capability of the network model while significantly reducing computational complexity.

[0151] In the fine-grained disaster identification model provided in this embodiment, the multi-level feature fusion module can be implemented using a Spatial Pyramid Pooling (SPP) module to enhance the multi-scale modeling capability of disaster areas. The SPP module extracts global and local features through pooling operations at different scales, and then concatenates them before inputting them into the corresponding layer of the decoder. The expression for the SPP module is:

[0152] F SPP =Concat(Pooling 1×1 (F), Pooling 3×3 (F), Pooling 5×5 (F));

[0153] In the above formula, F represents the original input of the multi-level feature fusion module; F SPP This represents the output of the multi-level feature fusion module; Pooling 1×1 and Pooling 3×3 These represent pooling operations with pooling window sizes of 1×1 and 3×3, respectively; Concat represents the feature concatenation operation.

[0154] Finally, the fine-grained disaster recognition model provided in this embodiment can encode and decode high-resolution images of the input local region to obtain the feature information of the disaster area, and calculate the disaster prediction category of each pixel through the Softmax activation function; then, it outputs a mask of the fault region with the same size as the input image to obtain the semantic segmentation result of the disaster area. Similar to the coarse-grained disaster recognition model, the fine-grained disaster recognition model in this embodiment also needs to be trained using real high-resolution images before practical application. During the training phase, a cross-entropy loss L is uniformly adopted. CE As a loss function.

[0155] Example 2

[0156] Based on the scheme in Example 1, this embodiment further provides an integrated air-space-ground geological disaster identification system, such as... Figure 4 As shown, it includes satellite receiving equipment, ground monitoring equipment, drones, and a command and control center.

[0157] The satellite receiving equipment is used to acquire monitoring data transmitted by geostationary satellites, synthetic aperture radar (SAR) satellites, and polar-orbiting satellites in orbit. The geostationary and polar-orbiting satellites acquire satellite images in different wavelengths within their respective orbits, including visible and infrared bands. The SAR satellite acquires SAR images. Ground-based monitoring equipment includes various multimodal sensors mounted on high towers to acquire environmental images of the surrounding area. The unmanned aerial vehicle (UAV) carries multimodal imaging equipment.

[0158] The command and control center is connected to satellite receiving equipment, ground monitoring equipment, and drones. It employs a multi-source fusion sensing integrated air-space-ground geological hazard identification method, as described above, to identify early geological hazard risks in target areas based on monitoring data from satellite receiving equipment and ground monitoring equipment, and sends dispatch instructions to drones.

[0159] The command and control center is equipped with a data preprocessing module, a coarse-grained disaster identification model, a drone scheduling component, and a fine-grained disaster identification model. The data processing module uses the sampling frequency of geostationary satellites as the time reference and a pre-trained image interpolator based on a generative diffusion model to perform temporal interpolation on three types of satellite images. Then, using the resolution of polar-orbiting satellites as the spatial reference, a pre-trained pixel interpolator based on a generative adversarial model is used to perform super-resolution reconstruction on the three types of satellite images. This results in a sequence of visible light image P1, infrared image P2, and SAR image P3 with spatiotemporal consistency and uniform resolution.

[0160] The coarse-grained disaster identification model adopts the ViT network with multimodal feature fusion. The network model includes a feature extraction module, a feature fusion module, and a prediction head. Based on the synchronously input P1, P2, and P3 images, the coarse-grained disaster identification model outputs a macroscopic probability map for the corresponding time period.

[0161] The drone scheduling component is used to identify the geographic information of high-risk local areas in the macro probability map, acquire high-resolution images of the corresponding areas from ground monitoring equipment, and call upon drones to conduct on-site image acquisition when the spatial coverage of high-pole monitoring equipment is insufficient. The data acquired by the drones and the data acquired by the ground monitoring equipment can be stitched together to produce high-resolution images of the high-risk local areas contained in the macro probability map.

[0162] The fine-grained disaster identification model employs a pre-trained image processing model based on U-net for disaster identification and segmentation. This model is used to output the final geological disaster identification result based on a stitched image of the input local region.

[0163] In practical applications, the various functional modules in the command and control center are mainly implemented through program code. This embodiment further improves a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the aforementioned multi-source fusion perception air-space-ground integrated geological disaster identification method.

[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A space-air-ground integrated geological disaster identification method based on multi-source fusion perception, characterized in that, It includes: Real-time acquisition of satellite images of the target area collected by geostationary satellites, synthetic aperture radar satellites, and polar-orbiting satellites; First, using the sampling frequency of geostationary satellites as a benchmark, a pre-trained image interpolator based on a generative diffusion model is used to perform temporal interpolation on the three types of satellite images. Then, using the resolution of polar orbit satellites as a benchmark, a pre-trained pixel interpolator based on a generative adversarial model is used to perform super-resolution reconstruction on the three types of satellite images. Finally, a sequence of visible light image P1, infrared image P2, and SAR image P3 with spatiotemporal consistency and uniform resolution is obtained. The image interpolator generates the interpolated image for the missing moments through a forward process of progressively adding noise and a reverse process of denoising and reconstruction; the forward process generates a series of intermediate images by adding noise. x t The reverse process estimates the denoised distribution using a neural network and gradually reconstructs the images at the missing time points. x 0; The pixel interpolator includes a generator and a discriminator; Generator to generate high resolution optical imagery I ref For reference, low resolution target modality images I low A corresponding high resolution image is generated; a discriminator judges the class of the image generated by the generator; P1, P2, and P3 are synchronously input into a pre-trained coarse-grained disaster identification model in a time sequence to obtain a macroscopic probability map representing the probability of geological disasters occurring in each region during the corresponding time period. The coarse-grained disaster identification model adopts a ViT network that incorporates multimodal feature fusion. The network model includes a feature extraction module, a feature fusion module, and a prediction head. The feature extraction module uses a ViT-based encoder as the backbone network and extracts image features from P1, P2, and P3. F optical , F IR and F SAR ;feature The fusion module employs a cross-modal attention mechanism and a spatial adaptive fusion strategy based on... F optical , F IR and F SAR Generate fusion features F final The prediction head is used to predict based on F final Generate a macroeconomic probability map for the corresponding time period; Identify the geographic information of high-risk local areas in the macro probability map, obtain high-resolution images of the corresponding areas collected by ground monitoring equipment; and call on drones to conduct on-site image supplementation when the spatial coverage of the high-pole monitoring equipment is insufficient; stitch together the high-resolution images collected by the ground monitoring equipment and / or drones to obtain high-resolution images of the local areas. High-resolution images of local areas are input into a pre-trained U-net-based fine-grained disaster identification model, and the model outputs the final geological disaster identification results.

2. The multi-source fusion sensing integrated air-space-ground geological disaster identification method as described in claim 1, characterized in that: wherein Loss function of the image interpolator in the pre-training phase L difusion f : ; wherein, E denotes an optimization objective expectation of the image interpolator; is a Gaussian noise, is a model predicted noise, t is an intermediate time instant.

3. The multi-source fusion perception space-earth-ground integrated geological disaster identification method of claim 2, wherein: Loss function of pixel interpolator in pre-training stage L total For: ; In the above formula, Indicates perceived loss; Indicates pixel-level reconstruction loss; This represents the generation adversarial loss of the pixel interpolator; express Weight in the final loss; Indicates the discriminator; This indicates a generator.

4. The multi-source fusion perception space-earth-ground integrated geological disaster identification method of claim 1, wherein: In the feature extraction module of the coarse-grained disaster recognition model, the multimodal images input from different channels are first segmented into fixed-size patches. Each patch is then linearly embedded to generate a high-dimensional feature representation incorporating positional encoding. These patches are then input into a Transformer encoder containing a multi-head self-attention mechanism and a feedforward network to extract image features from the corresponding channels. F optical , F IR and F SAR ; And / or, in the feature fusion module, the cross-modal attention mechanism first uses the query vector Q Key vector K Sum value vector V Interactive computation of cross-modal attention weights : ; In the above formula, denotes the scaling factor for the feature dimension; Then, using the bidirectional gating mechanism in the spatial adaptive fusion strategy, the global representation capability of the optical image and the detail reconstruction capability of the SAR image are dynamically adjusted according to the spatial distribution characteristics of the disaster area; thus generating the final fused features. F final : ; In the above formulae, Gate (·) denotes a gating function of a bidirectional gating mechanism; denotes an element-wise multiplication operation.

5. The multi-source fusion perception space-earth-ground integrated geological disaster identification method of claim 1, wherein: The drone performs adaptive cruise and image acquisition in a local area according to a preset path planning and image acquisition strategy; And / or, during the path planning phase, a heuristic search-based dynamic path planning algorithm combined with a real-time feedback mechanism is used to dynamically adjust the path; and the target area is rasterized into multiple sub-regions, with the following path planning optimization objective set to minimize the total flight distance. At the same time, priority should be given to covering high-risk areas: ; In the above formula, Subregion i and i Flight distance between +1; N Indicates the number of sub-regions; Indicates the first i The risk weight of each sub-region is determined by the fusion weight of each pixel in the macro-probability map generated by the coarse-grained disaster identification model. And / or, during the image acquisition phase, the UAV employs a serpentine or spiral scanning pattern within each sub-region; assuming the UAV is in the sub-region The flight path is Path coverage for: ; In the above formula, denotes the area actually covered by the drone within the sub-region , is the total area of the sub-region . The unmanned aerial vehicle dynamically adjusts the flight height and scanning mode when collecting images, so that the coverage rate reaches a preset threshold. reaches a preset threshold.

6. The multi-source fusion sensing integrated air-space-ground geological disaster identification method as described in claim 1, characterized in that: The fine-grained disaster identification model includes an encoder and a decoder; the encoder extracts multi-scale features of the disaster area through multi-layer convolution and downsampling; The decoder recovers spatial resolution through progressive upsampling, uses a multi-level feature fusion module to fuse global and local features, and combines skip connections to fuse high-resolution features from the encoder. And / or, the convolution module in the encoder section uses a dilated convolution module, the expression of which is: ; In the above formula, r for the void fraction, is the weight of the sampling position point in the convolution kernel; k is the weight of the sampling position point in the convolution kernel; K This indicates the number of sampling points in the convolution kernel; For the input feature map ( i , j The pixel value at (). The output of the dilated convolution module; For the input feature map ( i+rk , j+rk The pixel value at (). And / or, the decoder uses depthwise separable convolution instead of traditional convolution. Depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution, and its expression is: ; In the above formula, The original input represents the depthwise separable convolution; These are the first two convolutions, depthwise convolution and pointwise convolution, respectively. c The weight of each channel; C for the number of channels; denotes the output of a depthwise convolution; denotes the final output of a depthwise separable convolution; And / or, the multi-level feature fusion module is implemented using the Spatial Pyramid Pooling (SPP) module. The SPP module extracts global and local features through pooling operations at different scales, and then concatenates them before inputting them into the corresponding layer of the decoder; its expression is: ; In the above formula, F This represents the original input to the multi-level feature fusion module; F SPP This represents the output of the multi-level feature fusion module; and These represent pooling operations with pooling window sizes of 1×1 and 3×3, respectively. Concat This indicates a feature splicing operation.

7. The multi-source fusion sensing integrated air-space-ground geological disaster identification method as described in claim 1, characterized in that: The coarse-grained disaster identification model and the fine-grained disaster identification model employ cross-entropy loss during the training phase. L CE As a loss function, its expression is: ; In the above formula, Indicates the true label; H represents the predicted probability; H and W represent the height and width of the image, respectively.

8. A space-air-ground integrated geological disaster identification system, characterized in that: It includes: Satellite receiving equipment, used to acquire monitoring data transmitted by geostationary satellites, synthetic aperture radar satellites and polar-orbiting satellites in orbit; The geostationary satellite and the polar satellite are used to acquire satellite images of different bands in their respective orbits; the synthetic aperture radar is used to acquire SAR images. Ground monitoring equipment, including various multimodal sensors installed on high towers on the ground for acquiring environmental images of the area surrounding the installation; Unmanned aerial vehicles (UAVs) carry multimodal imaging equipment. The command and control center is communicatively connected to the satellite receiving equipment, ground monitoring equipment, and UAV; and is used to employ the multi-source fusion sensing integrated air-space-ground geological disaster identification method as described in any one of claims 1-7, to identify the geological disaster risk of the target area in the early stage based on the monitoring data of the satellite receiving equipment and ground monitoring equipment, and to send scheduling instructions to the UAV. 9.The space-air-ground integrated geological disaster identification system of claim 8, wherein: The command and control center is equipped with a data preprocessing module, a coarse-grained disaster identification model, a drone scheduling component, and a fine-grained disaster identification model. The data and processing module is used to perform temporal interpolation on three types of satellite images using the sampling frequency of geostationary satellites as the time reference and a pre-trained image interpolator based on a generative diffusion model; then, using the resolution of polar orbit satellites as the spatial reference, it performs super-resolution reconstruction on the three types of satellite images using a pre-trained pixel interpolator based on a generative adversarial model; thereby obtaining a sequence of visible light image P1, infrared image P2, and SAR image P3 with spatiotemporal consistency and uniform resolution. The coarse-grained disaster identification model employs a ViT network that incorporates multimodal feature fusion. The network model includes a feature extraction module, a feature fusion module, and a prediction head. Based on the synchronously input P1, P2, and P3 images, the coarse-grained disaster identification model outputs a macroscopic probability map for the corresponding time period. The UAV scheduling component is used to identify the geographic information of high-risk local areas in the macro probability map, acquire high-resolution images of the corresponding areas collected by ground monitoring equipment, and call UAVs to conduct on-site image supplementation when the spatial coverage of the high-pole monitoring equipment is insufficient; the data supplemented by UAVs and the data collected by ground monitoring equipment are used to stitch together a high-resolution image of the high-risk local areas contained in the macro probability map. The fine-grained disaster identification model adopts a pre-trained image processing model based on U-net for disaster identification and segmentation; the fine-grained disaster identification model is used to output the final geological disaster identification result based on the stitched image of the input local area.

10. A computer program product comprising a computer program, characterized in that: When the computer program is executed by the processor, it implements the multi-source fusion sensing integrated air-space-ground geological disaster identification method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • High-resolution radar echo extrapolation prediction method based on fused satellite data

    CN120559654A

  • Method and electronic device for performing image processing

    US20250173825A1