Highway landslide disease identification method and device based on multi-source image fusion and cross-modal cooperation, equipment and medium

By performing seasonal-weather joint enhancement and modal normalization on optical remote sensing images, combined with cross-modal feature processing and an improved U-Net network, the problem of weakened DEM terrain structure information in multi-source satellite image fusion was solved, and high-precision landslide disease identification under complex terrain was achieved.

CN121904541AActive Publication Date: 2026-04-21CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-03-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing multi-source satellite image fusion methods fail to effectively handle the characteristics of different modal data in landslide disease identification, resulting in weakened DEM topographic structure information, weak model generalization ability, and difficulty in meeting the high-precision identification requirements of highway landslide diseases under complex terrain.

Method used

By enhancing optical remote sensing images with seasonal-weather joint data, modal scale normalization and stitching are performed. Combined with cross-modal channel access control and spatial attention weighting, the data is input into an improved U-Net network for feature extraction, filtering, and reconstruction. Multi-level feature fusion and boundary refinement are achieved by employing dual-path heterogeneous downsampling, gated skip connections, and decoder feature calibration.

Benefits of technology

It improves the accuracy of identifying landslide hazards on highways in complex terrain environments and enhances the precision of boundary identification. It preserves terrain structure information, suppresses feature redundancy noise, and achieves high-precision landslide target identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904541A_ABST
    Figure CN121904541A_ABST
Patent Text Reader

Abstract

The invention discloses a highway landslide disease identification method, device and equipment based on multi-source image fusion and cross-modal cooperation and a medium, and relates to the technical field of satellite landslide disasters, and the method comprises the steps: obtaining registered optical remote sensing images and digital elevation model data, carrying out the seasonal-weather joint data enhancement of optical images, and obtaining the data of the digital elevation model; a four-channel fusion feature map is obtained through modal normalization splicing, and features are enhanced through cross-modal channel admission control and space attention weighting; and inputting the enhanced features into an improved U-Net network containing dual-path heterogeneous downsampling, gating jump connection and decoder feature calibration, and performing feature extraction, screening, calibration, reconstruction, pixel classification and boundary refinement to realize highway landslide disease identification. According to the method, landform structure information is reserved, feature redundant noise is suppressed, a landslide target under a complex landform can be accurately recognized, and the recognition accuracy and boundary recognition precision of highway landslide diseases under the complex landform environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite landslide disaster technology, and in particular to a method, device, equipment and medium for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration. Background Technology

[0002] Existing multi-source satellite image fusion methods for landslide disease identification often employ simple feature stitching or post-fusion approaches, lacking targeted processing for the characteristics of different modalities. They only use basic data augmentation without joint optimization for seasons and weather. Furthermore, during the fusion process, a single modality, such as optical images, often dominates due to differences in numerical distribution or high information density, weakening the topographic structure information in the DEM data. In the application of the Unet model, traditional downsampling methods and unfiltered skip connection mechanisms are widely used for feature extraction and reconstruction, further weakening the continuous expression of DEM topographic structure information and introducing redundant features.

[0003] However, it has the following drawbacks: the data augmentation method is singular and cannot adapt to seasonal and weather changes, resulting in weak model generalization ability; at the same time, when processing optical remote sensing images and DEM topographic data, simple channel stitching or post-feature overlay is often used for fusion, which easily leads to single-modal features dominating training and weakening the effective feature representation of another modality; the recognition method based on U-Net network has obvious structural defects, and the single max pooling downsampling method is difficult to retain both local salient features and continuous topographic structure information, resulting in insufficient topographic structure representation in deep semantic features; unfiltered skip connections introduce a large number of redundant features and noise, affecting the decoding and reconstruction accuracy; at the same time, direct stitching of multi-scale features has scale bias problems, resulting in low recognition rate of small-scale landslide targets and insufficient segmentation boundary accuracy, making it difficult to meet the high-precision and high-stability recognition requirements of highway landslide diseases under complex terrain.

[0004] Therefore, improving the accuracy of identifying landslide hazards on highways in complex terrain environments and the precision of boundary identification has become an urgent problem to be solved. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, equipment and medium for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration, aiming to solve the technical problem of how to improve the accuracy of identifying highway landslide hazards and the precision of boundary identification in complex terrain environments.

[0006] To achieve the above objectives, this application proposes a method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration, comprising: Acquire aligned optical remote sensing images and digital elevation model data for the same geographic region; Seasonal-weather joint data augmentation is performed on the optical remote sensing images to generate enhanced optical images simulating different seasons and weather conditions; Modal scaling is performed on the enhanced optical image and the digital elevation model data respectively, and then the data is stitched together to obtain a four-channel fused feature map. The four-channel fused feature map is sequentially subjected to cross-modal channel admission control and spatial attention weighting to obtain the enhanced fused features; The enhanced fusion features are input into the improved U-Net network to generate highway landslide disease identification results. The improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. The step of inputting the enhanced fusion features into the improved U-Net network to generate highway landslide disease identification results includes: The enhanced fusion features are input into the dual-path heterogeneous downsampling module. The features are downsampled and adaptively weighted and fused through parallel paths of max pooling and strip pooling to obtain coded features that preserve multi-scale terrain structure. The gated skip connection mechanism is used to perform cross-scale attention filtering on the encoded features to obtain effective skip features. The effective skip features are aligned in terms of channel number and feature distribution by a lightweight 1×1 convolution in the decoder feature calibration module to obtain scale-calibrated fused features. The scale-calibrated fused features are upsampled and reconstructed step by step to obtain a landslide feature map with multi-level feature fusion. The landslide feature map fused from the multi-level features is classified at the pixel level and its boundaries are refined to obtain the identification results of highway landslide diseases.

[0007] In one embodiment, the step of performing seasonal-weather joint data augmentation on the optical remote sensing image to generate enhanced optical images simulating different seasons and weather conditions includes: The baseline seasonal category is determined based on the acquisition time of the optical remote sensing images; The optical remote sensing image is transformed by the spectral characteristics of surface vegetation to simulate the changes in vegetation cover, hue and texture under the baseline seasonal category, thus obtaining a seasonally enhanced optical image. The illumination intensity and atmospheric visibility of the seasonally enhanced optical image are adjusted to simulate the changes in image brightness, contrast and detail recognition under a preset weather type in the baseline seasonal category, thereby obtaining a weather-enhanced optical image. The enhanced optical image of the weather is normalized and calibrated to obtain a calibrated enhanced optical image; The calibrated enhanced optical image is fused with the original optical remote sensing image at the feature layer to obtain a fused enhanced optical image. An edge-preserving filtering algorithm is used to enhance the details of the fused enhanced optical image to obtain an enhanced optical image.

[0008] In one embodiment, the step of sequentially performing cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map to obtain the enhanced fused features includes: The four-channel fused feature map is input into the cross-modal channel admission control module based on the extrusion structure. A first weight and a second weight set are generated through a weight network of preset dimensions. The first weight is the optical three-channel weight, and the second weight is the digital elevation model single-channel weight. The first and second weights are subjected to modal grouping normalization and then integrated to obtain the calibration channel weight set. The four-channel fusion feature map is subjected to channel-level weighted filtering using the calibration channel weight set to obtain a channel-enhanced fusion feature map. A local spatial convolution operation is performed on the channel enhancement fusion feature map using a convolution kernel of a preset size to generate a spatial attention weight map; The spatial attention weight map and the channel enhancement fusion feature map are multiplied by pixel-level weighting to obtain the enhanced fusion feature.

[0009] In one embodiment, the step of inputting the enhanced fused features into the dual-path heterogeneous downsampling module, downsampling the features through parallel paths of max pooling and strip pooling, and adaptively weighting and fusing them to obtain multi-scale terrain structure-preserving encoded features includes: The enhanced fusion features are input into the dual-path heterogeneous downsampling module to extract shallow semantic features, and the initial encoded features are obtained. The initial encoded features are distributed to the max pooling path and the strip pooling path, and local features and terrain structure features are downsampled simultaneously to obtain the first downsampled features and the second downsampled features. Adaptive weight allocation is performed on the first downsampled feature and the second downsampled feature to obtain the feature fusion weight; The first downsampled feature and the second downsampled feature are weighted and fused using the feature fusion weights to obtain a single-level fused coding feature; By repeatedly performing convolutional downsampling and dual-path fusion operations at each level, the single-level fusion coding features of each level are integrated to obtain multi-scale terrain structure-preserving coding features.

[0010] In one embodiment, the step of performing cross-scale attention filtering on the encoded features through the gated skip connection mechanism to obtain effective skip features includes: The encoded features are stratified by scale to obtain the encoded features to be screened at each level; Attention-aware evaluation is performed on the encoded features to be screened at each level to generate hierarchical attention weights; The hierarchical attention weights are used to weight the coding features to be screened at the corresponding level to obtain the weighted coding features at each level. Redundancy suppression is performed on the weighted coding features at each level to obtain the processed weighted coding features; By integrating the weighted coding features processed at each level, an effective skip feature is obtained.

[0011] In one embodiment, the step of performing progressive upsampling and feature reconstruction on the scale-calibrated fused features to obtain a landslide feature map with multi-level feature fusion includes: Perform transposed convolutional upsampling on the scale-calibrated fused features to restore the spatial resolution of the features and obtain single-level upsampled features; The single-level upsampling features are concatenated and fused with the effective skip features of the corresponding level to obtain single-level reconstructed features; Based on the integration of all the single-level reconstruction features, the reconstruction features of each level are obtained; Multi-scale feature fusion is performed on the reconstructed features at each level to integrate shallow details and deep semantic features, resulting in a fused feature map; Lightweight convolutional detail enhancement is performed on the fused feature map to obtain a landslide feature map with multi-level feature fusion.

[0012] In one embodiment, the step of performing pixel-level classification and boundary refinement on the landslide feature map fused with the multi-level features to obtain the highway landslide disease identification result includes: Pixel-level feature mapping is performed on the landslide feature map fused from the multi-level features to generate the category probability distribution of each pixel; The probability distribution of the categories is classified into pixels according to a preset probability threshold to obtain an initial binary segmentation map of the landslide. Edge contour detection is performed on the initial binary segmentation map of the landslide to extract the initial boundary features of the landslide area; The initial boundary features are refined by combining the terrain structure features of the digital elevation model to obtain an accurate landslide boundary; The precise landslide boundary and the landslide feature map fused with the multi-level features are combined to generate the highway landslide disease identification result.

[0013] Furthermore, to achieve the above objectives, this application also proposes a highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration, wherein the highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration includes: The acquisition module is used to acquire optical remote sensing images and digital elevation model data aligned to the same geographic area; The data augmentation module is used to perform seasonal-weather joint data augmentation on the optical remote sensing image to generate enhanced optical images that simulate different seasons and weather conditions; The fusion module performs modal scale normalization on the enhanced optical image and the digital elevation model data respectively and then stitches them together to obtain a four-channel fused feature map. The processing module sequentially performs cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map to obtain the enhanced fused features; The result module is used to input the enhanced fused features into the improved U-Net network to generate highway landslide disease identification results. The improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. It is also used to input the enhanced fused features into the dual-path heterogeneous downsampling module, downsample the features through parallel paths of max pooling and strip pooling, and adaptively weight and fuse them to obtain multi-scale terrain structure-preserving encoded features. The gated skip connection mechanism performs cross-scale attention filtering on the encoded features to obtain effective skip features. The lightweight 1×1 convolution in the decoder feature calibration module aligns the effective skip features in terms of channel number and feature distribution to obtain scale-calibrated fused features. The scale-calibrated fused features are then subjected to progressive upsampling and feature reconstruction to obtain a multi-level feature fused landslide feature map. Finally, the multi-level feature fused landslide feature map is subjected to pixel-level classification and boundary refinement to obtain the highway landslide disease identification results.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration as described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration as described above.

[0016] This application acquires registered optical remote sensing images and digital elevation model data, performs seasonal-weather joint data augmentation on the optical images, and obtains a four-channel fused feature map through modal normalization and stitching. Then, it enhances the features through cross-modal channel admission control and spatial attention weighting. The enhanced features are input into an improved U-Net network containing dual-path heterogeneous downsampling, gated skip connections, and decoder feature calibration. Through feature extraction, filtering, calibration, reconstruction, pixel classification, and boundary refinement, it achieves highway landslide hazard identification. It preserves terrain structure information, suppresses feature redundancy noise, and can accurately identify landslide targets in complex terrain, improving the accuracy of highway landslide hazard identification and boundary recognition in complex terrain environments. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the first embodiment of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in this application; Figure 2 This is a flowchart illustrating the second embodiment of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in this application; Figure 3 This is a schematic diagram of the module structure of the highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration in this application; Figure 4 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in the embodiments of this application.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] Existing multi-source satellite image fusion methods for landslide disease identification often employ simple feature stitching or post-fusion approaches, lacking targeted processing for the characteristics of different modal data. During the fusion process, a single modality, such as optical images, often dominates due to differences in numerical distribution or high information density, weakening the topographic structure information in the DEM data. Furthermore, only basic data augmentation is used, without joint optimization for different modal data based on seasons and weather. In the application of the Unet model, traditional downsampling methods and unfiltered skip connection mechanisms are widely used for feature extraction and reconstruction, further weakening the continuous expression of DEM topographic structure information and introducing redundant features.

[0023] However, it has the following drawbacks: the data augmentation method is singular and cannot adapt to seasonal and weather changes, resulting in weak model generalization ability; at the same time, when processing optical remote sensing images and DEM topographic data, simple channel stitching or post-feature overlay is often used for fusion, which easily leads to single-modal features dominating training and weakening the effective feature representation of another modality; the recognition method based on U-Net network has obvious structural defects, and the single max pooling downsampling method is difficult to retain both local salient features and continuous topographic structure information, resulting in insufficient topographic structure representation in deep semantic features; unfiltered skip connections introduce a large number of redundant features and noise, affecting the decoding and reconstruction accuracy; at the same time, direct stitching of multi-scale features has scale bias problems, resulting in low recognition rate of small-scale landslide targets and insufficient segmentation boundary accuracy, making it difficult to meet the high-precision and high-stability recognition requirements of highway landslide diseases under complex terrain.

[0024] Based on the above, this application also provides a method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in this application.

[0025] In this embodiment, the method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration includes steps S10 to S50: Step S10: Obtain optical remote sensing images and digital elevation model data aligned to the same geographic area.

[0026] It should be noted that optical remote sensing images are formed by capturing visible light, near-infrared, and other electromagnetic wave information reflected by ground objects using optical remote sensing sensors. Digital elevation model (DEM) data is a digital model that records surface elevation information using discrete grid cells. It is core topographic data obtained through topographic surveying technologies such as remote sensing and lidar, and can accurately reflect the topographic relief, terrain structure characteristics, slope aspect, and other topographic spatial information of a geographical area.

[0027] Specifically, optical remote sensing images of the same geographic area are acquired through a satellite remote sensing platform, while digital elevation model data of the same area is also collected. A geographic coordinate registration algorithm is used to spatially align the two types of data to ensure that each pixel of the optical image is precisely matched with the corresponding grid cell of the digital elevation model in terms of geographic location. This eliminates spatial offsets caused by differences in acquisition angle, resolution, or projection method, and provides a data foundation for subsequent cross-modal feature fusion.

[0028] Step S20: Perform seasonal-weather joint data augmentation on the optical remote sensing image to generate enhanced optical images simulating different seasons and weather conditions.

[0029] It should be noted that seasonal-weather joint data enhancement refers to a data enhancement method that determines the baseline seasonal category based on the acquisition time of the optical remote sensing image, sequentially performs surface vegetation spectral feature transformation and adjusts light intensity and atmospheric visibility to simulate different seasons and weather conditions, and obtains an enhanced optical image through pixel normalization calibration, feature layer fusion, and edge-preserving filtering. Step S20 includes: determining the baseline seasonal category based on the acquisition time of the optical remote sensing image; performing surface vegetation spectral feature transformation on the optical remote sensing image to simulate vegetation cover, tone, and texture changes under the baseline seasonal category to obtain a seasonal enhanced optical image; adjusting light intensity and atmospheric visibility on the seasonal enhanced optical image to simulate changes in image brightness, contrast, and detail recognition under a preset weather type under the baseline seasonal category to obtain a weather enhanced optical image; performing pixel value normalization calibration on the weather enhanced optical image to obtain a calibrated enhanced optical image; fusing the calibrated enhanced optical image with the original optical remote sensing image at the feature layer to obtain a fused enhanced optical image; and using an edge-preserving filtering algorithm to enhance the details of the fused enhanced optical image to obtain an enhanced optical image.

[0030] It's important to understand that the baseline season category refers to the natural season type of a geographic area determined based on the actual time of acquisition of the optical remote sensing image. This category, categorized according to the natural season in which the optical remote sensing image was acquired (e.g., spring, summer, autumn, or winter), corresponds to typical vegetation growth states, surface color variations, and other natural characteristics of the area. For example, in spring, vegetation is sparse and tender green; in summer, vegetation is lush, dense, and green; in autumn, vegetation is withered and sparse; and in winter, vegetation is dry and bare. This provides a clear direction for subsequent transformation of the spectral characteristics of surface vegetation. Surface vegetation spectral characteristic transformation is a feature adjustment operation performed on the spectral information of vegetation areas in optical remote sensing images. Spectral characteristics are inherent properties of vegetation; vegetation exhibits different spectral reflectance patterns in different seasons. By transforming these characteristics, the spectral performance of vegetation in remote sensing images under the corresponding season can be simulated, thereby restoring the visual characteristics of vegetation in different seasons. Preset weather types refer to various natural weather types selected in advance to improve the model's adaptability to different weather environments. These are typical weather categories pre-defined for highway landslide monitoring scenarios, namely sunny, cloudy, foggy, and rainy days. Different weather types have different impacts on the imaging effect of optical remote sensing images. Setting corresponding preset weather types based on the baseline seasonal category ensures the rationality of weather simulation and the realism of the scene. Pixel value normalization calibration is a standardization process performed on the pixel values ​​of the image after weather feature adjustment. Feature layer fusion is a fusion operation performed on two images at the feature extraction level. It is not a simple pixel-level superposition, but rather extracts the feature information of both the calibrated enhanced optical image and the original optical remote sensing image, and then organically integrates the two types of features so that the fused image retains both the core features of the original image and the multi-scene features of the enhanced image. Edge-preserving filtering algorithm is an image processing algorithm that combines image smoothing and edge detail preservation. While suppressing noise and smoothing the image, this algorithm can accurately preserve the edge contours and detailed features of ground objects in the image, avoiding the loss of key details such as landslide areas and terrain boundaries due to filtering operations. Seasonally enhanced optical images refer to images obtained by transforming the spectral characteristics of surface vegetation on the original optical remote sensing image. These images simulate the typical characteristics of vegetation cover, tone, and texture under the baseline seasonal category, achieving feature enhancement in the seasonal dimension. Weather-enhanced optical images refer to images obtained by adjusting the light intensity and atmospheric visibility of the seasonally enhanced optical image. Building upon seasonal enhancement, these images further simulate the imaging characteristics of a preset weather type under the baseline seasonal category, achieving feature enhancement in the weather dimension. Calibrated enhanced optical images refer to images obtained by normalizing and calibrating the pixel values ​​of the weather-enhanced optical image. These images eliminate pixel value deviations introduced by previous enhancement operations, ensuring the standardization of pixel values ​​and providing a unified numerical basis for subsequent feature fusion.Fusion-enhanced optical images refer to images obtained by fusing calibrated enhanced optical images with original optical remote sensing images at the feature layer. This image retains the core landslide features of the original image and the multi-seasonal and multi-weather scene features of the enhanced image, achieving complementarity and integration of image features.

[0031] Specifically, firstly, the baseline season category is determined based on the acquisition time of the optical remote sensing image. Using the actual acquisition time recorded in the image's metadata, and considering the climate and seasonal division patterns of the acquisition area, the baseline season category corresponding to the image is determined, providing a reference for subsequent seasonal feature simulation. Subsequently, the surface vegetation spectral feature transformation is performed on the optical remote sensing image to simulate vegetation cover, tone, and texture changes under the baseline season category, resulting in a seasonally enhanced optical image. For the determined baseline season category, typical spectral features of vegetation in that season are extracted, and the spectral features of the vegetation areas in the original optical remote sensing image are adjusted. Simultaneously, the vegetation cover, surface tone, and texture details corresponding to that season are simulated, completing the seasonal feature enhancement of the image and generating the seasonally enhanced optical image. Next, the illumination intensity and atmospheric visibility of the seasonally enhanced optical image are adjusted to simulate the changes in image brightness, contrast, and detail resolution under a preset weather type within the baseline seasonal category, thus obtaining a weather-enhanced optical image. Based on the determined baseline seasonal category, the corresponding preset weather type is selected. According to the imaging characteristics of this weather type, the illumination intensity of the seasonally enhanced optical image is increased or decreased, and the atmospheric visibility is adjusted accordingly. The overall brightness, contrast, and ground feature detail resolution of the image are adjusted simultaneously to simulate the image characteristics under this weather type, generating a weather-enhanced optical image. Next, pixel value normalization calibration is performed on the weather-enhanced optical image to obtain a calibrated enhanced optical image. Addressing the abnormal pixel value distribution issue that occurs in the weather-enhanced optical image after seasonal and weather simulation operations, a standardized numerical processing method is used to uniformly adjust the values ​​of all pixels in the image, standardizing them to a set range and eliminating numerical deviations. Then, the calibrated enhanced optical image is fused with the original optical remote sensing image at the feature layer to obtain a fused enhanced optical image. Multi-scene enhancement features from the calibrated enhanced optical image and the core landslide terrain features from the original optical remote sensing image are extracted separately. These two types of features are organically integrated at the feature level, allowing the fused image to simultaneously retain the true core features of the original image and the multi-seasonal, multi-weather scene features of the enhanced image. Finally, an edge-preserving filtering algorithm is used to enhance the details of the fused enhanced optical image, resulting in an enhanced optical image. This edge-preserving filtering algorithm processes the fused enhanced optical image, suppressing noise interference while accurately preserving key details such as landslide areas, terrain boundaries, and feature outlines, completing the image detail enhancement and ultimately generating the enhanced optical image.

[0032] Step S30: Modal scale normalization is performed on the enhanced optical image and digital elevation model data respectively, and the data are then stitched together to obtain a four-channel fused feature map.

[0033] It should be noted that modal scaling normalization is an operation that standardizes data from different sources with inconsistent numerical ranges. Augmented optical images and digital elevation model (DEM) data belong to different modalities, exhibiting significant differences in numerical magnitude, distribution range, and physical meaning. Modal scaling normalization maps both types of data to a unified numerical range, making the different modalities comparable and fusionable at the feature level. The four-channel fusion feature map is a combined feature map formed by stitching augmented optical images and DEM data along the channel dimension. Augmented optical images contain feature information from three channels, while DEM data contains feature information from one channel. The two types of data are stitched together in an orderly manner along the channel dimension to form a feature map with a total of four channels.

[0034] Specifically, firstly, modal scaling normalization is performed on the enhanced optical image. Following a preset standardization rule, the values ​​of each channel in the enhanced optical image are adjusted individually to unify the values ​​of all pixels into the same distribution range, eliminating numerical differences between different channels and obtaining normalized optical image features. Secondly, modal scaling normalization is performed on the digital elevation model (DEM) data. Using a standardization rule matching the enhanced optical image, the values ​​of all grid cells in the DEM data are adjusted to ensure that the numerical range of the elevation data is consistent with the optical image features, resulting in normalized terrain data features. Then, the normalized optical image features and the normalized terrain data features are concatenated along the channel dimension. The three feature channels of the optical image and one feature channel of the terrain data are combined in a fixed order without changing the spatial dimensions and positional correspondence of the data. Finally, a four-channel fused feature map containing both optical and terrain information is formed, serving as input data for subsequent cross-modal feature processing.

[0035] Step S40: Perform cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map in sequence to obtain the enhanced fused features.

[0036] It should be noted that step S40 includes: inputting the four-channel fused feature map into the cross-modal channel admission control module based on the extrusion structure; generating a first weight and a second weight set through a weight network of preset dimensions, wherein the first weight is the optical three-channel weight and the second weight is the digital elevation model single-channel weight; performing modal grouping normalization processing on the first weight and the second weight and integrating them to obtain a calibration channel weight set; using the calibration channel weight set to perform channel-level weighted filtering on the four-channel fused feature map to obtain a channel-enhanced fused feature map; performing local spatial convolution operation on the channel-enhanced fused feature map using a convolution kernel of preset size to generate a spatial attention weight map; and performing pixel-level weighted multiplication on the spatial attention weight map and the channel-enhanced fused feature map to obtain the enhanced fused feature.

[0037] It's important to understand that the squeeze-excitation structure is a network structure used for learning feature channel weights. This structure first compresses global information through a squeezing operation, then learns the importance of different channels through an excitation operation, and finally assigns corresponding weights to each channel. The cross-modal channel admission control module is a functional module used for channel weight allocation and filtering of multi-source modal features. This module is designed for two different modalities of data: optical remote sensing images and digital elevation models. It can distinguish the contribution level of each channel under different modalities, assigning higher weights to important channels and suppressing secondary channels.

[0038] Specifically, firstly, the four-channel fused feature map is input into the cross-modal channel admission control module based on a squeeze excitation structure. Global average pooling is used to compress the spatial dimension of each channel in the four-channel fused feature map, generating channel descriptors. Then, the channel descriptors are input into a 2-dimensional weight network in the bottleneck layer for dimensionality reduction and expansion, generating initial weights for the four channels. These weights are divided into a first set of weights corresponding to the three optical channels and a second set of weights corresponding to the single digital elevation model channel. This refines the importance ratio between optical and terrain information, adapting to a very small number of channels and ensuring the effectiveness of weight learning. Next, modal grouping normalization is performed on the first and second weights. The three optical channel weights are normalized as one group, and the single digital elevation model weights are normalized as another group. The two sets of normalized weights are integrated to obtain a calibration channel weight set. This prevents the numerical distribution difference caused by the dominance of optical channels from leading to single-modality dominance in training, thus balancing the dual-modal characteristics. The contribution of features is evaluated. Next, the four-channel fusion feature map is subjected to channel-level weighted filtering using a calibration channel weight set. The weight of each channel is multiplied with the feature map of the corresponding channel to obtain a channel-enhanced fusion feature map, thus preserving and strengthening effective features. Subsequently, a 5×5 convolution kernel is used to perform local spatial convolution on the channel-enhanced fusion feature map. By capturing the spatial context relationship between landslide areas and mountain boundaries in high-resolution remote sensing images through a larger receptive field, a spatial attention weight map is generated, and the weight values ​​are constrained to the range of 0.1 to 0.9 to avoid feature suppression or over-amplification caused by extreme weights. Finally, the spatial attention weight map and the channel-enhanced fusion feature map are multiplied at the pixel level, giving higher weights to key spatial locations such as landslide areas and topographic fault zones while suppressing background noise and abnormal pixels, thus obtaining enhanced fusion features. This improves the model's ability to perceive complex terrain environments under the dual enhancement of channel and spatial dimensions.

[0039] Step S50: Input the enhanced fusion features into the improved U-Net network to generate highway landslide disease identification results.

[0040] It should be noted that the improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module.

[0041] Specifically, the enhanced fused features are first fed into the input of the improved U-Net network. This improved U-Net network integrates three core improved modules on the basis of the classic U-Net encoder-decoder framework: a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. After entering the network, the enhanced fused features first flow through the encoder's dual-path heterogeneous downsampling module to extract multi-scale encoded features and preserve terrain structure information. Then, through the gated skip connection mechanism, cross-scale attention filtering is performed on the encoder's output encoded features, only transmitting effective skip features to the corresponding level of the decoder. Subsequently, the decoder feature calibration module aligns the effective skip features with the decoder's upsampled features in terms of channel count and feature distribution to eliminate scale bias. Finally, through step-by-step upsampling, feature reconstruction, pixel-level classification, and boundary refinement by the decoder, the final result for identifying highway landslide hazards is generated.

[0042] The dual-path heterogeneous downsampling module solves the problem of terrain structure information loss caused by single pooling, ensuring that encoded features retain both local salient features and continuous terrain features, providing complete structural support for landslide identification. The gated skip connection mechanism eliminates redundancy and noise through feature filtering, avoiding feature interference caused by the unfiltered skip connections in classic U-Net, and improving the accuracy of decoder feature reconstruction. The decoder feature calibration module solves the scale bias problem in multi-scale feature stitching, ensuring accurate matching of encoder and decoder features and improving the effect of multi-scale feature fusion. The overall improved U-Net network achieves end-to-end optimization from feature extraction, filtering, calibration to reconstruction, significantly improving the identification accuracy and boundary characterization ability of highway landslide hazards in complex terrain.

[0043] This embodiment acquires registered optical remote sensing images and digital elevation model data, performs seasonal-weather joint data augmentation on the optical images, and obtains a four-channel fused feature map through modal normalization and stitching. Then, it enhances the features through cross-modal channel admission control and spatial attention weighting. The enhanced features are input into an improved U-Net network containing dual-path heterogeneous downsampling, gated skip connections, and decoder feature calibration. Through feature extraction, filtering, calibration, reconstruction, pixel classification, and boundary refinement, it achieves highway landslide hazard identification. It preserves terrain structure information, suppresses feature redundancy noise, and can accurately identify landslide targets in complex terrain, improving the accuracy of highway landslide hazard identification and boundary recognition in complex terrain environments.

[0044] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration, step S50, further includes steps S201 to S205: Step S201: The enhanced fused features are input into the dual-path heterogeneous downsampling module. The features are downsampled and adaptively weighted and fused through parallel paths of max pooling and strip pooling to obtain the encoded features that preserve the multi-scale terrain structure.

[0045] It should be noted that step S201 includes: inputting the enhanced fusion features into the dual-path heterogeneous downsampling module to extract shallow semantic features, obtaining initial coding features; distributing the initial coding features to the max pooling path and the strip pooling path, and simultaneously performing downsampling of local features and terrain structure features to obtain first downsampling features and second downsampling features; performing adaptive weight allocation on the first downsampling features and second downsampling features to obtain feature fusion weights; using the feature fusion weights to perform weighted fusion of the first downsampling features and second downsampling features to obtain single-level fusion coding features; and repeating the convolutional downsampling and dual-path fusion operations level by level to integrate the single-level fusion coding features of each level to obtain multi-scale terrain structure preserved coding features.

[0046] Understandably, local features, obtained through max-pooling path downsampling, focus on salient information in local areas of the feature map, such as the edges of individual landslide blocks and fine-grained features like small-scale topographic abrupt changes. These are key features for identifying small-scale landslide targets. Topographic structure features, obtained through strip pooling path downsampling, focus on continuous topographic structure information across the entire feature map, such as the overall slope morphology of the landslide body and the topographic trend along highways. These are core features for identifying the spatial distribution of landslide areas. The first downsampling feature is obtained by downsampling the initial encoded features through max-pooling path downsampling, and its core carries the local salient features of the landslide area. The second downsampling feature is obtained by downsampling the initial encoded features through strip pooling path downsampling, and its core carries the continuous structural features of the terrain, with spatial dimensions consistent with the first downsampling feature.

[0047] Specifically, firstly, the enhanced fused features are input into a dual-path heterogeneous downsampling module. Shallow semantic features are extracted through three 3×3 convolutions, with batch normalization and ReLU activation following each convolution layer. The number of channels is gradually increased and the spatial dimension compressed to obtain initial encoded features, thereby capturing low-level edge and texture information in the fused features. Then, the initial encoded features are sent in parallel to the max-pooling path and the strip pooling path. The max-pooling path uses a 2×2 window with a stride of 2 to downsample and extract the most salient local features. The strip pooling path performs 1×W strip pooling in the horizontal direction and H×1 strip pooling in the vertical direction, then concatenates the features, simultaneously preserving the continuous structural features of the terrain in the dominant direction, resulting in the first and second downsampled features. This dual-path parallel mechanism balances local saliency with terrain continuity. Next, global average pooling is performed on the first and second downsampled features to generate channel descriptors, which are then combined... After concatenating the path descriptors, a nonlinear mapping is learned through a fully connected layer, outputting normalized feature fusion weights. The contribution ratio of the two path features is dynamically adjusted to adapt to different terrain complexities. The first and second downsampled features are weighted and fused using the feature fusion weights. After multiplying the weights with the corresponding path features, the channels are concatenated and dimensionality is reduced by 1×1 convolution to obtain single-layer fusion coding features, achieving adaptive fusion of local details and terrain structure. Finally, the above convolutional downsampling and dual-path fusion operations are repeated level by level. Each level uses the output of the previous level as input to continue extracting deeper abstract features, while maintaining the continuous attention of the strip pooling path to the terrain structure. The single-layer fusion coding features output from each level are integrated to obtain coding features that preserve multi-scale terrain structure. This allows the network to fully preserve multi-scale terrain information from shallow edges to deep semantics while compressing spatial resolution, providing a hierarchical feature foundation for accurate identification of subsequent landslide areas.

[0048] Step S202: The encoded features are subjected to cross-scale attention filtering through a gated skip connection mechanism to obtain effective skip features.

[0049] It should be noted that step S202 includes: stratifying the encoded features by scale to obtain the encoded features to be screened at each level; performing attention-aware evaluation on the encoded features to be screened at each level to generate hierarchical attention weights; using the hierarchical attention weights to weight the encoded features to be screened at the corresponding level to obtain weighted encoded features at each level; performing redundant feature suppression on the weighted encoded features at each level to obtain processed weighted encoded features; and integrating the processed weighted encoded features at each level to obtain effective skip features.

[0050] It's important to understand that scale-based stratification is an operation that divides coded features according to spatial resolution and semantic hierarchy. During the progressive downsampling process, coded features form feature hierarchies at different scales. Scale-based stratification, based on dimensions such as the spatial size and channel dimension of the feature map, decomposes the overall coded features into multiple independent hierarchical features, each corresponding to a scale of coded features to be screened. The coded features to be screened at each level are single-level features obtained after scale-based stratification, and each level possesses specific spatial resolution and semantic information. Attention-based evaluation is an importance assessment operation performed on the coded features to be screened at each level. Through a pre-defined attention evaluation network, the proportion of landslide-related features and feature response intensity in each level's features are analyzed to quantitatively evaluate the contribution of each level's features to landslide disease identification. Redundant feature suppression is a feature purification operation performed on weighted coded features. Through a pre-defined feature screening algorithm, redundant information such as background features and duplicate features unrelated to landslide disease in the weighted coded features is identified and removed, further strengthening the core landslide-related features and improving the discriminative power of the features. Effective jump features are a set of features obtained by integrating weighted encoded features after processing at all levels. They include core landslide features that have been filtered and enhanced at each scale and are key features passed from the encoder to the decoder.

[0051] Specifically, firstly, the encoded features are layered according to the encoder levels. The feature maps output from layers 1 to 4 of the encoder are extracted as the selected encoded features for layers 1 to 4, respectively. Each layer of features has different spatial resolution and semantic abstraction level, thus preserving multi-scale information for subsequent decoding and reconstruction. Then, attention-aware evaluation is performed on the selected encoded features of each layer. The selected encoded features of each layer are concatenated with the corresponding upsampled feature map of the decoder, and a single-channel attention map is generated through 3×3 convolution. Then, sigmoid activation is used to generate hierarchical attention weights, so that the gating signal is dynamically adjusted according to the current decoding state, realizing adaptive selection of encoded features. Next, the hierarchical attention weights are used to weight the selected encoded features of the corresponding layers, and the hierarchical attention weights are then combined with the selected encoded features. Features are multiplied at the pixel level to obtain weighted encoded features at each level, strengthening the feature responses relevant to the current decoding task. Subsequently, redundant feature suppression is performed on the weighted encoded features at each level. By setting an attention threshold of 0.1, pixels with weights below the threshold are set to zero. At the same time, channel sparsity constraints are used to suppress low-activation channels, resulting in processed weighted encoded features that filter out background noise and repetitive information irrelevant to landslide recognition. Finally, the processed weighted encoded features at each level are integrated. The weighted encoded features processed at layers 1 to 4 are convolved with 1×1 to unify the number of channels and then a skip connection is established with the corresponding decoder features to obtain effective skip features. This ensures that only highly relevant features after attention filtering participate in cross-layer transmission, improving the decoder's accuracy in reconstructing landslide boundaries and details.

[0052] Step S203: Align the number of channels and feature distribution of the effective skip features by using a lightweight 1×1 convolution in the decoder feature calibration module to obtain the scale-calibrated fused features.

[0053] Specifically, firstly, the effective skip features are input into the decoder feature calibration module. A 1×1 convolution is used to compress or expand the number of channels in the effective skip features, ensuring that the number of channels in the high-resolution skip features from the encoder is consistent with the number of channels in the upsampled features of the corresponding layer decoder. This resolves the dimensionality mismatch issue caused by differences in encoder and decoder channel settings. Secondly, the skip features with adjusted channel numbers are aligned in terms of feature distribution. Learnable scaling factors and bias terms are used to shift and scale the feature numerical distribution, making its mean and variance close to the statistical distribution of the corresponding layer decoder features. This alleviates fusion conflicts caused by differences in numerical ranges when directly stitching multi-scale features. Finally, the skip features with aligned channel numbers and feature distributions are stitched and fused with the upsampled features from the decoder to obtain scale-calibrated fused features. This ensures that the information transmitted across layers is compatible at the numerical level, improving the stability of feature fusion and the convergence speed of network training, while maintaining the low computational overhead of lightweight 1×1 convolutions to meet the efficiency requirements of large-scale remote sensing image batch processing.

[0054] Step S204: The scale-calibrated fused features are upsampled and reconstructed step by step to obtain a landslide feature map with multi-level feature fusion.

[0055] It should be noted that step S204 includes: performing transposed convolutional upsampling on the scale-calibrated fused features to restore the spatial resolution of the features and obtain single-level upsampled features; concatenating and fusing the single-level upsampled features with the effective skip features of the corresponding level to obtain single-level reconstructed features; integrating all single-level reconstructed features to obtain reconstructed features at each level; performing multi-scale feature fusion on the reconstructed features at each level to integrate shallow details and deep semantic features to obtain a fused feature map; and performing lightweight convolutional detail enhancement on the fused feature map to obtain a landslide feature map with multi-level feature fusion.

[0056] It's important to understand that single-level reconstructed features are obtained by concatenating and fusing single-level upsampled features with corresponding level effective skip features. They simultaneously contain deep semantic features from the decoder and shallow detail features from the encoder, representing the result of feature reconstruction at a single level. Level-specific reconstructed features are sets of reconstructed features across all scale levels obtained by repeatedly performing upsampling and fusion operations with features at the same level. Each level's reconstructed features correspond to a different spatial resolution. Lightweight convolutional detail enhancement is a feature enhancement operation performed on the fused feature map using convolutional kernels with small parameters and low computational cost (such as 1×1 or 3×3 convolutions). Without increasing model complexity, it enhances the detail feature responses in the feature map, improving the feature recognition of landslide boundaries and small-scale targets.

[0057] Specifically, firstly, transposed convolutional upsampling is performed on the scale-calibrated fused features. A 4×4 convolution kernel is used for deconvolution with a stride of 2 to gradually restore the spatial resolution of the feature map to a size matching the corresponding encoder level, resulting in single-level upsampled features. This allows for the gradual reconstruction of spatial details from deep abstract semantics. Next, the single-level upsampled features are concatenated and fused with the effective skip features of the corresponding level along the channel dimension. This merges the high-level semantic information recovered by the decoder with the low-level detail information retained by the encoder, resulting in single-level reconstructed features, achieving complementary enhancement of semantics and details at the same resolution. Finally, all single-level reconstructed features are integrated. The single-level reconstructed features generated iteratively from the deepest to the shallowest layer are stacked in hierarchical order to form a feature map containing multi-resolution information. The system first collects features at each level to construct a complete feature pyramid structure. Then, it performs multi-scale feature fusion on the reconstructed features at each level. Through an adaptive weight allocation mechanism, it integrates the edge details of shallow reconstructed features with the semantic category probabilities of deep reconstructed features to obtain a fused feature map. This comprehensively utilizes the advantages of different levels to improve the perception of small-scale landslide targets. Finally, it performs lightweight convolutional detail enhancement on the fused feature map. It uses two cascaded 3×3 depthwise separable convolutions to extract local details and reduce the number of parameters. After ReLU activation and batch normalization, it obtains a landslide feature map with multi-level feature fusion. This enhances the expressive ability of landslide boundaries and internal textures while maintaining computational efficiency, providing a high-quality feature foundation for subsequent pixel-level classification.

[0058] Step S205: Perform pixel-level classification and boundary refinement on the landslide feature map fused with multi-level features to obtain the identification results of highway landslide disease.

[0059] It should be noted that step S205 includes: performing pixel-level feature mapping on the landslide feature map fused with multi-level features to generate the category probability distribution of each pixel; classifying the category probability distribution according to a preset probability threshold to obtain an initial landslide binary segmentation map; performing edge contour detection on the initial landslide binary segmentation map to extract the initial boundary features of the landslide area; refining the initial boundary features by combining the terrain structure features of the digital elevation model to obtain the accurate landslide boundary; and fusing the accurate landslide boundary with the landslide feature map fused with multi-level features to generate the highway landslide disease identification result.

[0060] It's important to understand that the initial landslide binary segmentation image is a binary image obtained after pixel classification. The image contains only two pixel values, representing the landslide area and the non-landslide area respectively, providing a preliminary indication of the approximate location and extent of the landslide. Edge contour detection is the operation of extracting boundaries from the initial landslide binary segmentation image. This operation identifies the boundary line between the landslide area and the non-landslide area through gradient calculation and contour tracking, obtaining the initial boundary shape. Refinement correction is the operation of optimizing and adjusting the initial boundary features in conjunction with terrain structural characteristics. This operation smooths, fills in, and corrects the initial boundary based on the continuity and rationality of the terrain, eliminating jagged edges, breaks, and misalignments, making the boundary conform to the actual terrain trend.

[0061] Specifically, firstly, pixel-level feature mapping is performed on the landslide feature map fused from multiple levels. A 1×1 convolution reduces the number of feature map channels to the number of categories (landslide and non-landslide). Then, the map is normalized using the Softmax activation function to generate a probability distribution for each pixel belonging to the landslide category, achieving dense prediction. Next, pixel classification is performed based on a preset probability threshold of 0.5. Pixels with probabilities greater than the threshold are marked as landslide areas, and pixels with probabilities less than or equal to the threshold are marked as background areas, resulting in an initial binary landslide segmentation map, completing coarse-grained semantic segmentation. Then, edge contour detection is performed on the initial binary landslide segmentation map. The Canny operator is used to calculate the gradient magnitude and direction, and initial landslide regions are extracted through non-maximum suppression and double threshold filtering. Boundary features are used to identify the transition zone between the landslide and the background. Subsequently, the initial boundary features are refined by combining the topographic structure features of the digital elevation model. The elevation gradient and slope change rate at the initial boundary are calculated. False boundary points with elevation gradients less than the threshold (non-topographic fault zones) are removed, and missing boundary segments with abrupt slope changes but not forming complete closures are added to obtain accurate landslide boundaries consistent with the topographic undulations, thus solving the segmentation errors caused by optical image boundary blurring and noise. Finally, the accurate landslide boundary is used as a mask and fused with the landslide feature map fused with multi-level features. The original feature values ​​are retained within the boundary, while the weights are set to 0 or reduced outside the boundary, generating highway landslide disease identification results containing the precise boundary location, achieving an accurate quantitative expression of the landslide area range, morphology, and spatial distribution.

[0062] This embodiment inputs the enhanced fused features into a dual-path heterogeneous downsampling module, obtains multi-scale encoded features through parallel pooling and adaptive fusion, then achieves cross-scale attention filtering through gated skip connections, completes feature calibration using lightweight 1×1 convolutions, performs pixel classification and boundary refinement after stepwise upsampling reconstruction, and outputs landslide recognition results. This can effectively preserve the terrain structure, suppress redundant noise, and improve the accuracy of multi-scale feature matching and landslide recognition.

[0063] Based on the first embodiment of this application, this application also provides a highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration. Please refer to... Figure 3 The device includes: The acquisition module 10 is used to acquire optical remote sensing images and digital elevation model data aligned to the same geographic area.

[0064] The data augmentation module 20 is used to perform seasonal-weather joint data augmentation on optical remote sensing images to generate enhanced optical images that simulate different seasons and weather conditions.

[0065] The fusion module 30 performs modal scale normalization on the enhanced optical image and the digital elevation model data respectively and then stitches them together to obtain a four-channel fused feature map.

[0066] Processing module 40 sequentially performs cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map to obtain the enhanced fused features.

[0067] The result module 50 is used to input the enhanced fused features into the improved U-Net network to generate highway landslide disease identification results. The improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. It also inputs the enhanced fused features into the dual-path heterogeneous downsampling module, downsampling the features through parallel paths of max pooling and strip pooling and adaptively weighting the fusion to obtain encoded features that preserve multi-scale terrain structure. The gated skip connection mechanism performs cross-scale attention filtering on the encoded features to obtain effective skip features. The lightweight 1×1 convolution in the decoder feature calibration module aligns the effective skip features in terms of channel number and feature distribution to obtain scale-calibrated fused features. The scale-calibrated fused features are then subjected to progressive upsampling and feature reconstruction to obtain a multi-level feature fused landslide feature map. Finally, the multi-level feature fused landslide feature map is subjected to pixel-level classification and boundary refinement to obtain the highway landslide disease identification results.

[0068] The highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration provided in this application adopts the highway landslide hazard identification method based on multi-source image fusion and cross-modal collaboration in the above embodiments, which can solve the technical problem of how to improve the identification accuracy and boundary recognition precision of highway landslide hazards in complex terrain environments. Compared with the prior art, the beneficial effects of the highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration provided in this application are the same as those of the highway landslide hazard identification method based on multi-source image fusion and cross-modal collaboration provided in the above embodiments, and other technical features in the highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0069] This application provides a highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration. The highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration includes: at least one processor; and a memory communicatively connected to at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to execute the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in the above embodiment 1.

[0070] The following is for reference. Figure 4 This document illustrates a structural schematic diagram of a highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration, suitable for implementing embodiments of this application. The highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The illustrated highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0071] like Figure 4As shown, the highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the road landslide hazard identification device based on multi-source image fusion and cross-modal collaboration to exchange data with other devices wirelessly or via wired communication. Although various road landslide hazard identification devices based on multi-source image fusion and cross-modal collaboration are shown in the figures, it should be understood that implementation or possession of all of them is not required. More or fewer devices may be implemented alternatively.

[0072] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0073] The highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration provided in this application adopts the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in the above embodiments, which can solve the technical problem of how to improve the identification accuracy and boundary recognition precision of highway landslide diseases in complex terrain environments. Compared with the prior art, the beneficial effects of the highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration provided in this application are the same as the beneficial effects of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration provided in the above embodiments, and other technical features in the highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration are the same as the features disclosed in the previous embodiment method, and will not be repeated here.

[0074] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0075] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0076] This application provides a computer-readable medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration in the above embodiments.

[0077] The computer-readable medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, or any combination thereof. More specific examples of computer-readable media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable medium may be any tangible medium containing or storing a program that can be executed by instructions, used by a device, or used in conjunction with it. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0078] The aforementioned computer-readable medium may be included in a highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration; or it may exist independently and not be assembled into the highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration.

[0079] The aforementioned computer-readable medium carries one or more programs that, when executed by a highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration, enable the device to write computer program code for performing the operations of this application in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of this application. In this regard, all blocks in the flowcharts or block diagrams may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that all blocks in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using dedicated hardware-based implementations that perform the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0081] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0082] The readable medium provided in this application is a computer-readable medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration. This method can solve the technical problem of how to improve the accuracy of identifying highway landslide hazards and the precision of boundary identification in complex terrain environments. Compared with the prior art, the beneficial effects of the computer-readable medium provided in this application are the same as those of the highway landslide hazard identification method based on multi-source image fusion and cross-modal collaboration provided in the above embodiments, and will not be repeated here.

[0083] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration as described above.

[0084] The computer program product provided in this application can solve the technical problem of how to improve the accuracy of identifying and defining the boundaries of landslide hazards on highways in complex terrain environments. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the landslide hazard identification method based on multi-source image fusion and cross-modal collaboration provided in the above embodiments, and will not be repeated here.

[0085] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration, characterized in that, The method includes: Acquire aligned optical remote sensing images and digital elevation model data for the same geographic region; Seasonal-weather joint data augmentation is performed on the optical remote sensing images to generate enhanced optical images simulating different seasons and weather conditions; Modal scaling is performed on the enhanced optical image and the digital elevation model data respectively, and then the data is stitched together to obtain a four-channel fused feature map. The four-channel fused feature map is sequentially subjected to cross-modal channel admission control and spatial attention weighting to obtain the enhanced fused features; The enhanced fusion features are input into the improved U-Net network to generate highway landslide disease identification results. The improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. The step of inputting the enhanced fusion features into the improved U-Net network to generate highway landslide disease identification results includes: The enhanced fusion features are input into the dual-path heterogeneous downsampling module. The features are downsampled and adaptively weighted and fused through parallel paths of max pooling and strip pooling to obtain coded features that preserve multi-scale terrain structure. The gated skip connection mechanism is used to perform cross-scale attention filtering on the encoded features to obtain effective skip features. The effective skip features are aligned in terms of channel number and feature distribution by a lightweight 1×1 convolution in the decoder feature calibration module to obtain scale-calibrated fused features. The scale-calibrated fused features are upsampled and reconstructed step by step to obtain a landslide feature map with multi-level feature fusion. The landslide feature map fused from the multi-level features is classified at the pixel level and its boundaries are refined to obtain the identification results of highway landslide diseases.

2. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The step of performing seasonal-weather joint data augmentation on the optical remote sensing image to generate enhanced optical images simulating different seasons and weather conditions includes: The baseline seasonal category is determined based on the acquisition time of the optical remote sensing images; The optical remote sensing image is transformed by the spectral characteristics of surface vegetation to simulate the changes in vegetation cover, hue and texture under the baseline seasonal category, thus obtaining a seasonally enhanced optical image. The illumination intensity and atmospheric visibility of the seasonally enhanced optical image are adjusted to simulate the changes in image brightness, contrast and detail recognition under a preset weather type in the baseline seasonal category, thereby obtaining a weather-enhanced optical image. The enhanced optical image of the weather is normalized and calibrated to obtain a calibrated enhanced optical image; The calibrated enhanced optical image is fused with the original optical remote sensing image at the feature layer to obtain a fused enhanced optical image. An edge-preserving filtering algorithm is used to enhance the details of the fused enhanced optical image to obtain an enhanced optical image.

3. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The step of sequentially performing cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map to obtain the enhanced fused features includes: The four-channel fused feature map is input into the cross-modal channel admission control module based on the extrusion structure. A first weight and a second weight set are generated through a weight network of preset dimensions. The first weight is the optical three-channel weight, and the second weight is the digital elevation model single-channel weight. The first and second weights are subjected to modal grouping normalization and then integrated to obtain the calibration channel weight set. The four-channel fusion feature map is subjected to channel-level weighted filtering using the calibration channel weight set to obtain a channel-enhanced fusion feature map. A local spatial convolution operation is performed on the channel enhancement fusion feature map using a convolution kernel of a preset size to generate a spatial attention weight map; The spatial attention weight map and the channel enhancement fusion feature map are multiplied by pixel-level weighting to obtain the enhanced fusion feature.

4. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The step of inputting the enhanced fused features into the dual-path heterogeneous downsampling module, downsampling the features through parallel paths of max pooling and strip pooling, and adaptively weighting and fusing them to obtain the encoded features that preserve multi-scale terrain structure includes: The enhanced fusion features are input into the dual-path heterogeneous downsampling module to extract shallow semantic features, and the initial encoded features are obtained. The initial encoded features are distributed to the max pooling path and the strip pooling path, and local features and terrain structure features are downsampled simultaneously to obtain the first downsampled features and the second downsampled features. Adaptive weight allocation is performed on the first downsampled feature and the second downsampled feature to obtain the feature fusion weight; The first downsampled feature and the second downsampled feature are weighted and fused using the feature fusion weights to obtain a single-level fused coding feature; By repeatedly performing convolutional downsampling and dual-path fusion operations at each level, the single-level fusion coding features of each level are integrated to obtain multi-scale terrain structure-preserving coding features.

5. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The step of performing cross-scale attention filtering on the encoded features through the gated skip connection mechanism to obtain effective skip features includes: The encoded features are stratified by scale to obtain the encoded features to be screened at each level; Attention-aware evaluation is performed on the encoded features to be screened at each level to generate hierarchical attention weights; The hierarchical attention weights are used to weight the coding features to be screened at the corresponding level to obtain the weighted coding features at each level. Redundancy suppression is performed on the weighted coding features at each level to obtain the processed weighted coding features; By integrating the weighted coding features processed at each level, an effective skip feature is obtained.

6. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The step of performing stepwise upsampling and feature reconstruction on the scale-calibrated fused features to obtain a landslide feature map with multi-level feature fusion includes: Perform transposed convolutional upsampling on the scale-calibrated fused features to restore the spatial resolution of the features and obtain single-level upsampled features; The single-level upsampling features are concatenated and fused with the effective skip features of the corresponding level to obtain single-level reconstructed features; Based on the integration of all the single-level reconstruction features, the reconstruction features of each level are obtained; Multi-scale feature fusion is performed on the reconstructed features at each level to integrate shallow details and deep semantic features, resulting in a fused feature map; Lightweight convolutional detail enhancement is performed on the fused feature map to obtain a landslide feature map with multi-level feature fusion.

7. The method for identifying highway landslide hazards based on multi-source image fusion and cross-modal collaboration as described in claim 1, characterized in that, The steps of performing pixel-level classification and boundary refinement on the landslide feature map fused with the multi-level features to obtain the highway landslide disease identification results include: Pixel-level feature mapping is performed on the landslide feature map fused from the multi-level features to generate the category probability distribution of each pixel; The probability distribution of the categories is classified into pixels according to a preset probability threshold to obtain an initial binary segmentation map of the landslide. Edge contour detection is performed on the initial binary segmentation map of the landslide to extract the initial boundary features of the landslide area; The initial boundary features are refined by combining the terrain structure features of the digital elevation model to obtain an accurate landslide boundary; The precise landslide boundary and the landslide feature map fused with the multi-level features are combined to generate the highway landslide disease identification result.

8. A highway landslide disease identification device based on multi-source image fusion and cross-modal collaboration, characterized in that, The device includes: The acquisition module is used to acquire optical remote sensing images and digital elevation model data aligned to the same geographic area; The data augmentation module is used to perform seasonal-weather joint data augmentation on the optical remote sensing image to generate enhanced optical images that simulate different seasons and weather conditions; The fusion module performs modal scale normalization on the enhanced optical image and the digital elevation model data respectively and then stitches them together to obtain a four-channel fused feature map. The processing module sequentially performs cross-modal channel admission control and spatial attention weighting on the four-channel fused feature map to obtain the enhanced fused features; The result module is used to input the enhanced fused features into the improved U-Net network to generate highway landslide disease identification results. The improved U-Net network includes a dual-path heterogeneous downsampling module, a gated skip connection mechanism, and a decoder feature calibration module. It is also used to input the enhanced fused features into the dual-path heterogeneous downsampling module, downsample the features through parallel paths of max pooling and strip pooling, and adaptively weight and fuse them to obtain multi-scale terrain structure-preserving encoded features. The gated skip connection mechanism performs cross-scale attention filtering on the encoded features to obtain effective skip features. The lightweight 1×1 convolution in the decoder feature calibration module aligns the effective skip features in terms of channel number and feature distribution to obtain scale-calibrated fused features. The scale-calibrated fused features are then subjected to progressive upsampling and feature reconstruction to obtain a multi-level feature fused landslide feature map. Finally, the multi-level feature fused landslide feature map is subjected to pixel-level classification and boundary refinement to obtain the highway landslide disease identification results.

9. A highway landslide hazard identification device based on multi-source image fusion and cross-modal collaboration, characterized in that, The device includes: a memory, a processor, and a highway landslide disease identification program based on multi-source image fusion and cross-modal collaboration, stored in the memory and running on the processor, wherein the highway landslide disease identification program based on multi-source image fusion and cross-modal collaboration is configured to implement the steps of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration as described in any one of claims 1-7.

10. A storage medium, characterized in that, The storage medium stores a highway landslide disease identification program based on multi-source image fusion and cross-modal collaboration. When the processor executes the highway landslide disease identification program based on multi-source image fusion and cross-modal collaboration, it implements the steps of the highway landslide disease identification method based on multi-source image fusion and cross-modal collaboration as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent landslide identification method based on multi-source remote sensing image deep learning

    CN120219970A

  • Landslide identification method based on mixed attention mechanism of channel and space

    CN121392575A

  • Lightweight satellite landslide image intelligent detection method, apparatus and device, and medium

    CN121438136A