Landslide extraction method based on multi-task learning and multi-source data feature fusion

By adopting multi-task learning and multi-source data feature fusion methods in landslide extraction, the BFM-UNet model combined with optical remote sensing images and topographic data, the problem of low landslide extraction accuracy and efficiency in the existing technology is solved, and more efficient landslide identification and data utilization are achieved.

CN120014481AActive Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510092277.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The existing landslide extraction methods have shortcomings in data processing and feature extraction, especially in small sample learning and multi-source data fusion, resulting in low extraction accuracy and efficiency.

Method used

A landslide extraction method based on multi-task learning and multi-source data feature fusion is adopted, and optical remote sensing images and terrain data are combined through the BFM-UNet model, and auxiliary tasks are introduced to enhance the model's feature mining capabilities and data utilization.

Benefits of technology

It improves the accuracy and efficiency of landslide extraction, can make more efficient use of information in multi-source data, reduce dependence on high-quality data, and enhances the performance of the model under small samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014481A_ABST
    Figure CN120014481A_ABST
Patent Text Reader

Abstract

The invention relates to a landslide extraction method based on multi-task learning and multi-source data feature fusion, and belongs to the field of remote sensing image processing. The method comprises the following steps: firstly, acquiring multi-source data containing landslide characteristics, and preprocessing the multi-source data; then, the preprocessed data are input into a defined BFM-UNet model, and the model comprises a branch fusion encoder and a multi-task decoder; the branch fusion encoder extracts features of different source data through two branches, and fuses the features through the cross-stitch unit, thereby effectively utilizing information in multi-source data. The multi-task decoder increases recognition tasks of non-landslide ground objects which widely exist in the target area and are easy to recognize to assist landslide extraction, and the advantages of multi-task learning are exerted through optimal sharing among learning tasks of the cross-stitch unit. And finally, extracting the landslide by using the trained model to obtain a landslide distribution diagram. Experimental results show that the landslide extraction method provided by the invention can effectively improve the precision of landslide extraction and reduce the false drop rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of remote sensing image processing and relates to a landslide extraction method based on multi-task learning and multi-source data feature fusion. Background Art

[0002] Landslides are a high-frequency natural disaster with strong explosive power and destructiveness, causing serious damage to people's lives and property. The rapid extraction of landslide information is of great significance to emergency response to disasters and land use planning. Although the traditional landslide extraction method based on field investigation and measurement has high mapping accuracy and detailed geological information, it is time-consuming and labor-intensive, and there is still a certain degree of danger.

[0003] With the development of remote sensing technology, remote sensing data has been widely used in landslide extraction. Visual interpretation is a basic landslide extraction method based on remote sensing images. This method relies on the professional knowledge and experience of the interpreter to identify the landslide area. It has high accuracy but is highly subjective and difficult to process large-scale data. In addition, the automatic landslide extraction methods based on remote sensing can be roughly divided into feature threshold-based, shallow machine learning-based and deep learning-based methods according to the technical principles. Among them, the feature threshold-based method constructs a feature rule set by determining a suitable threshold to achieve landslide extraction. However, the suitable threshold is difficult to find and has low universality. The machine learning-based method uses machine learning algorithms for classification to achieve landslide extraction. This type of method has low demand for data and computing resources and high computational efficiency, but it is more dependent on feature selection, and its performance is also limited by the expression ability of the machine learning model. In contrast, the deep learning-based method shows great potential and advantages. UNet has a wide range of applications in landslide extraction based on remote sensing images due to its simple, flexible and easily scalable architecture design and wide applicability. In the application scenario of landslide extraction from remote sensing images, UNet is usually modified and optimized in a targeted manner. The deep learning method can automatically extract features from raw data, achieving end-to-end landslide extraction, which is more efficient when conducting large-scale landslide detection.

[0004] Deep learning methods have shown great potential and advantages in landslide extraction, but the particularity of landslides makes landslide identification more difficult, and existing methods still face some challenges. Due to its complex structure and large number of parameters, the good performance of deep learning models depends on a large number of high-quality data samples. However, high-quality landslide data samples are scarce, which makes the model face the problem of small sample learning and limits its performance. In addition, previous studies often only focus on the characteristics of the landslide itself, and simply regard the non-landslide area as the background, thereby ignoring the information that the background area may contain and is helpful for landslide identification. In addition, due to the similarity in texture and color between the slope area and the surrounding bare soil, rocks, etc., the irregular shape, and the interference of light, shadow, cloud and vegetation on the image quality, the effect of using only optical remote sensing images for landslide extraction is not ideal; although some existing studies have considered combining multi-source data for landslide extraction, simply splicing or adding data together often cannot fully utilize the unique information of each data type, nor can it deeply explore the potential correlation between data.

[0005] In response to the above problems, the optimization method is based on combining multi-task learning and making full use of multi-source data. On the one hand, the recognition task of non-landslide features that are widely present in the target area and easier to identify than landslides is added to assist landslide extraction, so that the model can more effectively use the feature information in the background (non-landslide area) when extracting landslides. At the same time, multi-task learning is used to enhance the feature mining ability of the network and reduce the dependence on data. On the other hand, the network structure is optimized to effectively extract and fuse the feature information in multi-source data, ensuring that the model can fully utilize the information contained in multi-source data when extracting landslides, further enhancing the model performance. Summary of the invention

[0006] In view of this, the object of the present invention is to provide a landslide extraction method based on multi-task learning and multi-source data feature fusion.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A landslide extraction method based on multi-task learning and multi-source data feature fusion includes the following steps:

[0009] S1: Data acquisition, including relevant multi-source data and vector label data, where the vector label data includes landslide labels and labels related to auxiliary tasks;

[0010] S2: Preprocess the acquired data, where the relevant multi-source data mainly needs to eliminate noise and interference as much as possible, and the relevant label data needs to be converted from vector to raster. Finally, it is necessary to ensure that all data have uniform resolution and are spatially aligned;

[0011] S3: Construct a multi-task deep learning model BFM-UNet (Branch-Fusion Multi-task UNet) for landslide extraction using multi-source data. The model includes a branch fusion encoder and a multi-task decoder.

[0012] S4: Training the model, including the calculation of the loss function and the setting of related hyperparameters. During the training process, for each task, a joint loss function is used, which combines the weighted cross entropy loss and the Dice loss, and a learnable dynamically changing task loss weight is used in multi-task learning. In addition, the appropriate batch size and training rounds are set, SGD (Stochastic Gradient Descent) is selected as the optimization algorithm, and a dynamically changing learning rate is used;

[0013] S5: Based on the optical remote sensing image and terrain data of the target area pre-processed by S2, the landslides are extracted using the trained BFM-UNet model to obtain a landslide distribution map.

[0014] Furthermore, in S1, the multi-source data to be acquired mainly include optical remote sensing images and digital elevation models, and other remote sensing data containing landslide characteristic information, such as SAR images, are added as appropriate; in addition, landslide vector labels and auxiliary task vector labels need to be acquired, wherein the auxiliary tasks select the objects that are widely present in the area and easier to identify than landslides as targets.

[0015] Furthermore, in S2, the preprocessing of multi-source data mainly includes: atmospheric correction of optical remote sensing images, exporting all or only some channels of data as required, such as RGB three channels; terrain analysis based on DEM data to obtain slope and aspect data; if there are other data, such as SAR images, radiation correction, denoising and terrain correction are required; to ensure the uniformity of the resolution of multi-source data, some of the data need to be resampled, generally with the highest resolution as the reference standard, the higher the resolution, the more conducive to landslide extraction; in addition, the landslide vector label and the auxiliary task vector label are converted from vectors to raster images with the same resolution as other data; finally, image registration is performed to ensure that the multi-source data and the corresponding label data are aligned in space.

[0016] Furthermore, in S3, the multi-task deep learning model BFM-UNet includes:

[0017] Branch fusion encoder, which extracts features from different source data through multiple branches, usually including branches for optical remote sensing images and terrain data. More branches are added as needed to process other data containing landslide features, such as SAR images. Richer multi-source data helps to expand feature expression and enable the model to cope with complex environments and changing geographical features more effectively. The features extracted by each branch are fused through the cross-stitch unit and then passed to the decoder for further processing;

[0018] The multi-task decoder, by introducing other ground feature recognition tasks besides landslide extraction, assists landslide extraction, enabling the model to more effectively utilize feature information in the background. At the same time, it uses multi-task learning to enhance the model's feature mining capabilities, reduce dependence on data, and utilize the best sharing between cross-stitch unit learning tasks, thereby better leveraging the advantages of multi-task learning.

[0019] Furthermore, in S3, both the branch fusion encoder and the multi-task decoder include a cross-stitch unit. The original cross-stitch unit is mainly used for adaptive sharing between different tasks. Here, its functional application is extended to use it for multi-source data feature fusion. In order to meet the flexibility required for multi-source data feature fusion, the cross-stitch unit is extended to no longer limit the number of input and output features, allowing them to be any value and not necessarily the same. In addition, considering the differences in different channels of the feature map, the cross-stitch unit is also improved to expand the weight elements in the original weight matrix from scalars to vectors with the same dimension as the number of feature map channels. The processing of the feature map is calculated according to formula (1):

[0020]

[0021] in, represents the i-th output feature map, F j , j = 1, 2, ..., m represents the jth input feature map; w ij It means that when calculating the i-th output feature map, the j-th input feature map corresponds to a weight. The weight is a vector whose dimension is the same as the number of channels of the input feature map, that is, each channel in the feature map corresponds to an element in the corresponding weight vector. The calculation of the i-th output feature map is shown in formula (2):

[0022]

[0023] The operation is performed on a channel basis, that is, F j Each element in the k-th channel feature map is related to the vector w ij Multiply the corresponding k-th element in .

[0024] Furthermore, in S3, the branch fusion encoder adopts a multi-branch structure to extract features from different source data respectively. The branch trunk consists of 1 input block and 4 downsampling blocks. Compared with the original UNet, the convolutional layer in the downsampling block is replaced by the residual block; the input and output of the input block are defined as and The height and width of the output feature map are the same as the input image, both are H in and W in , the number of channels is C data Convert to C in , C data is the number of channels of input data, C in is an adjustable parameter, usually set to 32 or 64. The processing of the input block is expressed as formula (3):

[0025] Y in =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X in )))))) (3)

[0026] Among them, ReLU is the activation function, BN means batch normalization, Conv 3×3 represents a 3×3 convolution operation; the processing process of the downsampling block on the feature map is expressed as formula (4):

[0027] Y d =Res(Res(MaxPool(X d ))) (4)

[0028] in, and Respectively represent the input feature map and output feature map of the downsampling block, H d ×W d ×C d is the shape of the input feature map, Res is the residual block, MaxPool is the maximum pooling operation with a window size of 2×2 and a step size of 2. After being processed by the downsampling block, the height and width of the feature map are reduced by half, while the number of channels is doubled; the processing of the feature map by the residual block is expressed by formula (5):

[0029] Y res =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X res ))))))+X res ) (5)

[0030] Among them, X res and Y res Represent the feature maps of input and output respectively;

[0031] The cross-stitch unit is applied at the jump connection and the bottom connection of the encoder and decoder to learn the association between multi-source data to more effectively fuse features and pass the fusion results to the decoder. At these locations, the number of input features and the number of output features of the cross-stitch unit correspond to the number of feature extraction branches in the encoder and the number of task branches in the decoder, respectively. The weight vector in the weight matrix is ​​initialized as shown in formula (6):

[0032]

[0033] Among them, N represents the number of branches in the encoder, and C is the dimension of the weight vector. Through this setting, the features extracted by each branch have the same weight at the beginning and are equally integrated.

[0034] Furthermore, in the S3, the multi-task decoder expands additional branches based on the original UNet decoder to identify suitable non-landslide features, and uses multi-task learning to enable the model to make better use of the information in the background. Each branch trunk consists of 4 upsampling blocks and an output block. Compared with the original UNet, the convolutional layer in the upsampling block is replaced by a residual block, and the output block of the auxiliary task branch is only enabled during model training; each upsampling block has two inputs, which are the feature maps from the last downsampling block in the encoder or the previous upsampling block. and the feature maps from the corresponding layers of the encoder via skip connections The output after upsampling block processing is Among them, H u ×W u ×C u It represents the shape of the feature map from the bottom downsampling block or the previous upsampling block. The processing of the upsampling block is as shown in formula (7):

[0035] Y u =Res(Res(Cat(X u2, Conv T (X u1 )))) (7)

[0036] Res represents the residual block, calculated according to formula (5), Cat represents the feature map according to the channel dimension, Conv T represents a transposed convolution; the output block consists of a 1×1 convolution whose input feature map has the shape of H out ×W out ×C in , the height and width of the output result remain unchanged, consistent with the height and width of the model input image, and the number of channels is 2, corresponding to the number of classification categories of each branch;

[0037] A cross-stitch unit is added after each upsampling block of each branch for feature sharing. This unit will learn the best sharing between tasks and further give play to the advantages of the multi-task structure. The number of its input and output features corresponds to the number of task branches. When initializing the parameters of this unit, the element values ​​of the weight vector on the diagonal of the weight matrix are all 1, and the element values ​​of the weight vector on the off-diagonal are set to all 0, so that the task branches are independent of each other at the beginning, as shown in formula (8):

[0038]

[0039] Among them, 1 C represents a vector of dimension C with all elements being 1, 0 C represents a vector of dimension C with all elements being 0.

[0040] Furthermore, in S4, each task corresponds to a joint loss function L, which combines the weighted cross entropy loss L CE and Dice loss L Dice , L is calculated according to formula (9):

[0041] L=L CE +L Dice (9)

[0042] Weighted cross entropy loss is an extension of cross entropy loss. Weights are introduced to deal with the problem of class imbalance. Its calculation formula is:

[0043]

[0044] Among them, B represents the total number of samples in a batch, C represents the total number of categories, and w j is the weight of category j, y ij is the one-hot encoded value of the true label of sample i in category j. If sample i is category j, then y ij is 1, otherwise it is 0. is the probability that the model predicts that sample i is of category j; the category weight is set according to the degree of category imbalance, using the common method based on inverse frequency, and the calculation formula is:

[0045]

[0046] Among them, N j is the total number of samples of category j; Dice loss is defined based on the Dice coefficient. Dice loss optimizes model performance by maximizing the Dice coefficient. Its calculation formula is:

[0047]

[0048] Among them, p ij and gij Respectively indicate whether sample i belongs to category j in the prediction result and the true label. If so, the value is 1, otherwise it is 0;

[0049] The above loss calculation is for a single task. When training for multi-task learning, the learnable dynamically changing task loss weights are used to learn the relative importance of different tasks from the data. The total loss L of multi-task learning is MT The calculation formula is:

[0050]

[0051] Among them, T represents the total number of tasks, L i is the loss of task i, calculated by formula (9), s i is a learnable parameter based on which a learnable dynamic loss weight is implemented.

[0052] Further, in S4, a data set is constructed based on the data obtained in S1 and processed according to S2, and the data is divided into a training set and a test set, and the training set is enhanced by rotation, vertical flipping and horizontal flipping. The training set is used for BFM-UNet model training, and the test set is used to evaluate the performance of the trained model. When training the model, a suitable batch size and training rounds are set, SGD is selected as the optimization algorithm, and a dynamically changing learning rate is used. The change of the learning rate is divided into two stages:

[0053] In the first stage, at the beginning of training, the learning rate is gradually increased linearly from a very low value to the set value to reduce the impact of the model's sensitivity to the initial weights. It is calculated according to formula (14):

[0054]

[0055] Among them, α is an adjustable parameter, i is the current number of iterations in the first stage, and num phase1 is the total number of iterations in this stage, lr init is the initial learning rate;

[0056] In the second stage, the learning rate gradually decays to close to 0 in a nonlinear manner, which helps the model to converge stably. The specific learning rate is calculated according to formula (15):

[0057]

[0058] Among them, β is an adjustable parameter, j is the current number of iterations in the second stage, and num phase2 is the total number of iterations in this phase.

[0059] Further, in S5, the multi-source data of the target area obtained by S1 and processed according to S2 is input into the trained multi-task deep learning model BFM-UNet for landslide extraction, thereby obtaining a corresponding landslide distribution map, in which the landslide area is displayed in white and the non-landslide area is displayed in black.

[0060] The beneficial effects of the present invention are:

[0061] (1) The multi-task decoder of the present invention assists landslide extraction by adding the recognition task of non-landslide features that are widely present in the target area and easier to identify than landslides through multi-task learning, so that the model can more effectively utilize background information, thereby improving the accuracy of landslide extraction.

[0062] (2) The multi-task decoder of the present invention utilizes multi-task learning to enhance the feature mining capability of the network, so that the model can mine more valuable features under limited samples, improve data utilization, and effectively deal with the small sample learning problem.

[0063] (3) The branch fusion encoder of the present invention extracts the features of multi-source data through different branches and uses the cross-stitch unit to learn the association between data, effectively extracting and fusing the feature information in the multi-source data, ensuring that the model can fully utilize the information contained in the multi-source data when performing landslide extraction.

[0064] (4) The present invention further improves the accuracy of landslide extraction through the collaborative work of the branch fusion encoder and the multi-task decoder, and can better distinguish landslides from other landforms, thereby reducing the false detection rate and improving the reliability of landslide extraction results.

[0065] (5) The framework of the landslide extraction method proposed in the present invention can be adjusted according to specific task requirements, has good adaptability and scalability, and can provide effective technical support for landslide monitoring, disaster warning, and disaster prevention and mitigation in complex geological environments.

[0066] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0068] Figure 1 This is the technical route of the present invention;

[0069] Figure 2 It is the BFM-UNet architecture;

[0070] Figure 3 Extract results and ground-truth labels for landslides. DETAILED DESCRIPTION

[0071] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0072] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0073] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0074] Technical route such as Figure 1As shown. First, the required data is obtained. Optical remote sensing images and DEM data are used as the basis for landslide extraction. In addition, landslide labels and labels of auxiliary tasks are used for model training; the auxiliary tasks select widely distributed and identified objects in the target area as targets, such as vegetation, water bodies, buildings, etc., and vegetation is taken as an example here. Subsequently, the acquired data is preprocessed. Optical remote sensing images need to be atmospherically corrected to obtain images that are closer to the actual reflection of the surface, and the data of the three RGB channels are exported; terrain analysis is performed based on DEM data to obtain terrain factors (slope and slope aspect); related label data needs to be converted from vector to raster. The above data processed separately are unified in resolution and image registration so that they have pixel-level correspondence in the same coordinate system. Then, the processed optical remote sensing images (RGB) and terrain data (elevation, slope and slope aspect) are used as inputs of the BFM-UNet model, and the model is trained in combination with the processed related label data. Finally, the processed optical remote sensing image (RGB) and terrain data (elevation, slope and aspect) are input into the trained model to obtain the landslide distribution map of the corresponding area.

[0075] In this embodiment, the BFM-Unet network architecture is as follows Figure 2 As shown in the figure. The branch fusion encoder adopts a dual-branch structure to extract features from 3-channel optical remote sensing images (Red, Green, Blue) and 3-channel terrain data (elevation, slope, and slope direction), and then fuses the features through the cross-stitch unit. The cross-stitch unit can learn the association between multi-source data, which is expressed as the corresponding weight. In the multi-task decoder, the recognition tasks of other objects (here vegetation is selected) besides landslide extraction are added to assist landslide extraction, and the optimal sharing between tasks is learned through the cross-stitch unit, which is also expressed as the corresponding weight.

[0076] (1) Branch Fusion Encoder

[0077] The encoder of the model adopts a dual-branch structure to extract features from 3-channel optical remote sensing images (Red, Green, Blue) and 3-channel terrain data (elevation, slope, and slope direction), respectively. Subsequently, the features of the two branches are fused through a cross-stitch unit, and the fusion result is passed to the decoder. Figure 1 The left side shows the specific structure of the encoder, where the trunks of the two branches of the encoder are composed of 1 input block and 4 downsampling blocks. The input block is mainly composed of two convolutional layers, and the input and output of the input block are defined as and H in and W in are the height and width of the input image respectively. The height and width of the output feature map are the same as the input image, and the number of channels is converted to 32. The processing of the input block can be expressed as formula (1):

[0078] Y in =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X in )))))) (1)

[0079] Among them, ReLU is the activation function, BN means batch normalization, Conv 3×3 represents a 3×3 convolution operation. The downsampling block contains a maximum pooling operation and two consecutive residual blocks. Its processing of the feature map can be expressed as formula (2):

[0080] Y d =Res(Res(MaxPool(X d ))) (2)

[0081] in, and Respectively represent the input feature map and output feature map of the downsampling block, H d ×W d ×C d is the shape of the input feature map, Res is the residual block, and MaxPool is the maximum pooling operation with a window size of 2×2 and a step size of 2. This operation reduces the height and width of the feature map by half. After processing by the downsampling block, the height and width of the feature map are reduced by half, while the number of channels is doubled. The processing of the feature map by the residual block can be expressed by formula (3):

[0082] Y res =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X res ))))))+X res ) (3)

[0083] Among them, X res and Y res Represent the input and output feature maps respectively. In addition, the cross-stitch unit is applied at the jump connection and the bottom connection of the encoder and decoder to learn the association between multi-source data to more effectively fuse features. At these locations, the number of input features and the number of output features of the cross-stitch unit correspond to the number of feature extraction branches in the encoder and the number of task branches in the decoder, both of which are 2 here. For the cross-stitch unit in the encoder, the weight vector in the weight matrix is ​​initialized as shown in formula (4):

[0084]

[0085] Where N is the number of branches in the encoder, C is the dimension of the weight vector, where i = 1, 2, j = 1, 2, and N = 2. This setting allows the features extracted by each branch to have the same weight at the beginning and can be equally integrated.

[0086] (2) Multi-task decoder

[0087] Based on the basic structure of UNet, additional branches are extended in the decoder to identify suitable non-landslide features, and the optimal sharing between cross-stitch units is utilized to learn tasks, thereby constructing a multi-task network for landslide extraction, so that the model can also make better use of the information in the background (non-landslide area). Figure 1 The structure of the decoder is shown on the right. The two branches correspond to the extraction tasks of landslide and non-landslide features (vegetation is used as an example here); the modules in the red dashed box are only used in the model training phase and do not need to be enabled in the actual landslide extraction application. Each branch trunk consists of 4 upsampling blocks and an output block. Each upsampling block has two inputs, which are the feature maps from the bottom downsampling block or the previous upsampling block. and the feature maps from the corresponding layers of the encoder via skip connections The output after upsampling block processing is Among them, H u ×W u ×C u Represents the shape of the feature map from the bottommost downsampling block or the previous upsampling block. The processing of the upsampling block is as shown in formula (5):

[0088] Y u =Res(Res(Cat(X u2 ,Conv T (X u1 )))) (5)

[0089] Res represents the residual block, and the processing method is shown in formula (3). Cat represents the feature map according to the channel dimension, Conv T Denotes transposed convolution. After the transposed convolution, the feature map is doubled in height and width, and the number of channels is halved. The output block of each branch in the decoder consists of a 1×1 convolution, and the shape of its input feature map is H out ×W out×32, the height and width of the output result remain unchanged, consistent with the height and width of the model input image, and the number of channels is 2, corresponding to the number of classification categories of each branch. Here, each branch is a binary classification task. In the decoder, in addition to utilizing the implicit mutual influence between different task branches in the multi-task network structure, a cross-stitch unit is added after each upsampling block of each branch for feature sharing, so that different task branches can explicitly influence each other, and the cross-stitch unit can learn the optimal sharing between tasks, further exerting the advantages of the multi-task structure. The role of the cross-stitch unit here is to realize adaptive feature sharing of different task branches, so the number of its input and output features corresponds to the number of task branches, both of which are 2. When initializing the parameters of the unit, the element values ​​of the weight vector on the diagonal of the weight matrix are all 1, and the element values ​​of the weight vector on the off-diagonal are set to all 0, so that each task branch is independent of each other at the beginning, as shown in formula (6):

[0090]

[0091] Among them, 1 C represents a vector of dimension C with all elements being 1, 0 C represents a vector of dimension C with all elements equal to 0, where i = 1, 2 and j = 1, 2.

[0092] (3) Loss Function

[0093] For each task, a joint loss function L is used, which combines the weighted cross entropy loss L CE and Dice loss L Dice , L is calculated according to formula (7):

[0094] L=L CE +L Dice (7)

[0095] Weighted cross entropy loss is an extension of cross entropy loss. Weights are introduced to deal with the problem of class imbalance. Its calculation formula is:

[0096]

[0097] Among them, B represents the total number of samples in a batch (here refers to the pixels in the image), C represents the total number of categories, and w j is the weight of category j, y ij is the one-hot encoded value of the true label of sample i in category j. If sample i is category j, then y ij is 1, otherwise it is 0. is the probability that the model predicts that sample i is of category j; the category weight is set according to the degree of category imbalance, using the common method based on inverse frequency, and the calculation formula is:

[0098]

[0099] Among them, N j is the total number of samples of category j. Dice loss is defined based on the Dice coefficient. Dice loss optimizes model performance by maximizing the Dice coefficient. Its calculation formula is:

[0100]

[0101] Among them, p ij and g ij They respectively indicate whether sample i belongs to category j in the predicted result and the true label. If so, the value is 1; otherwise, it is 0.

[0102] The above loss calculation is for a single task. In multi-task learning, a learnable dynamically changing task loss weight is used to learn the relative importance of different tasks from the data. In this way, the total loss L of multi-task learning is MT The calculation formula is:

[0103]

[0104] Among them, T represents the total number of tasks, L i is the loss of task i, calculated by formula (7), s i is a learnable parameter based on which a learnable dynamic loss weight is implemented.

[0105] (4) Hyperparameter setting

[0106] When training the model, the batch size is set to 4, the training rounds are 300, SGD is selected as the optimization algorithm and a dynamically changing learning rate is used. The change of the learning rate is divided into two stages. In the first stage, at the beginning of training, the learning rate is gradually increased linearly from an extremely low value to the set learning rate (0.01) to reduce the impact of the model's sensitivity to the initial weights. It is calculated according to formula (12):

[0107]

[0108] Among them, α is 0.001, i is the current iteration number in the first stage (starting from 1), num phase1 is the total number of iterations in this stage, lr init is the initial learning rate (i.e. 0.01); in the second stage, the learning rate gradually decays to close to 0 in a nonlinear manner, which helps the model to converge stably. The specific learning rate is calculated according to formula (13):

[0109]

[0110] Among them, β is 0.9, j is the current iteration number in the second stage (starting from 1), num phase2 is the total number of iterations in this stage; the first training round is the first stage, and the remaining training rounds constitute the second stage, and each training round contains 199 iterations.

[0111] A dataset is constructed based on the existing labeled data, and the training set and test set are divided into training set and test set in a ratio of 7:3. The training set is enhanced by rotation, vertical flipping and horizontal flipping. The training set is used for BFM-UNet model training.

[0112] (5) Landslide extraction

[0113] The three-channel optical remote sensing images (Red, Green, Blue) and three-channel terrain data (elevation, slope, and slope aspect) of the target area are input into the trained multi-task deep learning model BFM-UNet for landslide extraction, thereby obtaining the corresponding landslide distribution map. In the distribution map, the landslide area is displayed in white, and the non-landslide area is displayed in black.

[0114] In this embodiment, the above method is applied to extract landslides in Jiuzhaigou area. Figure 3 Several results of this method and the corresponding true labels are shown, and it can be seen that the two are very close.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A landslide extraction method based on multi-task learning and multi-source data feature fusion, characterized by: The following steps are involved: S1: Data acquisition, including relevant multi-source data and vector label data, where the vector label data includes landslide labels and labels related to auxiliary tasks; S2: Preprocess the acquired data, where the relevant multi-source data needs to eliminate noise and interference, and the relevant label data needs to be converted from vector to raster. Finally, it is necessary to ensure that all data have uniform resolution and are spatially aligned; S3: Construct a multi-task deep learning model BFM-UNet (Branch-Fusion Multi-task UNet) for landslide extraction using multi-source data. The model includes a branch fusion encoder and a multi-task decoder. S4: Training the model, including the calculation of the loss function and the setting of related hyperparameters. During the training process, for each task, a joint loss function is used, combining weighted cross entropy loss and Dice loss, and learnable dynamically changing task loss weights are used in multi-task learning. In addition, the appropriate batch size and training rounds are set, SGD (Stochastic Gradient Descent) is selected as the optimization algorithm, and a dynamically changing learning rate is used; S5: Based on the optical remote sensing images and terrain data of the target area preprocessed by S2, the landslides are extracted using the trained BFM-UNet model to obtain the landslide distribution map.

2. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S1, the multi-source data to be acquired include optical remote sensing images and digital elevation models, and SAR images are added in a timely manner according to actual needs; landslide vector labels and auxiliary task vector labels are also required, wherein the auxiliary task selects non-landslide features in the area as targets.

3. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S2, the preprocessing of multi-source data includes: performing atmospheric correction on the optical remote sensing image, and exporting data of all or only part of the channels as required; performing terrain analysis based on DEM data to obtain slope and aspect data; if there is a SAR image, performing radiation correction, denoising and terrain correction; to ensure the uniformity of the resolution of the multi-source data, taking the data with the highest resolution as the standard, resampling the data with a resolution lower than the highest resolution; converting the landslide vector label and the auxiliary task vector label from vector to a raster image with a resolution consistent with other data; finally, performing image registration to ensure that the multi-source data and the corresponding label data are aligned in space.

4. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S3, the multi-task deep learning model BFM-UNet includes: The branch fusion encoder extracts features of different source data through multiple branches, including branches for optical remote sensing images and terrain data. More branches are added as needed to process SAR images. The features extracted by each branch are fused through the cross-stitch unit and then passed to the decoder for processing; The multi-task decoder, by introducing other land feature recognition tasks besides landslide extraction, utilizes new feature information in the background to assist landslide extraction, enhances the feature mining capability of the model and utilizes the best sharing between cross-stitch unit learning tasks.

5. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 4 is characterized in that: The branch fusion encoder and the multi-task decoder both include a cross-stitch unit, and the original cross-stitch unit is used for adaptive sharing between different tasks; in the branch fusion encoder, it is used for multi-source data feature fusion; in order to meet the flexibility required for multi-source data feature fusion, the cross-stitch unit is expanded, the number of input and output features is not limited, and it is allowed to be any value and does not have to be the same, and the weight elements in the original weight matrix are expanded from scalars to vectors with the same dimension as the number of feature map channels. The processing of the feature map is calculated according to formula (1): in, represents the i-th output feature map, F j , j = 1, 2, ..., m represents the jth input feature map; w ij It means that when calculating the i-th output feature map, the j-th input feature map corresponds to a weight. The weight is a vector whose dimension is the same as the number of channels of the input feature map, that is, each channel in the feature map corresponds to an element in the corresponding weight vector. The calculation of the i-th output feature map is shown in formula (2): The operation is performed on a channel basis, that is, F j Each element in the k-th channel feature map is related to the vector w ij Multiply the corresponding k-th element in .

6. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 4 is characterized in that: The branch fusion encoder adopts a multi-branch structure to extract features from different source data respectively. The branch trunk consists of 1 input block and 4 downsampling blocks. Compared with the original UNet, the convolutional layer in the downsampling block is replaced by the residual block; the input and output of the input block are defined as and The height and width of the output feature map are the same as the input image, both are H in and W in , the number of channels is C data Convert to C in , C data is the number of channels of input data, C in is an adjustable parameter, set to 32 or 64, and the processing of the input block is expressed as formula (3): Y in =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X in )))))) (3) Among them, ReLU is the activation function, BN means batch normalization, Conv 3×3 represents a 3×3 convolution operation; the processing process of the downsampling block on the feature map is expressed as formula (4): Y d =Res(Res(MaxPool(X d ))) (4) in, and Respectively represent the input feature map and output feature map of the downsampling block, H d ×W d ×C d is the shape of the input feature map, Res is the residual block, MaxPool is the maximum pooling operation with a window size of 2×2 and a step size of 2. After being processed by the downsampling block, the height and width of the feature map are reduced by half, while the number of channels is doubled; the processing of the feature map by the residual block is expressed by formula (5): Y res =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (X res )))))+X res ) (5) Among them, X res and Y res Represent the feature maps of input and output respectively; The cross-stitch unit is applied at the jump connection and the bottom connection of the encoder and decoder to learn the association between multi-source data to fuse features and pass the fusion results to the decoder; at these locations, the number of input features and the number of output features of the cross-stitch unit correspond to the number of feature extraction branches in the encoder and the number of task branches in the decoder, respectively. The weight vector in the weight matrix is ​​initialized as shown in formula (6): Among them, N represents the number of branches in the encoder, and C is the dimension of the weight vector. Through this setting, the features extracted by each branch have the same weight at the beginning and are equally integrated.

7. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 4 is characterized in that: The multi-task decoder expands additional branches to identify suitable non-landslide features based on the decoder of the original UNet, and uses multi-task learning to make the model more fully utilize the information in the background. Each branch trunk consists of 4 upsampling blocks and an output block. Compared with the original UNet, the convolutional layer in the upsampling block is replaced by a residual block, and the output block of the auxiliary task branch is only enabled during model training; each upsampling block has two inputs, which are the feature maps from the last downsampling block in the encoder or the previous upsampling block. and the feature maps from the corresponding layers of the encoder via skip connections The output after upsampling block processing is Among them, H u ×W u ×C u It represents the shape of the feature map from the bottom downsampling block or the previous upsampling block. The processing of the upsampling block is as shown in formula (7): Y u =Res(Res(Cat(X u2 ,Conv T (X u1 )))) (7) Res represents the residual block, calculated according to formula (5), Cat represents the feature map according to the channel dimension, Conv T represents a transposed convolution; the output block consists of a 1×1 convolution whose input feature map has the shape of H out ×W out ×C in , the height and width of the output result remain unchanged, consistent with the height and width of the model input image, and the number of channels is 2, corresponding to the number of classification categories of each branch; A cross-stitch unit is added after each upsampling block of each branch for feature sharing. This unit will learn the best sharing between tasks. The number of its input and output features corresponds to the number of task branches. When initializing the parameters of this unit, the element values ​​of the weight vector on the diagonal of the weight matrix are all 1, and the element values ​​of the weight vector on the off-diagonal are set to all 0, so that each task branch is independent of each other at the beginning, as shown in formula (8): Among them, 1 C represents a vector of dimension C with all elements being 1, 0 C represents a vector of dimension C with all elements being 0.

8. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S4, each task corresponds to a joint loss function L, which is combined with the weighted cross entropy loss L CE and Dice loss L Dice , L is calculated according to formula (9): L=L CE +L Dice (9) Weighted cross entropy loss is an extension of cross entropy loss. Weights are introduced to deal with the problem of class imbalance. Its calculation formula is: Among them, B represents the total number of samples in a batch, C represents the total number of categories, and w j is the weight of category j, y ij is the one-hot encoded value of the true label of sample i in category j. If sample i is category j, then y ij is 1, otherwise it is 0. is the probability that the model predicts that sample i is of category j; the category weight is set according to the degree of category imbalance, using the common method based on inverse frequency, and the calculation formula is: Among them, N j is the total number of samples of category j; Dice loss is defined based on the Dice coefficient. Dice loss optimizes model performance by maximizing the Dice coefficient. Its calculation formula is: Among them, p ij and g ij Indicates whether sample i belongs to category j in the prediction result and the true label. If so, the value is 1, otherwise it is 0; The above loss calculation is for a single task. When training with multi-task learning, the learnable dynamically changing task loss weights are used to learn the relative importance of different tasks from the data. The total loss L of multi-task learning is MT The calculation formula is: Among them, T represents the total number of tasks, L i is the loss of task i, calculated by formula (9), s i is a learnable parameter based on which a learnable dynamic loss weight is implemented.

9. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S4, a data set is constructed based on the data obtained in S1 and processed according to S2, and the data is divided into a training set and a test set. The training set is enhanced by rotation, vertical flipping and horizontal flipping. The training set is used for BFM-UNet model training, and the test set is used to evaluate the performance of the trained model. When training the model, an appropriate batch size and training rounds are set, SGD is selected as the optimization algorithm, and a dynamically changing learning rate is used. The change of the learning rate is divided into two stages: In the first stage, at the beginning of training, the learning rate is gradually increased linearly from a very low value to the set value to reduce the impact of the model's sensitivity to the initial weights. It is calculated according to formula (14): Among them, α is an adjustable parameter, i is the current number of iterations in the first stage, and num phase1 is the total number of iterations in this stage, lr init is the initial learning rate; In the second stage, the learning rate gradually decays to close to 0 in a nonlinear manner, so that the model converges stably. The specific learning rate is calculated according to formula (15): Among them, β is an adjustable parameter, j is the current number of iterations in the second stage, and num phase2 is the total number of iterations in this phase.

10. The landslide extraction method based on multi-task learning and multi-source data feature fusion according to claim 1 is characterized in that: In S5, the multi-source data of the target area obtained by S1 and processed by S2 are input into the trained multi-task deep learning model BFM-UNet for landslide extraction, thereby obtaining a corresponding landslide distribution map, in which the landslide area is displayed in white and the non-landslide area is displayed in black.

Citation Information

Patent Citations

  • Landslide identification method and device based on satellite data UNet network model

    CN115131684A

  • Semantic segmentation method for landslide detection by using mid-resolution multi-source remote sensing data

    CN115588138A

  • Intelligent landslide identification method based on comprehensive multi-modal remote sensing image

    CN118941961A

  • Double-flow coding landslide automatic identification method based on improved U-Net

    CN119226891A

Cited By

  • SAR image landslide information extraction method based on deep learning algorithm

    CN120747600A

  • Landslide risk multi-criterion depth prediction method fusing geographic and physical information

    CN121170617A

  • Landslide extraction method based on multi-source remote sensing data fusion

    CN121661503A

  • Landslide extraction method based on multi-source remote sensing data fusion

    CN121661503B