Road change detection method and system

By adopting a feature extraction method based on bar convolution in road change detection, the problem of difficulty in capturing the linear features of the road in the prior art is solved, and a higher road change detection accuracy is achieved.

CN120164098AActive Publication Date: 2025-06-17STATE GRID LOCATION BASED SERVICE CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510191863.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the linear characteristics of the road, resulting in low accuracy of road change detection.

Method used

The road change feature extraction method based on strip convolution is adopted, and the two-time phase image is extracted through a pre-trained feature extraction model. The enhancement encoder is used to extract features in multiple directions to reduce the blurring of road edges in unrelated areas and improve the integrity and continuity of road change features.

Benefits of technology

It effectively avoids the fuzzy problem of traditional square convolution on the edges of slender objects, improves the integrity and continuity of road extraction, and thus improves the accuracy of road change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164098A_ABST
    Figure CN120164098A_ABST
Patent Text Reader

Abstract

The invention provides a road change detection method and system, and the method comprises the steps: carrying out the feature extraction of dual-time-phase images of a target region at different moments through a prediction training feature extraction model, and obtaining key road features corresponding to the dual-time-phase images; then based on key road features corresponding to the double-time-phase images, change features are calculated through a pixel-by-pixel difference algorithm, feature fusion is carried out on the change features, road change features of the target area at different moments are generated, and a pre-trained feature extraction model comprises a convolution layer, an enhancement encoder and a decoder which are connected in sequence. The enhanced encoder is used for carrying out feature extraction on a double-time-phase image in multiple different directions so as to avoid the problem of blurring of the edge of a slender object in the prior art, and then the integrity and continuity of road extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing change detection, and in particular to a road change detection method and system. Background Art

[0002] The purpose of remote sensing change detection (RSCD) is to locate and segment surface changes from co-registered image pairs taken at different times in the same area. It is an important issue in remote sensing understanding tasks and a key step in many real-world tasks, such as global resource monitoring, land use change detection, damage assessment, and urban management. Currently, most deep learning-based RSCD backbones regard RSCD as a binary segmentation task and use encoder-decoder structures such as Fully Convolutional Network (FCN) and U-Net to classify the extracted dual-temporal features into different categories to detect changes. For example, Li et al. proposed a dense attention refinement network based on a U-shaped encoder-decoder architecture, which supports progressive adjustment of the predicted change map at each stage of the decoding process.

[0003] Road change detection based on dual-temporal remote sensing images identifies newly built, disappeared, damaged, or rebuilt roads according to images collected at different times. Existing deep learning-based road change detection methods are mainly based on general change detection neural networks. Since roads in remote sensing images are usually distributed throughout the image and have the characteristics of large spans and narrow shapes, while general change detection methods use square convolution kernels, they cannot well capture the linear features of roads.

[0004] Therefore, how to accurately extract the linear features of roads is of great significance.

[0005] In order to adapt to the long and narrow characteristics of roads, this study focuses on a road change feature extraction method based on strip convolution, which can better adapt to the road distribution characteristics, reduce the blurring of road edges caused by irrelevant regions, and improve the integrity and continuity of road change features. Summary of the Invention

[0006] In order to solve the problem that the prior art cannot accurately capture the linear features of roads, the present invention proposes a road change detection method and system.

[0007] In a first aspect, a road change detection method is provided, including:

[0008] Feature extraction is performed on the dual-temporal images of the target region at different times through a pre-trained feature extraction model to obtain the key road features corresponding to the respective dual-temporal images. Among them, the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence. The enhanced encoder is used to extract features from the dual-temporal images in multiple different directions;

[0009] Based on the key road features corresponding to the respective dual-temporal images, change features are calculated through a pixel-by-pixel difference algorithm, and the change features are feature-fused to generate the road change features of the target region between different times.

[0010] Preferably, the performing feature extraction on the dual-temporal images of the target region at different times through a pre-trained feature extraction model to obtain the key road features corresponding to the respective dual-temporal images includes:

[0011] Feature extraction is respectively performed on the dual-temporal images of the target region at different times through the convolutional layer in the pre-trained feature extraction model to obtain the initial image features of the three-dimensional tensors corresponding to the respective dual-temporal images;

[0012] Feature extraction is respectively performed on different directions of the initial image features through the enhanced encoder in the pre-trained feature extraction model to obtain the key road features corresponding to the respective dual-temporal images.

[0013] Preferably, the enhanced encoder includes a plurality of hybrid strip convolutional feature enhancement modules and local convolutional kernels. The performing feature extraction on different directions of the initial image features corresponding to the respective dual-temporal images through the enhanced encoder in the pre-trained feature extraction model to obtain the key road features corresponding to the respective dual-temporal images includes:

[0014] Feature extraction is performed on different directions of the initial image features corresponding to the respective dual-temporal images through the plurality of hybrid strip convolutional feature enhancement modules of the enhanced encoder in the pre-trained feature extraction model to generate the initial features in a plurality of directions corresponding to the respective dual-temporal images;

[0015] The receptive fields of the initial features in the plurality of directions are enhanced by using the local convolutional kernels in the hybrid strip convolutional feature enhancement modules to obtain the feature outputs in a plurality of directions corresponding to the respective dual-temporal images;

[0016] The feature outputs in a plurality of directions corresponding to the respective dual-temporal images are mixed, and the mixed feature outputs are feature-fused and dimension-reduced;

[0017] Apply an activation function to perform non - linear transformation on the feature outputs corresponding to the fused and dimension - reduced dual - temporal images to generate the weight maps corresponding to the dual - temporal images respectively;

[0018] Based on the weight maps corresponding to the dual - temporal images respectively, perform weighted processing on the initial image features corresponding to the dual - temporal images respectively to obtain the key road features corresponding to the dual - temporal images respectively.

[0019] Preferably, the feature fusion of the change features to generate the road change features between different times of the target area includes:

[0020] Upsample the change features through the decoder of the feature extraction model, and perform change feature aggregation on the change features of different feature dimensions after upsampling through a feature aggregation strategy from high to low to generate the road change features with a fixed feature dimension between different times of the target area.

[0021] Preferably, the training process of the feature extraction model includes:

[0022] Use the training dual - temporal images of the target area at different times as inputs, use the real road change feature labels as the outputs corresponding to the training dual - temporal images, input the training dual - temporal images into the feature extraction model to be trained, and obtain the predicted road change features;

[0023] Calculate the total loss value between the predicted road change features and the real road change feature labels based on the hybrid loss function, and optimize the parameters of the feature extraction model to be trained based on the total loss value until the preset training end condition is met, and then obtain the trained feature extraction model;

[0024] Among them, the hybrid loss function is composed of a binary cross - entropy loss function and a hierarchical entropy minimization loss function, and the total loss value is the sum of the binary cross - entropy loss value and the hierarchical entropy minimization loss value between the predicted road change features and the real road change feature labels.

[0025] Preferably, the calculation formula of the binary cross - entropy loss is as follows:

[0026]

[0027] Among them, P is the predicted road change feature, Y is the real road change feature label corresponding to the predicted road change feature, L BCE (P,Y) represents the binary cross - entropy loss between the predicted road change feature and the real road change feature label, W is the width of the initial image feature, H is the height of the initial image feature, p ij represents the predicted road change feature at position i, j, y ijRepresents the true road change feature label at positions i and j.

[0028] Preferably, the calculation formula of the hierarchical entropy minimization loss is as follows:

[0029]

[0030] Where P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L HEM (P, Y) represents the hierarchical entropy minimization loss between the predicted road change feature and the true road change feature label, both α and β represent hyperparameters, and γ represents an exponent.

[0031] In a second aspect, the present invention also provides a road change detection system, including:

[0032] An extraction module, configured to extract features from the dual-temporal images of the target area at different times through a pre-trained feature extraction model, to obtain the key road features corresponding to each of the dual-temporal images, where the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence, and the enhanced encoder is configured to extract features from the dual-temporal images in multiple different directions;

[0033] A generation module, configured to calculate change features based on the key road features corresponding to each of the dual-temporal images through a per-pixel difference algorithm, and perform feature fusion on the change features to generate the road change features between different times in the target area.

[0034] Preferably, the extraction module is further configured to:

[0035] Extract features from the dual-temporal images of the target area at different times through the convolutional layer in the pre-trained feature extraction model, to obtain the initial image features of the three-dimensional tensors corresponding to each of the dual-temporal images;

[0036] Extract features from different directions of the initial image features through the enhanced encoder in the pre-trained feature extraction model, to obtain the key road features corresponding to each of the dual-temporal images.

[0037] Preferably, the enhanced encoder includes a plurality of hybrid bar convolutional feature enhancement modules and local convolutional kernels, and the extraction module is further configured to:

[0038] Extract features from different directions of the initial image features corresponding to each of the dual-temporal images through the plurality of hybrid bar convolutional feature enhancement modules in the enhanced encoder of the pre-trained feature extraction model to generate the initial features in multiple directions corresponding to each of the dual-temporal images;

[0039] Enhance the receptive fields of the initial features in multiple directions by using local convolutional kernels in the hybrid bar-shaped convolutional feature enhancement module to obtain the feature outputs in multiple directions corresponding to each of the dual-temporal images;

[0040] Mix the feature outputs in multiple directions corresponding to each of the dual-temporal images, and perform feature fusion and dimensionality reduction on the mixed feature outputs;

[0041] Apply an activation function to perform non-linear transformation on the feature outputs corresponding to each of the dual-temporal images after fusion and dimensionality reduction to generate the weight maps corresponding to each of the dual-temporal images;

[0042] Perform weighted processing on the initial image features corresponding to each of the dual-temporal images based on the weight maps corresponding to each of the dual-temporal images to obtain the key road features corresponding to each of the dual-temporal images.

[0043] Preferably, the generation module is further configured to:

[0044] Upsample the change features through the decoder of the feature extraction model, and perform change feature aggregation on the change features with different feature dimensions after upsampling through a feature aggregation strategy from high to low to generate the road change features with a fixed feature dimension between different moments of the target region.

[0045] Preferably, the training process of the feature extraction model in the extraction module includes:

[0046] Use the training dual-temporal images of the target region at different moments as inputs, use the real road change feature labels as the outputs corresponding to the training dual-temporal images, input the training dual-temporal images into the feature extraction model to be trained, and obtain the predicted road change features;

[0047] Calculate the total loss value between the predicted road change features and the real road change feature labels based on the hybrid loss function, and optimize the parameters of the feature extraction model to be trained based on the total loss value until the preset training end condition is satisfied, and then obtain the trained feature extraction model;

[0048] Wherein, the hybrid loss function is composed of a binary cross-entropy loss function and a hierarchical entropy minimization loss function, and the total loss value is the sum of the binary cross-entropy loss value and the hierarchical entropy minimization loss value between the predicted road change features and the real road change feature labels.

[0049] Preferably, the calculation formula of the binary cross-entropy loss in the model training sub-module is as follows:

[0050]

[0051] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L BCE (P, Y) represents the binary cross-entropy loss between the predicted road change feature and the true road change feature label, W is the width of the initial image feature, H is the height of the initial image feature, and p ij represents the predicted road change feature at position i, j, and y ij represents the true road change feature label at position i, j.

[0052] Preferably, the calculation formula of the hierarchical entropy minimization loss in the model training sub-module is as follows:

[0053]

[0054] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L HEM (P, Y) represents the hierarchical entropy minimization loss between the predicted road change feature and the true road change feature label, both α and β represent hyperparameters, and γ represents an exponent.

[0055] On the other hand, the present application also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;

[0056] The memory is used to store one or more programs;

[0057] When the one or more programs are executed by the at least one processor, the above-mentioned road change detection method is implemented.

[0058] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the above-mentioned road change detection method is implemented.

[0059] Compared with the prior art, the beneficial effects of the present invention are:

[0060] The present invention provides a method and system for road change detection. The method extracts features from dual-temporal images of a target area at different times through a feature extraction model trained by prediction, obtains the key road features corresponding to each of the dual-temporal images, then calculates change features through a pixel-by-pixel difference algorithm based on the key road features corresponding to each of the dual-temporal images, and performs feature fusion on the change features to generate road change features between different times in the target area. Moreover, the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence. The enhanced encoder is used to extract features from the dual-temporal images in multiple different directions to avoid the problem of blurring the edges of slender objects in the prior art, thereby improving the integrity and continuity of road extraction. Description of the Drawings

[0061] Figure 1 is a flowchart of the road change detection method of the present invention;

[0062] Figure 2 is a schematic diagram of the overall model architecture of the road change detection method of the present invention;

[0063] Figure 3 is a schematic diagram of the structure of the feature enhancement module of the road change detection method of the present invention;

[0064] Figure 4 is a schematic diagram of the visual comparison of the road change detection method of the present invention;

[0065] Figure 5 is a schematic diagram of the structure of the road change detection system of the present invention;

[0066] Figure 6 is a schematic diagram of the structure of an electronic device of the present invention. Detailed Embodiments

[0067] The present invention proposes a mixed strip convolution feature enhance module (MSCFE). This module extracts the linear features of roads through four strip convolution kernels in different directions and uses multi-scale strip convolution kernels to detect the long-distance correlation of road features. This method effectively avoids the problem of blurring the edges of slender objects by traditional square convolutions, thereby improving the integrity and continuity of road extraction, and further effectively improving the accuracy of road change detection.

[0068] To better understand the present invention, the content of the present invention will be further described below in conjunction with the accompanying drawings of the specification and embodiments.

[0069] Example 1:

[0070] A road change detection method, asFigure 1 As shown in

[0071] Step 1: Extract features from the dual-temporal images of the target area at different times through a pre-trained feature extraction model to obtain the key road features corresponding to each of the dual-temporal images. Among them, the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence. The enhanced encoder is used to extract features from the dual-temporal images in multiple different directions;

[0072] Step 2: Calculate the change features based on the key road features corresponding to each of the dual-temporal images through a pixel-by-pixel difference algorithm, and perform feature fusion on the change features to generate the road change features between different times in the target area.

[0073] In this embodiment, in the process of extracting the key road features corresponding to each of the dual-temporal images by using the pre-trained feature extraction model to extract features from the dual-temporal images of the target area at different times in Step 1, it includes:

[0074] Extract features from the dual-temporal images of the target area at different times respectively through the convolutional layer in the pre-trained feature extraction model to obtain the initial image features of the three-dimensional tensors corresponding to each of the dual-temporal images;

[0075] Extract features from different directions of the initial image features respectively through the enhanced encoder in the pre-trained feature extraction model to obtain the key road features corresponding to each of the dual-temporal images.

[0076] In this embodiment, the enhanced encoder includes a plurality of hybrid strip convolutional feature enhancement modules and local convolutional kernels. Based on this, in the process of extracting features from different directions of the initial image features corresponding to each of the dual-temporal images respectively through the enhanced encoder in the pre-trained feature extraction model and obtaining the key road features corresponding to each of the dual-temporal images, it includes:

[0077] Extract features from different directions of the initial image features corresponding to each of the dual-temporal images respectively through the plurality of hybrid strip convolutional feature enhancement modules in the enhanced encoder of the pre-trained feature extraction model to generate the initial features in multiple directions corresponding to each of the dual-temporal images;

[0078] Use the local convolutional kernels in the hybrid strip convolutional feature enhancement modules to enhance the receptive fields of the initial features in multiple directions to obtain the feature outputs in multiple directions corresponding to each of the dual-temporal images;

[0079] Mix the feature outputs in multiple directions corresponding to each of the dual-temporal images, and perform feature fusion and dimensionality reduction on the mixed feature outputs;

[0080] Apply an activation function to perform non-linear transformation on the feature outputs corresponding to the fused and dimension-reduced dual-temporal images to generate the weight maps corresponding to the dual-temporal images respectively;

[0081] Based on the weight maps corresponding to the dual-temporal images respectively, perform weighted processing on the initial image features corresponding to the dual-temporal images respectively to obtain the key road features corresponding to the dual-temporal images respectively.

[0082] In this embodiment, in step 2, during the process of performing feature fusion on the change features to generate the road change features between different moments of the target area, it includes:

[0083] Upsample the change features through the decoder of the feature extraction model, and perform change feature aggregation on the change features of different feature dimensions after upsampling through a feature aggregation strategy from high to low to generate the road change features with a fixed feature dimension between different moments of the target area.

[0084] In this embodiment, during the process of training the feature extraction model, it includes:

[0085] Use the training dual-temporal images of the target area at different moments as inputs, and use the real road change feature labels as the outputs corresponding to the training dual-temporal images to train the feature extraction model to be trained, and obtain the trained feature extraction model.

[0086] In this embodiment, during the process of training the feature extraction model to be trained to obtain the trained feature extraction model, it includes:

[0087] Input the training dual-temporal images into the feature extraction model to be trained to obtain the predicted road change features;

[0088] Calculate the total loss value between the predicted road change features and the real road change feature labels based on the hybrid loss function, and optimize the parameters of the feature extraction model to be trained based on the total loss value until the preset training end condition is met, and then obtain the trained feature extraction model;

[0089] Among them, the hybrid loss function is composed of a binary cross-entropy loss function and a hierarchical entropy minimization loss function, and the total loss value is the sum of the binary cross-entropy loss value and the hierarchical entropy minimization loss value between the predicted road change features and the real road change feature labels.

[0090] In one embodiment, the calculation formula of the binary cross-entropy loss is as follows:

[0091]

[0092] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L BCE (P, Y) represents the binary cross-entropy loss between the predicted road change feature and the true road change feature label. W is the width of the initial image feature, H is the height of the initial image feature, and p ij represents the predicted road change feature at position i, j, and y ij represents the true road change feature label at position i, j.

[0093] In one embodiment, the calculation formula of the hierarchical entropy minimization loss is as follows:

[0094]

[0095] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L HEM (P, Y) represents the hierarchical entropy minimization loss between the predicted road change feature and the true road change feature label. Both α and β represent hyperparameters, and γ represents an exponent.

[0096] Embodiment 2:

[0097] See Figure 2 , Figure 2 which is the schematic diagram of the overall network architecture of the road change detection method of the present invention. The road change detection method of the present invention will be introduced below in combination with Figure 2 herein.

[0098] The overall network architecture proposed by the present invention is as shown in Figure 2 and mainly consists of a dual-branch enhancement encoder and a decoder. First, the enhanced MobileNetV2 is used as the backbone network, and the average pooling layer and the fully connected layer of MobileNetV2 are removed. And in order to adapt to the characteristics of the road shape, a feature enhancement module MSCFE based on hybrid strip convolution is used to efficiently extract the key road features of the dual-temporal images (i.e., the T1 temporal image and the T2 temporal image). Next, the change features are calculated by pixel-by-pixel difference, and the change features at different levels are upsampled and spliced, adopting a feature aggregation strategy from high to low. At the same time, through the deep supervision mechanism (i.e., Supervision), the change features at each level can be gradually obtained to improve the overall detection performance. Among them, stage1, stage2... stage5 represent convolutional layer 1, convolutional layer 2... convolutional layer 5, d2, d3... d5 represent change feature 2, change feature 3... change feature 5, Decoder Block represents the decoder, output represents the output, and gt represents the true value (i.e., the aforementioned true road change feature label).

[0099] SeeFigure 3 , Figure 3 is a schematic diagram of the feature enhancement module structure of the road change detection method of the present invention. The feature enhancement in the present invention will be described below in conjunction with Figure 3 the present invention.

[0100] In road change detection, the shape of the road usually exhibits the characteristics of coexistence of large span and narrowness. However, when dealing with road features, traditional square convolution kernels have some obvious defects. First, the square convolution kernel will calculate non-road areas during feature extraction, thus affecting the capture effect of road linear features. Second, the square convolution kernel lacks sensitivity to the direction change of the road. Especially when extracting oblique features, the effect is significantly poor.

[0101] To solve these problems, the present invention proposes a feature enhancement module based on strip convolution. This module uses four strip convolution kernels: horizontal, vertical, left diagonal, and right diagonal, which can better adapt to road distributions in different directions. It not only optimizes the extraction ability of linear features but also helps to reduce the computational amount, thereby improving the overall detection efficiency.

[0102] As Figure 3 shown, the input of MSCFE is a three-dimensional tensor with a feature dimension of f in ∈R H×W×D , where H and W represent the height and width of the feature map respectively, D is the number of channels, R is the set of real numbers, strip pool represents strip pooling, Verticalkernel represents the vertical strip convolution kernel, Horizontal kernel represents the horizontal strip convolution kernel, Channel Avg represents channel average, Channel max represents channel maximum, Left diagonal kernel represents the left diagonal strip convolution kernel, Right diagonal kernel represents the right diagonal strip convolution kernel, and f out represents the output feature.

[0103] First, MSCFE extracts two key feature representations: f in from f h ∈R H×1×D and f w ∈R 1×W×DThis step effectively captures the global context information of the feature map in two main axis directions. Subsequently, to balance global and local information fusion, MSCFE adopts local convolutional kernels of specific sizes (w×1 and 1×w respectively), which further expands the receptive field of the features while maintaining local details. In addition, considering the possible complex directional features in natural scenes, MSCFE also introduces convolutional operations in the left diagonal and right diagonal directions. These operations can capture special directional information that traditional methods may ignore, thereby further enhancing the integrity and accuracy of feature representation.

[0104] Finally, MSCFE mixes the feature outputs from four different directions (horizontal, vertical, left diagonal, right diagonal) and performs feature fusion and dimensionality reduction through a 1×1 convolutional layer. Subsequently, a sigmoid activation function is applied to generate a weight map that reflects the relative importance of each position in the input feature map. Finally, this weight map is used to weight the original input feature f in to obtain the enhanced feature map f out ∈R H×W×D , highlighting the feature regions that are more critical for the road change detection task.

[0105] In one embodiment, RSCD can be regarded as a pixel-level classification task. Traditionally, such tasks rely on binary cross-entropy loss to optimize network parameters, and this loss function equally evaluates the contribution of each pixel. However, in the remote sensing image road change detection scenario, a significant problem is class imbalance - that is, the road change regions (positive samples) are significantly fewer in number compared to the unchanged regions (negative samples). This imbalance causes the binary cross-entropy loss to possibly over-focus on negative samples during training, thus ignoring the importance of positive samples and further affecting the model's ability to identify change regions.

[0106] Furthermore, due to the complex background characteristics of high-resolution remote sensing images, such as occlusion, shadow and other factors, it is difficult to accurately identify some roads, which is usually manifested as fewer false positives (FPs) and a significant increase in false negatives (FNs). To improve the recall rate of the model for change regions, that is, to reduce FNs, a mechanism needs to be designed to increase the penalty for FNs. Here, the HEM loss function flexibly adjusts the attention to FNs and FPs by introducing hyperparameters α and β to balance the error costs between the two.

[0107] In addition, the HEM loss also cleverly utilizes an exponent γ, which aims to enhance the model's attention to difficult-to-identify samples. This mechanism helps the model better cope with interference factors such as occlusion and shadow in the complex and changeable VHR remote sensing image environment and improve the overall detection performance.

[0108] In summary, to overcome the challenges of class imbalance and complex backgrounds in the remote sensing image road change detection task, we propose a hybrid loss function that combines binary cross-entropy loss \(L\) BCE and HEM loss \(L\) HEM . This combination can not only effectively balance the learning weights of positive and negative samples, but also strengthen the learning of difficult-to-identify samples through the special design of the HEM loss, thereby comprehensively improving the accuracy and robustness of road change detection.

[0109]

[0110] where · represents the dot product operation, |·| represents the L1 norm, \(P\) is the predicted change map, \(Y\) is the corresponding ground truth label, and \(W\) and \(H\) are the width and height of the image, respectively. The values of \(\alpha\), \(\beta\), and \(\gamma\) are 0.3, 0.7, and 0.75, respectively.

[0111] The total loss \(L\) total can be expressed as:

[0112] \(L\) total = \(L\) BCE + \(L\) HEM

[0113] Example 3:

[0114] The present invention is experimented on the road change detection dataset WRCD.

[0115] 3.1 Dataset

[0116] WRCD (Wuhan Road Change Detection, WRCD): The WRCD dataset is located in Jiangxia District, Wuhan City, China. It is 17-level remote sensing images downloaded from Google Earth, and the acquisition times are 2012, 2014, and 2016. This dataset only contains road changes. The resolution of each time-phase image is 10884×13655. The images are cropped into 512×512-sized patches with an overlap rate of 0.25. The first 48 columns of each time-phase are taken as test data, and the rest are used as training data, resulting in a total of 1995 pairs of training images and 980 pairs of test images.

[0117] 3.2 Implementation Details

[0118] This invention uses MobileNetV2 as the backbone, performs data augmentation on the input images using random cropping and random flipping, uses the Adam optimizer with a momentum of 0.9, β1 of 0.9, β2 of 0.99, an initial learning rate lr of 5e-4, and uses the poly strategy to decay the learning rate. As the training progresses, the learning rate is adjusted to (1-(curiteration / maxiteration)) power ×lr. The total number of training epochs is set to 100, the power value is 0.9, and the batch size is set to 10. All the code is implemented based on the PyTorch framework and experiments are conducted on a server equipped with GeForce RTX 4090, with a device video memory of 24GB, an operating system version of Ubuntu22.04, and a cuda version of 2.1.2. Here, curiteration / maxiteration represents the current iteration number / maximum iteration number.

[0119] 3.3 Evaluation Metrics

[0120] Four widely used evaluation metrics, Intersection over Union (IoU), F1-score (F1), Recall (Rec), and Precision (Pre), are adopted to evaluate the performance of RSCD. Among them, Pre reflects the proportion of correctly predicted positive samples in the predicted results and measures the precision of the model; Rec reflects the proportion of actual positive samples predicted as positive samples and measures the recall of the model; F1 is the weighted average of the two; IoU reflects the proportion of all correctly classified pixel points in the total pixels. The specific definitions of the above evaluation metrics are as follows:

[0121]

[0122] Where Tp, Fp, Tn, and Fn represent the numbers of true positives, false positives, true negatives, and false negatives respectively.

[0123] 3.4 Comparative Experiments

[0124] We compared the proposed model with five state-of-the-art RSCD methods, including three CNN-based methods: TinyCD, USSFCNet, and a Transformer-based method: BiT. To verify the robustness of the proposed method, experiments were conducted on the road change detection dataset WRCD respectively.

[0125] The visualization of the change detection results on the WRCD dataset is as Figure 4As shown, the quantization results are shown in Table 1. To make the visualization more intuitive, we use four different colors to represent different results: true positive (white), false positive (blue), true negative (black), and false negative (red). Among them, Label represents the annotated image, BiT represents a transfer learning model, TinyCD represents a lightweight model for change detection, USSFCNet represents the ultra-light hyperspectral feature collaboration network, and ours represents the detection result of the road change detection model of the present invention.

[0126] From the visualization Figure 4 it can be seen that the present invention performs better in road detection accuracy.

[0127] Advantages in detection accuracy: Compared with other methods, the road change detection accuracy of the present invention is higher, and it can locate the change target more accurately. In the first, fifth, and sixth rows of the figure, BiT, TinyCD, and USSFCNet can hardly identify road changes. In contrast, the present invention can identify the change area more accurately.

[0128] Advantages in road continuity: The roads in remote sensing images are long and narrow, resulting in incomplete identification of road changes. Compared with other methods, the present invention has more advantages in road continuity and integrity. In the fourth and sixth rows of the figure, the continuity of the road change areas identified by many methods is poor. On the contrary, the present invention can well identify the edges of road changes.

[0129] Table 1. Quantitative comparison on the WRCD dataset

[0130] Method IoU F1 Rec Pre BiT 0.4818 0.6503 0.6058 0.7019 TinyCD 0.5138 0.6788 0.6424 0.7196 USSFCNet 0.4188 0.5903 0.6271 0.5577 The method in this paper 0.5535 0.7126 0.7005 0.7252

[0131] From Table 1, we can draw the conclusion that the model proposed by the present invention achieves the best recall rate (0.7005), IoU (0.5535), and F1 score (0.7126) on the WRCD road dataset. Compared with TinyCD, the recall rate of the present invention has increased by 5.81%, indicating that our model can detect more complete road changes. Compared with TinyCD, our IoU has increased by 3.97%, indicating that the road changes detected by the present invention are more in line with the real situation.

[0132] 3.5 Ablation experiment

[0133] To verify the effectiveness of MSCFE, a series of ablation experiments are conducted on the WRCD dataset.

[0134] Table 2 shows the quantitative results. The model with MSCFE added obtains the best IoU, F1, and recall rate on the WRCD dataset, proving that the model can detect complete road changes.

[0135] Ablation results on the WRCD dataset in Table 2

[0136]

[0137] After deleting MSCFE, both the F1 score and recall rate of the model decreased significantly. The F1 score decreased from 71.26% to 70.88%, and the recall rate decreased from 73.95% to 70.87%, indicating that MSCFE is useful for road change detection.

[0138] Example 4:

[0139] The present invention based on the same inventive concept also provides a road change detection system, as Figure 5 shown, including:

[0140] An extraction module, configured to extract features from the dual-temporal images of the target area at different times through a pre-trained feature extraction model, to obtain the key road features corresponding to each of the dual-temporal images, wherein the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence, and the enhanced encoder is configured to extract features from the dual-temporal images in multiple different directions;

[0141] A generation module, configured to calculate change features based on the key road features corresponding to each of the dual-temporal images through a pixel-by-pixel difference algorithm, and perform feature fusion on the change features to generate the road change features of the target area between different times.

[0142] Preferably, the extraction module is further configured to:

[0143] Extract features from the dual-temporal images of the target area at different times through the convolutional layer in the pre-trained feature extraction model, to obtain the initial image features of the three-dimensional tensors corresponding to each of the dual-temporal images;

[0144] Extract features from different directions of the initial image features through the enhanced encoder in the pre-trained feature extraction model, to obtain the key road features corresponding to each of the dual-temporal images.

[0145] Preferably, the enhanced encoder includes a plurality of hybrid strip convolutional feature enhancement modules and local convolutional kernels, and the extraction module is further configured to:

[0146] Extract features from different directions of the initial image features corresponding to each of the dual-temporal images through the plurality of hybrid strip convolutional feature enhancement modules of the enhanced encoder in the pre-trained feature extraction model to generate the initial features in multiple directions corresponding to each of the dual-temporal images;

[0147] Enhance the receptive field of the initial features in multiple directions by using local convolution kernels in the hybrid strip convolution feature enhancement module to obtain the feature outputs in multiple directions corresponding to each of the two-temporal-phase images;

[0148] Mix the feature outputs in multiple directions corresponding to each of the two-temporal-phase images, and perform feature fusion and dimensionality reduction on the mixed feature outputs;

[0149] Apply an activation function to perform non-linear transformation on the feature outputs corresponding to each of the two-temporal-phase images after fusion and dimensionality reduction to generate the weight maps corresponding to each of the two-temporal-phase images;

[0150] Perform weighted processing on the initial image features corresponding to each of the two-temporal-phase images based on the weight maps corresponding to each of the two-temporal-phase images to obtain the key road features corresponding to each of the two-temporal-phase images.

[0151] Preferably, the generation module is further configured to:

[0152] Upsample the change features through the decoder of the feature extraction model, and perform change feature aggregation on the change features with different feature dimensions after upsampling through a feature aggregation strategy from high to low to generate the road change features with a fixed feature dimension between different moments of the target area.

[0153] Preferably, the training process of the feature extraction model in the extraction module includes:

[0154] Use the training two-temporal-phase images of the target area at different moments as inputs, and use the real road change feature labels as the outputs corresponding to the training two-temporal-phase images. Input the training two-temporal-phase images into the feature extraction model to be trained to obtain the predicted road change features;

[0155] Calculate the total loss value between the predicted road change features and the real road change feature labels based on the hybrid loss function, and optimize the parameters of the feature extraction model to be trained based on the total loss value until the preset training end condition is met, and then obtain the trained feature extraction model;

[0156] Wherein, the hybrid loss function is composed of a binary cross-entropy loss function and a hierarchical entropy minimization loss function, and the total loss value is the sum of the binary cross-entropy loss value and the hierarchical entropy minimization loss value between the predicted road change features and the real road change feature labels.

[0157] Preferably, the calculation formula of the binary cross-entropy loss in the model training sub-module is as follows:

[0158]

[0159] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L BCE (P, Y) represents the binary cross-entropy loss between the predicted road change feature and the true road change feature label. W is the width of the initial image feature, H is the height of the initial image feature, and p ij represents the predicted road change feature at position i, j, and y ij represents the true road change feature label at position i, j.

[0160] Preferably, the calculation formula of the hierarchical entropy minimization loss in the model training sub-module is as follows:

[0161]

[0162] Among them, P is the predicted road change feature, Y is the true road change feature label corresponding to the predicted road change feature, and L HEM (P, Y) represents the hierarchical entropy minimization loss between the predicted road change feature and the true road change feature label. Both α and β represent hyperparameters, and γ represents an exponent.

[0163] Example 5

[0164] As Figure 6 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.

[0165] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a road change detection method in the above embodiment.

[0166] Example 6

[0167] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in an electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more executable programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a road change detection method in the above embodiments can be realized.

[0168] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0169] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the function specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the function in the flow Figure 1One process or multiple processes and / or boxes Figure 1 The functions specified in one box or multiple boxes.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 One process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one box or multiple boxes.

[0172] The above are only embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A road change detection method, characterized in that: include: Extracting features from dual-phase images of a target area at different times through a pre-trained feature extraction model to obtain key road features corresponding to each of the dual-phase images, wherein the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder, and a decoder connected in sequence, and the enhanced encoder is used to extract features from the dual-phase images in multiple different directions; Based on the key road features corresponding to the dual-phase images, the change features are calculated by a pixel-by-pixel difference algorithm, and the change features are fused to generate the road change features of the target area at different times.

2. The method according to claim 1, characterized in that: The feature extraction of the dual-phase images of the target area at different times is performed by using a pre-trained feature extraction model to obtain key road features corresponding to each of the dual-phase images, including: The convolutional layers in the pre-trained feature extraction model are used to extract features from the dual-phase images of the target area at different times, so as to obtain initial image features of the three-dimensional tensors corresponding to the dual-phase images; The enhanced encoder in the pre-trained feature extraction model performs feature extraction on different directions of the initial image features respectively, so as to obtain key road features corresponding to each of the dual-phase images.

3. The method according to claim 2, characterized in that The enhanced encoder includes a plurality of mixed strip convolution feature enhancement modules and local convolution kernels. The enhanced encoder in the pre-trained feature extraction model extracts features from different directions of the initial image features corresponding to the dual-phase images to obtain key road features corresponding to the dual-phase images, including: Extracting features in different directions of initial image features corresponding to each of the dual-phase images through multiple mixed strip convolution feature enhancement modules of the enhancement encoder in the pre-trained feature extraction model generates initial features in multiple directions corresponding to each of the dual-phase images; The local convolution kernel in the mixed strip convolution feature enhancement module is used to enhance the receptive field of the initial features in the multiple directions, so as to obtain feature outputs in the multiple directions corresponding to the dual-phase images; Mixing the feature outputs in multiple directions corresponding to the dual-phase images, and performing feature fusion and dimensionality reduction on the mixed feature outputs; Applying an activation function to perform nonlinear transformation on the feature outputs corresponding to the dual-phase images after fusion and dimensionality reduction to generate weight maps corresponding to the dual-phase images; Based on the weight maps corresponding to the dual-phase images, weighted processing is performed on the initial image features corresponding to the dual-phase images to obtain the key road features corresponding to the dual-phase images.

4. The method according to claim 1, characterized in that: The step of fusing the change features to generate the road change features of the target area at different times includes: The change features are upsampled by the decoder of the feature extraction model, and the change features of different feature dimensions after upsampling are aggregated through a high-to-low feature aggregation strategy to generate road change features of the target area with fixed feature dimensions at different times.

5. The method according to claim 2, characterized in that: The training process of the feature extraction model includes: Taking the training dual-phase images of the target area at different times as input, taking the real road change feature labels as outputs corresponding to the training dual-phase images, inputting the training dual-phase images into the feature extraction model to be trained, and obtaining predicted road change features; Calculating the total loss value of the predicted road change feature and the actual road change feature label based on the mixed loss function, and optimizing the parameters of the feature extraction model to be trained based on the total loss value until a preset training end condition is met, thereby obtaining a trained feature extraction model; Among them, the hybrid loss function is composed of a binary cross entropy loss function and a hierarchical entropy minimization loss function, and the total loss value is the sum of the binary cross entropy loss value and the hierarchical entropy minimization loss value between the predicted road change feature and the real road change feature label.

6. The method according to claim 5, characterized in that The calculation formula of the binary cross entropy loss is as follows: Among them, P is the predicted road change feature, Y is the real road change feature label corresponding to the predicted road change feature, and L BCE (P,Y) represents the binary cross entropy loss between the predicted road change feature and the real road change feature label, W is the width of the initial image feature, H is the height of the initial image feature, and p ij represents the predicted road change characteristics at position i, j, y ij Represents the real road change feature label at position i, j.

7. The method according to claim 5, characterized in that The calculation formula of the hierarchical entropy minimization loss is as follows: Among them, P is the predicted road change feature, Y is the real road change feature label corresponding to the predicted road change feature, and L HEM (P,Y) represents the hierarchical entropy minimization loss between the predicted road change features and the true road change feature labels, α and β both represent hyperparameters, and γ represents an exponent.

8. A road change detection system, characterized in that: include: An extraction module, used to extract features from dual-phase images of a target area at different times through a pre-trained feature extraction model to obtain key road features corresponding to each of the dual-phase images, wherein the pre-trained feature extraction model includes a convolutional layer, an enhanced encoder and a decoder connected in sequence, and the enhanced encoder is used to extract features from the dual-phase images in multiple different directions; The generation module is used to calculate the change characteristics based on the key road features corresponding to each of the dual-phase images through a pixel-by-pixel difference algorithm, and perform feature fusion on the change characteristics to generate the road change characteristics of the target area at different times.

9. An electronic device, characterized in that: include: at least one processor and memory; The memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the road change detection method according to any one of claims 1 to 7 is implemented.

10. A readable storage medium, characterized in that: An execution program is stored thereon, and when the execution program is executed, the road change detection method as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Low-visibility road target detection method based on bimodal fusion

    CN115205651A

  • SAR image ship target detection method based on multistage feature fusion and mixed attention

    CN116469002A

  • Remote sensing image change detection method based on deep learning

    CN117934395A

  • Remote sensing image dense road segmentation method based on weaving type feature extraction

    CN119360349A

  • Method of predicting road attributer, data processing system and computer executable code

    US20230266144A1