Cross-domain image change detection method based on style randomization and similarity difference

By using a weighted twin convolutional neural network and a structural style randomization and multi-scale similarity difference module, the problem of insufficient adaptability to style differences and cross-domain generalization ability in remote sensing image change detection is solved, achieving higher detection accuracy and precision.

CN120976770AActive Publication Date: 2025-11-18SHANDONG FENGSHI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511483221.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-18
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing remote sensing image change detection methods are easily affected by lighting, climate and sensor differences when facing complex scenes, resulting in many false changes, insufficient cross-domain generalization ability, and lack of effective modeling of structural and multi-scale features, which affects the stability and accuracy of the detection results.

Method used

A cross-domain image change detection method based on style randomization and similarity difference is adopted. Features are extracted by a Siamese convolutional neural network with shared weights. Combined with structural style randomization and multi-scale structural similarity difference modules, the structural diversity and cross-domain robustness of features are enhanced, and subtle changes of boundaries and small targets are captured.

Benefits of technology

It significantly improves the accuracy and generalization ability of cross-domain remote sensing change detection, reduces the false change rate, improves the detection accuracy of complex scenes and small-scale targets, and generates more complete and accurate change detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976770A_ABST
    Figure CN120976770A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-domain image change detection method based on style randomization and similarity difference, and belongs to the technical field of remote sensing image processing and artificial intelligence. Inputting the two stages of images, extracting multi-layer features by using a twinborn convolutional neural network encoder sharing weight, and inputting the extracted multi-layer features step by step from a high layer to a low layer along a low-to-high transmission direction of a feature pyramid to realize fusion of shallow details and deep semantics; performing structural style randomization processing on the multi-layer fused features obtained in each period; inputting the two stages of feature maps after the structural styles are randomized under the same scale into a multi-scale structural similarity difference module to obtain fused difference features; performing convolutional decoding on the fused difference features to output a pixel-level change probability graph, and performing post-processing to obtain a final change detection result; and calculating a loss function by using the change probability graph output by the decoder, and carrying out optimization training to obtain a final model. According to the method, the accuracy and generalization ability of cross-domain remote sensing change detection can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a cross-domain image change detection method based on style randomization and similarity difference, belonging to the technical field of remote sensing image processing and artificial intelligence. BACKGROUND

[0002] Remote sensing image change detection is an important technology for identifying and analyzing changes in ground objects using images of the same area acquired at different times, and has wide applications in land use dynamic monitoring, ecological environment management, disaster assessment and urban planning. Traditional methods such as image difference, change vector analysis and principal component transformation, although simple to implement and computationally efficient, have obvious shortcomings when faced with complex scenes and high-resolution images, especially being susceptible to light, climate and sensor differences, resulting in many false changes and lack of stability in the detection results. With the development of deep learning, models based on convolutional neural networks and attention mechanisms have gradually become the mainstream method for change detection, achieving high accuracy on public datasets, but still have some problems to be solved in real-world applications.

[0003] Firstly, the problem of false changes caused by imaging condition differences is still prominent. Remote sensing images often come from different sensors or different time periods, and even after geometric correction, there are still significant differences in brightness, color, contrast and texture distribution. Such style differences are mistaken by existing models as actual ground object changes, resulting in an increase in false detection rate, especially in areas with complex textures such as urban construction and farmland cultivation, affecting the reliability of the detection results.

[0004] Secondly, the distribution of image data from different regions and different sensors is significantly different, and the trained model performs well on local data but has a significant drop in accuracy on new regions or new data sources, lacking cross-domain generalization ability, limiting the application value of the model in large-scale and long-term monitoring tasks.

[0005] Finally, existing methods still lack modeling of structure and multi-scale features. Most methods rely on pixel-level difference or simple feature difference to determine changes, lacking deep use of spatial structure similarity and context information. In practical applications, this deficiency manifests as insufficient sensitivity to small-scale change targets (such as temporary buildings and ponds) and boundary regions, leading to false positives or false negatives, thus reducing the completeness and precision of the detection results.

[0006] In summary, although deep learning has driven the rapid development of remote sensing image change detection, existing methods still have obvious shortcomings in style difference adaptability, label dependence, cross-domain generalization ability and structure modeling, and new technical means are urgently needed to overcome these bottlenecks. SUMMARY

[0007] The purpose of the present application is to overcome the above-mentioned deficiencies, and provide a cross-domain image change detection method based on style randomization and similarity difference, which can significantly improve the accuracy and generalization ability of cross-domain remote sensing change detection.

[0008] The technical scheme adopted by the present application is: The cross-domain image change detection method based on style randomization and similarity difference comprises the following steps: S1. Selecting multiple pairs of images of the same area in two periods, preprocessing and dividing the data set; S2. Inputting the two-period images into a twin convolutional neural network encoder with shared weights to extract multi-layer features, and inputting the extracted multi-layer features from high layer to low layer along the transmission direction from top to bottom of the feature pyramid to realize the fusion of shallow details and deep semantics, and obtaining two-period multi-layer fused features; S3. Performing structure style randomization processing on the multi-layer fused features obtained in each period; S4. Inputting the structure style randomized feature maps of the two periods at the same scale into a multi-scale structure similarity difference module, calculating the structure similarity index SSIM of the feature maps at different window sizes, obtaining the structure difference, and performing weighted fusion on the difference results of all windows to obtain the difference map at this scale level, and obtaining the fused difference features by upsampling and splicing the difference maps at all scale levels of the two periods; S5. Decoding the fused difference features to output a pixel-level change probability map, and post-processing to obtain the final change detection result; S6. Calculating the loss function using the change probability map output by the decoder, and optimizing the training to obtain the final model.

[0009] In the above method, the feature pyramid in step S2 is "top-down feature fusion": first, the high semantic features of the deepest layer are reduced in dimension through 1x1 convolution, then upsampled to expand the spatial size to be consistent with the high semantic features of the next deeper layer, and then added element by element with the high semantic features of the next deeper layer, and then smoothed through 3x3 convolution to obtain a layer of fused features; the fused features of the high layer are then reduced in dimension through 1x1 convolution, then upsampled to expand the spatial size to be consistent with the high semantic features of the next deeper layer, and so on, gradually passing the high layer features to the low layer to realize the fusion of shallow details and deep semantics.

[0010] The structure style randomization processing in step S3 is performed on each layer of fused features in each period. For the multi-layer fused features obtained in the same period, the deepest layer features are upsampled to the same spatial size as the shallow layer features, then KNN clustering is performed to generate K semantic segmentation maps, and K masks are obtained , , …, corresponding to different semantic regions, the features of each semantic region are extracted from the shallow features using the masks, the mean μ( ) and the standard deviation σ( ) of each region feature are calculated and used as the style feature of the region, the style feature of the current region is randomly replaced with the style features of other regions from the same feature, and new region features are generated, and all processed region features are combined to obtain a new feature map with local style disturbance. In the style randomization processing of the deepest feature structure, both inputs are the same feature, i.e., the deepest feature, one is used to extract the mask, and the other is used to exchange regions.

[0011] The structure style randomization calculation process is as follows: (1) the features of each semantic region are extracted from the shallow features using the masks: , where ⊙ represents element-wise multiplication, and represents the region feature of category c. (2) the mean μ( ) and the standard deviation σ( ) of each region feature are calculated and used as the style feature of the region, the style of the current region is randomly replaced with the style of other regions from the same feature, and the calculation process is as follows: , where μ and σ are the mean and standard deviation of the randomly selected other regions, respectively, to generate new region features ; (3) finally, all processed region features are combined to obtain a new feature map with local style disturbance: .

[0012] The process of obtaining the difference map in step S4 is as follows: (1) input two feature maps at the same scale, denoted as and , both with size B×C×h×w, first perform local statistical analysis on the feature maps under different window sizes, the window set is {1, 3, 5}, corresponding to 1×1, 3×3 and 5×5 sliding windows, respectively; (2) for any feature fragments x and y in a window, calculate their structural similarity index SSIM, defined as follows: , where and SSIM (x, y) = (2μxμy+C1) / (μx2+μy2+C1) and SSIM (x, y) = (2μxμy+C1) / (μx2+μy2+C1) SSIM (x, y) = (2μxμy+C1) / (μx2+μy2+C1) and C1 is a constant term to avoid zero denominator, usually C1 = 0.01², C1 = 0.03²; (3) After obtaining the structural similarity, further define the structural difference: , where w represents the window size, if the structure of the two features at this position is very similar, the SSIM value is close to 1, and the difference is close to 0; if the difference is significant, the SSIM value is lower, and the difference is close to 1; (4) Weighted fusion of the difference results of all windows: , where , , are the difference results of the window size of 1, 3, and 5 respectively, , , are learnable weights, satisfying .

[0013] The decoding in step S5 first uses a 3x3 convolution to expand the channel number, and cooperates with batch normalization and ReLU activation function to improve the feature expression ability: , Then through the up-sampling operation, the feature map is restored to the original image size HxW: , , obtain a feature map with the same size as the input image, Finally, use a 1x1 convolution layer to compress the channel number to 1, and map it to the [0, 1] interval through the Sigmoid function, output the pixel-level change probability map.

[0014] The post-processing in step S5 adopts Otsu adaptive threshold method to convert the change probability map into a binary change map, and then performs morphological closing operation and connected component analysis.

[0015] The loss function in step S6 adopts cross-entropy loss and Dice loss function, The cross-entropy loss function is defined as: , wherein N represents the number of all pixels, ∈ {0,1} is the true label of the i-th pixel, ∈ [0,1] is the corresponding predicted probability; The Dice loss function is defined as: , wherein ε is a very small constant, used to avoid the denominator being zero; The final total loss function is: ,

[0016] wherein λ is a balance coefficient, usually taking 1.

[0017] The beneficial effects of the present application are: The structure style randomization randomly replaces and disturbs the local style of different semantic regions at the shallow feature level, increases the structural diversity of the features, and improves the generalization ability of the model in complex scenes. The multi-scale structure similarity difference module calculates the structure similarity index on the multi-scale window, captures the subtle changes of the boundary and small target region, and reduces the false detection and missed detection. Compared with the existing method, the structure style randomization effectively weakens the style difference caused by cross-sensor and cross-time phase in remote sensing images, and reduces the pseudo change rate. The multi-scale structure similarity difference module calculates the structure similarity index on the pixel neighborhood and the multi-scale window, effectively captures the subtle differences of the boundary and small scale target; the two modules work together, so that the model has cross-domain robustness and fine change description ability, forming complementary advantages, thereby significantly improving the accuracy and generalization ability of cross-domain remote sensing change detection. The present application designs a change detection framework with overall creativity through the cooperation of the two modules. The feature pyramid structure is introduced to realize the step-by-step fusion of different level features, so as to balance the global scene perception and local detail description, and improve the overall feature expression ability. After obtaining the change probability map, the Otsu adaptive threshold method is used to convert it into a binary change map, and further through morphological closing operation and connected domain analysis, the isolated noise is removed and the small cavity is filled, so that a more complete and accurate change detection result is obtained. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is the method flowchart of the present application; Figure 2 is the method model network architecture diagram of the present application; Figure 3 is the structure style randomization processing process diagram; Figure 4 is the multi-scale structure similarity difference module processing process diagram. DETAILED DESCRIPTION

[0019] The application will be further described in connection with specific embodiments.

[0020] Example 1: A cross-domain image change detection method based on style randomization and similarity difference, including the following steps: S1. Selecting multiple pairs of two-phase images of the same area, pre-processing and dividing the data set; Select two-phase remote sensing images covering the same area, denoted as X1e R^{Bx3xH1xW1} and X2e R^{Bx3xH2xW2}, where B is the batch size, and the channel number is 3, corresponding to the red, green, and blue three bands. First, according to the geographic reference coordinates of the two-phase images, the completely overlapping area is extracted to ensure that the subsequent processing is within the same spatial range. In order to eliminate the spatial misalignment caused by imaging differences, geometric registration is performed on the two-phase images. Then the two-phase images are unified to the same spatial resolution. For images with low resolution, bilinear or bicubic interpolation method is used to resample them to the reference size HxW. At this time, the two-phase images obtained are X1' and X2', both with uniform spatial size. In order to further reduce the influence of brightness and contrast differences under different imaging conditions, the mean μ and standard deviation σ are calculated on each image in the three channels, and a standardization transformation is performed to make the images tend to be consistent in statistical distribution. After the above processing, the input pair (X1', X2') after geometric alignment and radiation normalization is obtained, which provides a stable input basis for subsequent feature extraction and difference calculation. All input pairs are divided into training set, validation set and test set according to the ratio of 7:1:2.

[0021] S2. Input the two-phase images into a twin convolutional neural network encoder with shared weights to extract multi-layer features, and the extracted multi-layer features are input in the transmission direction from top to bottom along the feature pyramid from high layer to low layer, realizing the fusion of shallow details and deep semantics, and obtaining two-phase multi-layer fused features: After preprocessing, the two-phase images X1' and X2' have been unified to the same size and distribution, which will be used as the input of the subsequent feature extraction network. In order to be able to extract the details and semantic information of ground objects at different scales, the application adopts a twin convolutional neural network with shared weights as the encoder. The so-called "twin" means that the two network structures are exactly the same and the parameters are shared, which respectively process X1' and X2', so that the extracted features have comparability.

[0022] The backbone network of the encoder selects ResNet-50, which contains multiple convolutional layers and residual blocks. After each convolution and pooling, the spatial resolution of the feature map is reduced, but the semantic information is gradually enhanced. For an image with an input size of Bx3xHxW, four levels of feature maps can be obtained after ResNet-50: : the number of channels is 256, the spatial size is about H / 4xW / 4, and mainly contains low-level features such as edges and textures; : the number of channels is 512, the spatial size is about H / 8xW / 8, and can capture larger structural information such as roads or small areas; : the number of channels is 1024, the spatial size is about H / 16xW / 16, and contains more semantic information such as building groups or farmland blocks; : the number of channels is 2048, the spatial size is about H / 32xW / 32, and mainly represents global semantic information.

[0023] Since a single scale of features cannot simultaneously consider details and semantics, the application adds a feature pyramid structure after the encoder to fuse information at different levels. The principle of the feature pyramid FPN is "top-down feature fusion": the high semantic features of the deepest layer (L4) are reduced to 256 channels through 1x1 convolution, then upsampled by one, so that their spatial size is expanded to be consistent with L3 , and are added element by element with L3 F 4, and then smoothed through 3x3 convolution to obtain a layer of fused features. The fused features of the higher layer are reduced through 1x1 convolution, then upsampled, so that their spatial size is expanded to be consistent with the high semantic features of the next deeper layer, and so on. The features of the higher layer are gradually transferred to the lower layer to realize the fusion of shallow details and deep semantics. Finally, four sets of fused features are obtained: ∈R^{Bx256xH / 4xW / 4}, ∈R^{Bx256xH / 8xW / 8}, ∈R^{Bx256xH / 16xW / 16}, ∈R^{Bx256xH / 32xW / 32}.

[0024] Under the twin structure, X1' and X2' will undergo the above encoding and pyramid process respectively, and the output feature sets are respectively denoted as and , where l ∈ {2, 3, 4, 5}. In this way, the two-phase images are represented by features with a unified number of channels at multiple scales, providing a foundation for subsequent style randomization and difference calculation.

[0025] S3. Perform structural style randomization processing on the obtained multi-layer fused features of each phase: The structural style randomization module is introduced to enhance the structural diversity of features and improve the robustness of the model in different scenarios. The module mainly consists of an unsupervised semantic perception unit and a local style exchange unit.

[0026] The structural style randomization processing is performed on each layer of the fused features of each phase. Finally, four features are obtained for each phase. Taking phase 1 as an example, the deep layer feature is , and the shallow layer feature is , , After obtaining the mask from , perform three times of style randomization processing on , , , and perform one time of style randomization processing on itself.

[0027] Taking the fourth layer (pyramid output feature) output feature f 4 as an example, upsample it to the same spatial size as the shallow layer feature f 1, then perform KNN clustering to generate K semantic segmentation maps, denoted as M ∈ R^{H / 4 × W / 4}. Where each pixel position belongs to {1, 2, …, K}, indicating the semantic class it is divided into. Through the semantic map, K masks { , , …, } can be obtained, corresponding to different semantic regions.

[0028] Next, extract the features of each semantic region from the shallow layer feature : , where ⊙ represents element-wise multiplication, and represents the region feature of class c.

[0029] For each region feature, calculate its mean μ( ) and standard deviation σ( ) as the style feature of the region. To increase structural diversity, the style of the current region is randomly replaced with the style from other regions in the same feature. The calculation process is as follows: , where μ and σ are the mean and standard deviation of the randomly selected other region, respectively. After this processing, the style of different semantic regions is randomized, thereby generating new region features .

[0030] Finally, all the processed region features are recombined to obtain a new feature map with local style disturbance: , The new feature map has higher style diversity on the semantic region, can significantly expand the structural distribution range of the training data, and alleviate the overfitting problem caused by single data. The processed feature will continue to be input into the subsequent layers of the encoder for deep feature extraction and finally participate in the output of decoding and change detection. Compared with the existing method, the application adopts the way of semantic region mask and local style random replacement instead of global style disturbance, which is more suitable for complex scenes in remote sensing images.

[0031] S4. The feature maps of the two periods after structure style randomization at the same scale are input into a multi-scale structure similarity difference module, the structure similarity index SSIM of the feature maps is calculated at different window sizes, the structure difference is obtained, the difference results of all windows are weighted and fused to obtain the difference map at this scale level, and after upsampling and splicing of the difference maps of all scale levels of the two periods, the fused difference feature is obtained: After processing by the twin encoder and the structure style randomization module, the feature maps of the two periods of images already have strong semantic consistency and diversified structure performance. In order to accurately extract the change information from them, the application introduces a multi-scale structure similarity difference module to obtain change representation by calculating the structure similarity of the features of the two periods at multiple spatial scales.

[0032] The input of the module is the feature maps of the two periods at the same scale, which are respectively denoted as and , and the size is BxCxhxw. The module first performs local statistical analysis on the feature maps at different window sizes, and the window set is {1, 3, 5}, which respectively corresponds to 1x1, 3x3 and 5x5 sliding windows.

[0033] For any feature fragments x and y in a window, the structure similarity index SSIM of them is calculated, which is defined as follows: , where, and respectively represent the mean, ² and denote the variance, denote the covariance, and is a constant term to avoid zero denominator, usually = 0.01², = 0.03².

[0034] After obtaining the structural similarity, the structural difference is further defined as: , where w denotes the window size. In this way, if the structure of the two periods of features at this position is very similar, the SSIM value is close to 1, and the difference is close to 0; if the difference is significant, the SSIM value is lower, and the difference is close to 1.

[0035] In order to take into account the structural differences at different scales, the module weights and fuses the difference results of all windows: , where , , are the difference results of the window sizes of 1, 3, and 5, respectively, , , are learnable weights that satisfy . The weights are automatically learned through backpropagation during training, so as to flexibly adjust the importance of each scale under different data scenarios.

[0036] Finally, the module will output a difference map at each scale level, which has the same size as the input feature. After upsampling and splicing of all levels of difference maps, they will be further fused in the decoder to generate a complete change probability map.

[0037] After processing by the structural style randomization module and the multi-scale structural similarity difference module, the difference maps of the two periods of images at different scales have been obtained, denoted as , whose spatial dimensions are H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32, respectively. In order to generate the final change detection result, these difference maps need to be uniformly fused and decoded.

[0038] First, the lower resolution difference maps are upsampled to the size of H / 4×W / 4 level by level. The upsampling method can be implemented by bilinear interpolation or deconvolution, denoted as Up(·). After processing, all difference maps have the same size, which are: , , , Where scale represents the upsampling factor. Then... These upsampling results are concatenated along the channel dimension to obtain the fused differential features: , at this time The number of channels is 4, and the spatial dimensions are H / 4×W / 4.

[0039] By introducing multi-scale structural similarity difference, this invention can simultaneously focus on fine-grained texture and large-scale contours, maintaining high detection accuracy for small targets, boundaries, and complex backgrounds. Based on the Structural Similarity Index (SSIM), this invention calculates and fuses differences across windows of H / 4, H / 8, and H / 16 scales, enabling better characterization of small targets and boundary variations.

[0040] S5. Decode the fused differential features to output a pixel-level change probability map, and then perform post-processing to obtain the final change detection result: First, a 3×3 convolution is used to expand the number of channels to 64, and batch normalization and ReLU activation function are used to improve feature representation capability: , Then, through two upsampling operations, the feature map is restored to the original image size H×W: , , A feature map of the same size as the input image is obtained. Finally, a 1×1 convolutional layer is used to compress the number of channels to 1, and the map is mapped to the [0,1] interval using the Sigmoid function, outputting a pixel-level probability map of changes. , in, ∈ R^{B×1×H×W} represents the probability value of each pixel being "changed".

[0041] After obtaining the probability map of change, post-processing is required to improve the usability of the results. First, the Otsu adaptive thresholding method is used to... Convert to binary transformation graph : , in, The threshold is calculated automatically. Next, morphological closing operations are used to fill the small holes and smooth the boundaries of the changing regions. Specifically, this involves... The morphological operation is performed by using a 5*5 structural element, and then dilating and eroding are performed. Finally, noise regions are removed by connected component analysis, and isolated spots with an area less than 50 pixels are deleted, to obtain the final change detection result .

[0042] Through the decoding and outputting process, the application can effectively fuse multi-scale differential information, generate a change detection map with complete boundaries and less noise, and provide reliable support for subsequent practical applications.

[0043] S6. Calculate the loss function by using the change probability map output by the decoder, and optimize the training to obtain a final model: After the decoder outputs the change probability map , a reasonable loss function needs to be designed to optimize the model. The application adopts a combination of cross-entropy loss and Dice loss to simultaneously consider pixel-level classification accuracy and overall region overlap.

[0044] The cross-entropy loss is defined as: , where N represents the number of all pixels, ∈ {0,1} is the true label of the i-th pixel, ∈ [0,1] is the corresponding predicted probability. The cross-entropy loss can effectively constrain the prediction accuracy at the pixel level.

[0045] The Dice loss is defined as: , where ε is a small constant to avoid a zero denominator. The Dice loss directly measures the overlap between the predicted result and the true label region, and has good robustness to the imbalance between foreground and background classes.

[0046] The final total loss function is: , where λ is a balance coefficient, usually taking 1.

[0047] In the training process, the AdamW optimizer is used to update the network parameters, and the initial learning rate is set to , and the cosine annealing strategy is used to dynamically adjust the learning rate. The batch size is set to 8, and the total number of training rounds is 200. The convolutional layer parameters of the model can be initialized using the ImageNet pre-trained weights to speed up the convergence speed.

[0048] In the training stage, firstly, the two images are input into the twin encoder and the structural style randomization module to extract and enhance the multi-scale features; then the multi-scale structural similarity difference module is used to generate the difference feature map; then the decoder is used to obtain the change probability map ; finally, the total loss is calculated according to the real label , and the network parameters are updated through back propagation. The process is repeated in each training round until the model converges.

[0049] In the inference stage, only the two images are input into the trained model, and the change probability map and the final binary change map can be directly output without additional post-processing steps, ensuring the efficiency and practicality of the method.

[0050] The above is a further description of the application in combination with the embodiments, and the protection scope of the application is not limited thereto.

Claims

1. A cross-domain image change detection method based on style randomization and similarity difference, characterized by: The steps include the following: S1. Select multiple pairs of two-period images of the same region, preprocess them, and divide the dataset; S2. The two images are input using a twin convolutional neural network encoder with shared weights to extract multi-layer features. The extracted multi-layer features are input step by step from high to low along the feature pyramid from low to high, so as to achieve the fusion of shallow details and deep semantics and obtain the multi-layer fused features of the two images. S3. Perform structural style randomization on the multi-layer fusion features obtained in each period; S4. Input the randomized feature maps of the structural styles of the two periods at the same scale into the multi-scale structural similarity difference module. Calculate their structural similarity index (SSIM) for the feature maps under different window sizes to obtain structural differences. Perform weighted fusion on the difference results of all windows to obtain the difference map at that scale level. After upsampling and stitching the difference maps of all scale levels of the two periods, obtain the fused difference features. S5. Decode the fused differential features to output a pixel-level change probability map, and then perform post-processing to obtain the final change detection result; S6. Calculate the loss function using the probability map output by the decoder, and optimize the training to obtain the final model.

2. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, Step S2, the feature pyramid, is a "top-down feature fusion": First, the deepest high semantic features are reduced in dimensionality through 1×1 convolution, then upsampled to expand their spatial size to be consistent with the second deepest high semantic features, and then element-wise added to the second deepest high semantic features, and then smoothed through 3×3 convolution to obtain a fused feature layer. After the high-level features are fused, they are then subjected to 1×1 convolution for dimensionality reduction, and then upsampled to expand their spatial size to match the high semantic features of the next deeper layer. This process is repeated step by step to pass the high-level features to the low-level layers, thereby achieving the fusion of shallow details and deep semantics.

3. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, Step S3, structural style randomization, processes the fused features of each layer in each period. For the multi-layer fused features obtained in the same period, the deepest layer features are upsampled to the same spatial size as the shallowest layer features, and then KNN clustering is performed to generate K categories of semantic segmentation maps, resulting in K masks. , , …, }, each corresponding to a different semantic region. These masks are used to extract features of each semantic region from the shallow features. For each region feature, its mean μ is calculated. ) and standard deviation σ( As the style feature of the current region, the style feature of the current region is randomly replaced with the style feature of other regions from the same shallow feature layer to generate a new regional feature. All processed regional features are recombined to obtain a new feature map of the shallow layer with local style perturbation. When the deepest feature structure style is randomized, both inputs are the same feature, namely the deepest feature. One is used to extract the mask, and the other is used to exchange regions.

4. The cross-domain image change detection method based on style randomization and similarity difference according to claim 3, characterized in that, The structural style randomization calculation process is as follows: (1) Using a mask to extract shallow features Extract features from each semantic region: , Where ⊙ represents element-wise multiplication. Represents the regional characteristics of category c; (2) For each regional feature, calculate its mean μ( ) and standard deviation σ( The style of the current region is used as a style feature of that region. The style of the current region is then randomly replaced with the style of other regions from the same feature. The calculation process is as follows: , Where μ and σ are the mean and standard deviation of other randomly selected regions, respectively, thus generating new regional features. ; (3) Finally, all the processed regional features are recombined to obtain a new feature map with local style perturbations: 。 5. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, The process of obtaining the difference map in step S4 is as follows: (1) Input the feature maps of the two periods at the same scale, and denote them as follows: and All windows have dimensions of B×C×h×w. First, local statistical analysis is performed on the feature maps under different window sizes. The window set is {1, 3, 5}, which correspond to sliding windows of 1×1, 3×3 and 5×5 respectively. (2) For any feature segments x and y within a window, calculate their structural similarity index SSIM, defined as follows: , in, and These represent the mean within the window, ² and ² represent the variances, Describing covariance, and This is a constant term used to avoid the denominator being zero; it is usually taken as... = 0.01², = 0.03²; (3) After obtaining the structural similarity, the structural difference is further defined: , Where w represents the window size. If the structures of the two features at this position are very similar, the SSIM value is close to 1 and the difference is close to 0; if the difference is significant, the SSIM value is low and the difference is close to 1. (4) Perform weighted fusion of the difference results of all windows: , in , , These are the difference results for window sizes of 1, 3, and 5, respectively. , , For learnable weights, satisfying .

6. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, The decoding described in step S5 first uses a 3×3 convolution to expand the number of channels, and then combines batch normalization and the ReLU activation function to improve feature representation capability. , Then, through upsampling, the feature map is restored to the original image size H×W: , , Obtain a feature map of the same size as the input image. Finally, a 1×1 convolutional layer is used to compress the number of channels to 1, and the Sigmoid function is used to map it to the [0,1] interval, outputting a pixel-level change probability map.

7. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, The post-processing described in step S5 uses the Otsu adaptive thresholding method to convert the change probability map into a binary change map, and then performs morphological closing operations and connected component analysis.

8. The cross-domain image change detection method based on style randomization and similarity difference according to claim 1, characterized in that, The loss function described in step S6 uses both cross-entropy loss and Dice loss function. The cross-entropy loss function is defined as follows: , Where N represents the total number of pixels. ∈ {0,1} represents the true label of the i-th pixel. ∈ [0,1] represents the corresponding predicted probability; The Dice loss function is defined as follows: , Where ε is a local constant used to avoid the denominator being zero; The final total loss function is: , Where λ is the balance coefficient, which is usually taken as 1.

Citation Information

Patent Citations

  • Remote sensing image change detection method combining feature fusion and enhancement difference

    CN115908254A

  • Cross-domain underwater target detection method of concerning invariant information

    CN118262227A

  • Multi-modal remote sensing image change detection method and system and readable storage medium

    CN120259299A

  • Double-flow remote sensing image change detection method fused with Mmba enhancement

    CN120298906A

  • Remote sensing image change detection method and system based on style alignment and edge constraint

    CN120339709A

Cited By

  • Dam break accident rapid change area detection method based on optical remote sensing image

    CN122265872A

  • A method for detecting rapidly changing areas in dam break accidents based on optical remote sensing images

    CN122265872B