High-resolution change detection method for multi-level difference feature grouping fusion
By introducing differential global Transformer and spatial change perception module in change detection, and adopting a new feature grouping and fusion paradigm, the shortcomings of existing methods in capturing global information and mining two-time phase change information are solved, and the accuracy and robustness of change detection are achieved.
Patent Information
- Application Number
- CN202510281190.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing high-resolution change detection method based on Transformer is difficult to effectively capture global information, and the change information between the two-phase images is not sufficiently mined, resulting in limited performance in complex scenarios.
A high-resolution change detection method for multi-level differential feature grouping and fusion is proposed. Multi-level features of two-time phase remote sensing images are extracted through differential global Transformer (DiffFormer), and a spatial change perception module (SCAM) is introduced to enhance the model's perception ability of complex shapes. Finally, a new feature grouping and fusion paradigm is used to efficiently fusion extract multi-level differential features.
By effectively extracting and integrating multi-level and multi-scale significant feature information, the model's ability to understand local details and global semantics is enhanced, the recognition accuracy of complex change areas is improved, and the accuracy and robustness of change detection are improved.
Smart Images

Figure CN120219958A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a high-resolution change detection method for multi-level differential feature grouping and fusion. Background Art
[0002] Remote sensing change detection identifies change areas by comparing remote sensing images of the same geographical area at different times, providing key information for land cover change observation. With the wide application of high-resolution satellites such as QuickBird, IKONOS, and WorldView, change detection technology has been further promoted and widely used in fields such as urban planning, environmental monitoring, disaster assessment, and resource management. Therefore, obtaining a high-precision change prediction map is crucial for accurately interpreting the observed scene.
[0003] Traditional change detection methods usually rely on pixel-level differential analysis to identify change areas by comparing remote sensing images at different time points. Common methods include principal component analysis, independent component analysis, conditional random field mixture models, and support vector machines, etc. However, these methods rely on manually designed features and are often affected by factors such as noise, image registration accuracy, and illumination changes in complex scenes, resulting in certain limitations in the accuracy and reliability of the detection results. In recent years, with the continuous development of deep learning, change detection methods based on convolutional neural networks have been widely applied. However, these methods are difficult to effectively capture global information, resulting in difficulty in maintaining structural integrity when dealing with large-scale change areas. To solve this problem, researchers have applied Transformer to the change detection task, further improving the detection accuracy by effectively aggregating context information. However, these methods only use Transformer as a single-temporal feature extractor and fail to fully exploit the change information between two-temporal images. In addition, the existing high-resolution change detection methods based on Transformer still need to improve the integration ability of different-level features and fail to fully utilize the correlation and complementarity between different-level features, resulting in limited performance in complex scenes. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a high-resolution change detection method for multi-level differential feature grouping and fusion, which extracts multi-level features of two-phase remote sensing images through a differential global Transformer (DiffFormer), integrates differential information while capturing global information, thereby generating richer feature representations and enhancing the model's perception ability. Subsequently, a spatial change awareness module (SCAM) is introduced to enhance the model's perception ability of complex shapes and avoid blurring or misjudgment when dealing with details or boundary regions. Finally, in order to efficiently fuse the extracted multi-level differential features, a new feature grouping and fusion paradigm is introduced, enabling the model to simultaneously focus on local details and global semantic information, thereby achieving high-precision change detection.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A high-resolution change detection method for multi-level differential feature grouping and fusion, the method comprising the following steps: S1: Preprocess the high-resolution images used as the training set to obtain training samples; S2: Input the training samples into a multi-level differential feature grouping and fusion network for training; S3: After training is completed, perform change detection on the test samples to obtain results.
[0007] Further, in step S1, preprocess the two-phase high-resolution images used as the training set. First, perform image registration to ensure the alignment of the two-phase images; then, perform atmospheric correction, radiometric correction, and geometric correction to eliminate atmospheric effects and radiometric differences under different sensors and shooting conditions, and ensure the geometric accuracy of the images; then, crop the images using non-overlapping windows of 256×256, and perform data augmentation operations, including random flipping, random rescaling, Gaussian blur, and random color jitter, where the random rescaling ratio range is (0.8 - 1.2); finally, perform normalization processing on the images to ensure that the images are comparable on the same scale.
[0008] Further, in step S2, input the training samples into a multi-stage DiffFormer to extract rich multi-scale features. DiffFormer consists of repeated consecutive blocks. In the initial block, the two-phase features are first normalized by layer normalization, and then a differential operation is performed to calculate the differential operator. The normalized features of the corresponding phase and the differential operator are input into a differential attention mechanism (DW-MSA) for processing. The formula is as follows:
[0009]
[0010] Among them, are the two-phase input features, LN is layer normalization, f l is the differential operator, represents the features output by DW-MSA.
[0011] Then, the processed two-phase features are again subjected to a difference operation, and together with the corresponding phase features, they are input into the DW-MSA for processing to obtain a residual, which is then added to the input features. The formula is as follows:
[0012]
[0013] Among them, is the processed two-phase feature, and m l is the difference operator, represents the feature after addition.
[0014] Finally, the feature after addition is subjected to a layer normalization operation once again, and then passed through a multi-layer perceptron. During this process, the Gaussian error linear unit is used as the non-linear activation function and added to the feature after addition. The formula is as follows:
[0015]
[0016] Among them, MLP is the multi-layer perceptron, represents the feature output by the initial block.
[0017] The feature output by the initial block is input into the subsequent block and a similar operation is performed. The difference is that the second attention operation uses the global difference attention mechanism (GDW-MSA), which introduces a global query component to make up for the defects of local attention without increasing the computational complexity. The formula is as follows:
[0018]
[0019] Among them, represents the feature output by the initial block, represents the feature output by the DW-MSA of the subsequent block, and f l+1 and m l+1 are the difference operators of the DW-MSA and GDW-MSA in the subsequent block respectively, q g1 is the global query component, represents the feature after addition by the GDW-MSA, represents the feature output by the subsequent block.
[0020] Furthermore, in the DW-MSA, first, the input feature is transformed into a query value through a linear layer, and the difference operator is transformed into a key value and an attribute value. Taking the first attention operation of the initial block as an example, the formula is as follows:
[0021]
[0022] Among them, represents the input feature, and f lDenote the difference operator, and Q, K, V denote the obtained query value, key value, and attribute value, respectively. W q , W k , W v denote the projection weights of the query value, key value, and attribute value, respectively.
[0023] Then, after dividing the input into h heads along the channel dimension, window attention calculation is performed using non-overlapping windows, and the calculated attention is concatenated along the channel dimension to obtain the final output result of DW-MSA. The formula is as follows:
[0024]
[0025] where, Q n , K n , V n denote the query value, key value, and attribute value of the n-th head, A n denotes the attention of the n-th head, softmax() denotes the activation function that maps values to the interval [0, 1], d denotes the dimension of each head, [] denotes the concatenation operation along the channel dimension, and A denotes the attention value obtained by DW-MSA.
[0026] Different from DW-MSA, GDW-MSA introduces a global query component. When calculating the attention of each head, the attention weight guided by the global query is added. The formula is as follows:
[0027]
[0028] where, q g1 denotes the global query component, denotes the attention of the n-th head of GDW-MSA.
[0029] Furthermore, in GDW-MSA, the global query component is pre-computed by the global query generator at each stage and shared among all blocks to interact with local key-value pairs. First, the height and width of the input feature map are transformed into the height and width of the local window through the repeated feature matching module, and then the global query component with the same resolution size as the local query is obtained through reshaping and repetition. The feature matching module mainly includes a convolutional layer and a squeeze-and-excitation activation module, where the GELU non-linear activation function is applied, and a residual connection is combined to enhance the feature transfer efficiency. A max-pooling layer is also used for downsampling to achieve step-by-step dimensionality reduction of the features. The formula is as follows:
[0030] FM(X) = MaxPool(f1(SE(GELU(f3(X)))) + X)
[0031] Among them, X is the input feature map, FM(X) is the feature map obtained after passing through the feature matching module, MaxPool() represents the max pooling operation, SE represents the squeeze-and-excitation module, and f1 and f3 respectively represent convolutional operations with convolutional kernel sizes of 1×1 and 3×3.
[0032] In the squeeze-and-excitation module, squeeze, excitation, and multiplicative feature fusion operations are adopted to adjust the importance of each channel, thereby enhancing the representation ability of the module and reducing the computational complexity. The formula is as follows:
[0033] SE(X) = res(σ(f l (GELU(f l (res(AvgPool(X)))))))·X
[0034] Among them, X is the input feature map, SE(X) is the feature map obtained after passing through the squeeze-and-excitation module, AvgPool() represents the average pooling operation, f l represents the linear layer, σ() represents the Sigmoid activation function, and res() represents the shaping operation.
[0035] Furthermore, the feature maps before and after the first-stage DiffFormer extraction are respectively input into SCAM. In this module, the feature map first undergoes preliminary feature transformation through two 3×3 convolutional layers, and then feature enhancement is performed through a multi-scale dilated deformable convolutional block. The enhanced feature map is concatenated with the input feature map in the channel dimension, and then passes through a 3×3 convolutional layer to generate the residual feature. Finally, the residual feature is added to the original input feature map to optimize the feature. The formula is as follows:
[0036] SCAM(X) = f3[MDDC(f3(ReLU(f3(X)))),X]+X
[0037] Among them, X is the input feature map, SCAM(X) is the feature map obtained after passing through the spatial change awareness module, and MDDC is the multi-scale dilated deformable convolutional block.
[0038] In the multi-scale dilated deformable convolutional block, the input feature map is respectively processed through a 1×1 convolutional layer and three dilated deformable convolutional blocks with convolutional kernel sizes of 3×3 and dilation rates of 3, 6, and 12. The feature maps processed by each branch are concatenated along the channel dimension to form the output of this module. The formula is as follows:
[0039]
[0040] Among them, X is the input feature map, MDDC(X) is the feature map obtained after passing through the multi-scale dilated deformable convolutional block, They respectively represent dilated deformable convolution blocks with a convolution kernel size of 3×3 and dilation rates of 3, 6, and 12.
[0041] Inside each dilated deformable convolution block, first, two dilated convolution layers with a convolution kernel size of 3×3 and a ReLU activation function are used to expand the receptive field; then, the transformed feature map is added to the input feature map to ensure that the original information details are not lost; finally, after being processed by a 3×3 convolution layer, a deformable convolution layer is introduced to dynamically adjust the position and shape of the convolution kernel, so as to more precisely adapt to the complex and irregular contour features of the changing region. Taking the dilation rate of 3 as an example, the formula is as follows:
[0042]
[0043] Among them, X is the input feature map, and f 3,3 represents the dilated convolution operation with a convolution kernel size of 3×3 and a dilation rate of 3, and f d represents the deformable convolution operation.
[0044] Furthermore, a differential operation is performed on the two-phase multi-scale features extracted by the multi-stage DiffFormer and processed by SCAM to obtain a multi-scale differential feature map. The formula is as follows:
[0045]
[0046] Among them, T1 and T2 represent two different time phases, respectively represent the features extracted by the DiffFormer in the i-th stage of different time phases, represents the feature map obtained by differentiating the features of the DiffFormer in the i-th stage between the two time phases, respectively represent the features processed by SCAM in different time phases, represents the feature map obtained by differentiating the features of the i-th stage processed by SCAM between the two time phases.
[0047] Furthermore, the low-level difference features are input into the low-level feature fusion module for adaptive fusion. In the low-level feature fusion module, first, is projected through a convolution module composed of a convolution layer, a batch normalization layer, and a ReLU activation function and added to . The formula is as follows:
[0048]
[0049] Among them, F is the added feature map, and f cDenote the convolutional module; then the added features are enhanced by the channel and spatial attention modules respectively, and the enhanced features are added together, and then a weight map is generated through the Sigmoid activation function. The formula is as follows:
[0050] w = σ(CA(F) + SA(F))
[0051] where F is the added feature map, w is the weight map, σ() represents the Sigmoid activation function, CA represents the channel attention module, and SA represents the spatial attention module; and the weight map is multiplied with the input feature respectively, and the obtained features are added to the input feature again, and finally, a 1×1 convolution is performed to generate the final fused feature. The formula is as follows:
[0052]
[0053] where w is the weight map, f1 represents the 1×1 convolution, and F l is the output feature map of the low-level feature fusion module.
[0054] In the channel attention module, the input feature map first undergoes max-pooling and average-pooling operations based on length and width respectively. Then, the pooled features are input into two-layer multi-layer perceptrons respectively, and added together, and then processed by the Sigmoid activation function to obtain the weighted coefficient. Finally, the weighted coefficient is multiplied element-wise with the input feature map to generate the final output of this module. The formula is as follows:
[0055] CA(X) = σ(MLP(AvgPool(X)) + MLP(MaxPool(X)))·X
[0056] where CA(X) is the feature map obtained after passing through the channel attention module, AvgPool(X) and MaxPool(X) respectively represent the feature maps obtained after average-pooling operation and max-pooling operation, and MLP represents the two-layer multi-layer perceptron.
[0057] In the spatial attention module, the input feature map first undergoes max-pooling and average-pooling operations based on channels respectively. Then, the pooled features are concatenated along the channel dimension. Next, a 7×7 convolution is performed for dimensionality reduction. Finally, after being processed by the Sigmoid activation function, it is multiplied with the input feature map to obtain the final output of this module. The formula is as follows:
[0058] SA(X) = σ(f7[AvgPool(X), MaxPool(X)])·X
[0059] Among them, SA(X) is the feature map obtained after passing through the spatial attention module, and f7 represents the convolution operation with a convolution kernel size of 7×7.
[0060] Furthermore, the high-level difference features are input into the high-level feature fusion module for fusion. In the high-level feature fusion module, the input features are first processed by the convolution module, and then the resolution is increased through the upsampling module composed of convolutional layers. The upsampled features are concatenated with another feature in the channel dimension, and then processed by the same convolution module and upsampling module and concatenated with another feature Finally, the obtained features are adjusted in resolution through the convolution module and processed by the normalization-based attention module to output the final fusion features. The formula is as follows:
[0061]
[0062] Among them, U represents the upsampling module, NAM represents the normalization-based attention module, and F h is the output feature map of the high-level feature fusion module.
[0063] Furthermore, the features after high- and low-level fusion are input into the feature merging module for effective interaction to form a more comprehensive feature representation. In the feature merging module, first, the low-level fusion features pass through the average pooling layer and two 1×1 convolutional layers, then perform an element-wise multiplication operation with themselves, and then pass through the Softmax activation function to obtain the enhanced shallow fusion features. The formula is as follows:
[0064] F l ′ = softmax(f 1×2 (AvgPool(F l )) · F l )
[0065] Among them, f 1×2 represents the two 1×1 convolutional operations, and F l ′ is the enhanced shallow fusion feature.
[0066] At the same time, the high-level fusion features are respectively input into the dilated convolution modules with a convolution kernel size of 3×3 and dilation rates of 1, 3, and 5. The features output by each module are concatenated along the channel dimension, and then passed through a 1×1 convolution module for feature integration to reduce the number of channels and enhance the feature expression ability, finally obtaining the enhanced deep fusion features. The formula is as follows:
[0067] F h ′ = f c1 [f c3,1 (Fh ), f c3,3 (F h ), f c3,5 (F h )]
[0068] Among them, f c3,1 , f c3,3 , f c3,5 respectively represent dilated convolution modules with a convolution kernel size of 3×3 and dilation rates of 1, 3, and 5. f c1 represents a 1×1 convolution module, and F h ′ is the enhanced deep fusion feature.
[0069] Finally, after processing the enhanced low-level fusion feature through a 1×1 convolution, it is matrix-multiplied with the enhanced high-level fusion feature. Then, the result of the matrix multiplication is normalized through the Softmax activation function and added to the high-level fusion feature. The formula is as follows:
[0070] F l ″ = softmax(f1(F l ′) × F h ′) + F h
[0071] Among them, F l ″ is the processed low-level fusion feature.
[0072] Similarly, the same operation process is performed on the enhanced high-level fusion feature. Finally, the features obtained from the above two steps are concatenated in the channel dimension and integrated through a 1×1 convolution module, and then input into the normalization-based attention module to generate the final merged feature, which is then passed through the upsampling module and the classification layer to obtain the final change detection result. The formula is as follows:
[0073] ″″
[0074] F c = C(U(NAM(f c1 [F l , F h )))
[0075] Among them, F h ″ is the processed high-level fusion feature, C represents the classification layer, and F c is the final change detection result.
[0076] Furthermore, the loss function of the overall network is set as follows:
[0077] Loss = -[X·logY + (1 - X)·log(1 - Y)]
[0078] Among them, X represents the predicted change detection result, and Y represents the ground truth label of the given sample. Calculate the loss function, and optimize the model parameters of the change detection framework according to the loss function and the backpropagation process. After training is completed, a trained change detection framework is obtained; the input sample is discriminated through the trained change detection framework, and a change detection result map is output.
[0079] The beneficial effects of the present invention are as follows:
[0080] The multi-level difference feature grouping fusion network (MDGF) proposed by the present invention effectively extracts and integrates multi-level and multi-scale significant feature information, which not only enhances the model's ability to understand local details and global semantics, but also improves the recognition accuracy of complex change regions, and enhances the accuracy and robustness of change detection. The present invention proposes a new Transformer structure (DiffFormer), which simultaneously introduces change information and global information in the attention mechanism, can effectively model the long-distance interdependent relationship between different objects, and significantly enhances the feature representation of change information. The present invention proposes a spatial change perception module (SCAM), which can capture the overall structure of the change region and can better adapt to the complex shapes and edges of the change region. The present invention proposes a new multi-level feature fusion method to adaptively combine spatial detail information and rich semantic information, thereby improving the model's performance in identifying interesting changes in complex scenarios. Experimental results on three publicly available high-resolution change detection datasets show that the performance of the proposed MDGF is superior to the current state-of-the-art high-resolution change detection methods.
[0081] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0082] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably with reference to the accompanying drawings, where:
[0083] Figure 1 is the flowchart of the method of the present invention;
[0084] Figure 2 is the diagram of the multi-level difference feature grouping fusion network (MDGF) for high-resolution change detection;
[0085] Figure 3Structural diagram of the differential global Transformer (DiffFormer) of the present invention, where (a) is the structural diagram of the continuous block in DiffFormer, (b) is the structural diagram of DW-MSA, and (c) is the structural diagram of GDW-MSA;
[0086] Figure 4 Structural diagram of the global query generator of the present invention;
[0087] Figure 5 Structural diagram of the spatial change awareness module (SCAM) of the present invention;
[0088] Figure 6 Structural diagram of the low-level feature fusion module of the present invention;
[0089] Figure 7 Structural diagram of the high-level feature fusion module of the present invention;
[0090] Figure 8 Structural diagram of the feature merging module of the present invention;
[0091] Figure 9 Several examples in the LEVIR-CD dataset, where the Image1 column is the pre-change image, the Image2 column is the post-change image, and the MDGF column is the change detection result of the method described in the present invention;
[0092] Figure 10 Visualization results of different methods on the LEVIR-CD dataset, where the methods include FC-EF, FC-Siam-Di, FC-Siam-Conc, DTCDSCN, SNUNet, BIT, ChangeFormer, RDPNet, MDGF, and the GT column is the ground truth map. Detailed implementation manner
[0093] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0094] Figure 1 Flowchart of the method of the present invention. The present invention provides a multi-level differential feature grouping and fusion method for high-resolution change detection. As shown in the figure, in the image preprocessing stage, image registration is quickly achieved by means of radiometric correction, geometric correction, etc., and data augmentation operations are performed. The deep learning network for change detection is as Figure 2As shown, it can effectively extract and integrate multi-scale significant feature information, improving the accuracy and robustness of change detection. The network consists of a differential global Transformer (DiffFormer), a spatial change awareness module (SCAM), and a new feature grouping and fusion paradigm. First, the multi-scale differential features are obtained from the dual-temporal high-resolution images through the DiffFormer and the spatial change awareness module designed for low-level features. Then, a new feature grouping and fusion paradigm is introduced to efficiently fuse local details and global semantic information, thereby achieving high-precision change detection. The differential global Transformer (DiffFormer) designed in the present invention is used to effectively model long-range interdependencies and significantly enhance the feature representation of change information. The spatial change awareness module SCAM designed in the present invention is used to capture the overall structure of the changed area and better adapt to the complex shape of the changed area. The feature grouping and fusion paradigm designed in the present invention is used to adaptively combine spatial detail information and rich semantic information, improving the performance of the model in identifying changes of interest in complex scenarios. Specifically, the technical solution of the present invention includes the following contents:
[0095] 1. Preprocessing of high-resolution images: Preprocess the dual-temporal high-resolution images used as the training set. First, perform image registration to ensure the alignment of the two temporal images. Then, perform atmospheric correction, radiometric correction, and geometric correction to eliminate the atmospheric effect and the radiometric differences under different sensors and shooting conditions, and ensure the geometric accuracy of the images. Then, crop the images using non-overlapping windows of 256×256 and perform data augmentation operations, including random flipping, random rescaling, Gaussian blur, and random color jitter, where the random rescaling ratio range is (0.8 - 1.2). Finally, normalize the images to ensure that the images are comparable on the same scale.
[0096] 2. Input the training samples into the multi-stage DiffFormer to extract rich multi-scale features. The DiffFormer consists of repeated consecutive blocks. As shown in Figure 3 (a), in the initial block, the two-temporal features are first normalized through layer normalization, and then the differential operation is performed to calculate the difference operator. The normalized features of the corresponding temporal phases and the difference operator are input into the differential attention mechanism (DW-MSA) for processing. The formula is as follows:
[0097]
[0098] where, are the two-temporal input features, LN is layer normalization, f l is the difference operator, represents the features output by DW-MSA.
[0099] Then, the processed two-phase features are subjected to a difference operation again, and the corresponding phase features are input into the DW-MSA for processing again to obtain a residual, which is then added to the input features. The formula is as follows:
[0100]
[0101] Among them, is the processed two-phase feature, and m l is the difference operator, represents the feature after addition.
[0102] Finally, the feature after addition is subjected to a layer normalization operation again, and then passed through a multi-layer perceptron. During this process, the Gaussian error linear unit is used as the non-linear activation function and added to the feature after addition. The formula is as follows:
[0103]
[0104] Among them, MLP is the multi-layer perceptron, represents the feature output by the initial block.
[0105] The feature output by the initial block is input into the subsequent block and a similar operation is performed. The difference is that the second attention operation uses the global difference attention mechanism (GDW-MSA), which introduces a global query component to make up for the defects of local attention without increasing the computational complexity. The formula is as follows:
[0106]
[0107] Among them, represents the feature output by the initial block, represents the feature output by the DW-MSA of the subsequent block, and f l+1 and m l+1 are the difference operators of DW-MSA and GDW-MSA in the subsequent block respectively, and q g1 is the global query component, represents the feature after addition by GDW-MSA, represents the feature output by the subsequent block.
[0108] 3. The schematic diagram of DW-MSA is as shown in Figure 3 (b). In DW-MSA, first, the input feature is transformed into a query value through a linear layer, and the difference operator is transformed into a key value and an attribute value. Taking the first attention operation of the initial block as an example, the formula is as follows:
[0109]
[0110] Among them, represents the input feature, and fl Denotes the difference operator, Q, K, and V denote the obtained query value, key value, and attribute value, and W q , W k , W v respectively denote the projection weights of the query value, key value, and attribute value.
[0111] Then, after dividing the input into h heads in the channel dimension, non-overlapping windows are used for window attention calculation, and the calculated attention is concatenated in the channel dimension to obtain the final output result of DW-MSA. The formula is as follows:
[0112]
[0113] Among them, Q n , K n , V n denote the query value, key value, and attribute value of the nth head, A n denotes the attention of the nth head, softmax() denotes the activation function that maps values to the interval [0, 1], d denotes the dimension of each head, [] denotes the concatenation operation in the channel dimension, and A denotes the attention value obtained by DW-MSA.
[0114] The schematic diagram of GDW-MSA is as shown in Figure 3 (c). Different from DW-MSA, GDW-MSA introduces a global query component. When calculating the attention of each head, the attention weight guided by the global query is added. The formula is as follows:
[0115]
[0116] Among them, q g1 denotes the global query component, denotes the attention of the nth head of GDW-MSA.
[0117] 4. In GDW-MSA, the global query component is pre-computed by the global query generator at each stage and shared among all blocks to interact with local key-value pairs. The schematic diagram of the global query generator is as shown in Figure 4 . First, the height and width of the input feature map are transformed into the height and width of the local window through the repeated feature matching module, and then the global query component with the same resolution size as the local query is obtained through shaping and repetition. The feature matching module mainly includes a convolutional layer and a compression activation module, where the GELU non-linear activation function is applied, and a residual connection is combined to enhance the feature transfer efficiency. A max-pooling layer is also used for downsampling to achieve step-by-step dimensionality reduction of the features. The formula is as follows:
[0118] FM(X) = MaxPool(f1(SE(GELU(f3(X)))) + X)
[0119] Where X is the input feature map, FM(X) is the feature map obtained after passing through the feature matching module, MaxPool() represents the max pooling operation, SE represents the squeeze-and-excitation module, and f1 and f3 represent convolutional operations with convolutional kernels of size 1×1 and 3×3 respectively.
[0120] In the squeeze-and-excitation module, squeeze, excitation, and multiplicative feature fusion operations are adopted to adjust the importance of each channel, thereby enhancing the representation ability of the module and reducing the computational complexity. The formula is as follows:
[0121] SE(X) = res(σ(f l (GELU(f l (res(AvgPool(X)))))))·X
[0122] Where X is the input feature map, SE(X) is the feature map obtained after passing through the squeeze-and-excitation module, AvgPool() represents the average pooling operation, f l represents a linear layer, σ() represents the Sigmoid activation function, and res() represents an integer operation.
[0123] 5. Input the feature maps before and after the first-stage DiffFormer extraction into SCAM respectively. The schematic diagram of SCAM is as Figure 5 shown. In this module, the feature map first undergoes preliminary feature transformation through two 3×3 convolutional layers, and then feature enhancement is performed through a multi-scale dilated deformable convolutional block. The enhanced feature map is concatenated with the input feature map in the channel dimension, and then passes through a 3×3 convolutional layer to generate a residual feature. Finally, this residual feature is added to the original input feature map to achieve feature optimization. The formula is as follows:
[0124] SCAM(X) = f3[MDDC(f3(ReLU(f3(X)))), X] + X
[0125] Where X is the input feature map, SCAM(X) is the feature map obtained after passing through the spatial change awareness module, and MDDC is the multi-scale dilated deformable convolutional block.
[0126] In the multi-scale dilated deformable convolutional block, the input feature map is processed through a 1×1 convolutional layer and three dilated deformable convolutional blocks with convolutional kernels of size 3×3 and dilation rates of 3, 6, and 12 respectively. The feature maps processed by each branch are concatenated along the channel dimension to form the output of this module. The formula is as follows:
[0127]
[0128] Among them, X is the input feature map, and MDDC(X) is the feature map obtained after passing through the multi-scale dilated deformable convolution block. They respectively represent dilated deformable convolution blocks with a convolution kernel size of 3×3 and dilation rates of 3, 6, and 12.
[0129] Inside each dilated deformable convolution block, first, two dilated convolution layers with a convolution kernel size of 3×3 and the ReLU activation function are used to expand the receptive field; then, the transformed feature map is added to the input feature map to ensure that the original information details are not lost; finally, after being processed by a 3×3 convolution layer, a layer of deformable convolution layer is introduced to dynamically adjust the position and shape of the convolution kernel, so as to more precisely adapt to the complex and irregular contour features of the changing region. Taking the dilation rate of 3 as an example, the formula is as follows:
[0130]
[0131] Among them, X is the input feature map, and f 3,3 represents the dilated convolution operation with a convolution kernel size of 3×3 and a dilation rate of 3, and f d represents the deformable convolution operation.
[0132] 6. Perform a difference operation on the two-phase multi-scale features extracted by the multi-stage DiffFormer and processed by SCAM to obtain a multi-scale difference feature map. The formula is as follows:
[0133]
[0134] Among them, T1 and T2 represent two different time phases, respectively represent the features extracted by the DiffFormer in the i-th stage of different time phases, represents the feature map obtained by the difference of the two-phase features after the DiffFormer in the i-th stage, respectively represent the features processed by SCAM in different time phases, represents the feature map obtained by the difference of the two-phase features after being processed by SCAM in the i-th stage.
[0135] 7. Input the low-level difference features into the low-level feature fusion module for adaptive fusion. As Figure 6 shown, in the low-level feature fusion module, first, is projected through a convolution module composed of a convolution layer, a batch normalization layer, and the ReLU activation function and added to The formula is as follows:
[0136]
[0137] Among them, F is the added feature map, and f c represents the convolutional module; then the added features are enhanced through the channel and spatial attention modules respectively, and the enhanced features are added together, and then a weight map is generated through the Sigmoid activation function. The formula is as follows:
[0138] w = σ(CA(F) + SA(F))
[0139] Among them, F is the added feature map, w is the weight map, σ() represents the Sigmoid activation function, CA represents the channel attention module, and SA represents the spatial attention module; and the weight map is multiplied with the input feature respectively, and the obtained features are added to the input feature again, and finally, a 1×1 convolution is performed to generate the final fused feature. The formula is as follows:
[0140]
[0141] Among them, w is the weight map, f1 represents the 1×1 convolution, and F l is the output feature map of the low-level feature fusion module.
[0142] In the channel attention module, the input feature map first undergoes max-pooling and average-pooling operations based on length and width respectively. Then, the pooled features are input into two-layer multi-layer perceptrons respectively, and added together, and then processed through the Sigmoid activation function to obtain the weighted coefficient. Finally, the weighted coefficient is multiplied with the input feature map element-wise to generate the final output of this module. The formula is as follows:
[0143] CA(X) = σ(MLP(AvgPool(X)) + MLP(MaxPool(X)))·X
[0144] Among them, CA(X) is the feature map obtained after passing through the channel attention module, AvgPool(X) and MaxPool(X) respectively represent the feature maps obtained after average-pooling operation and max-pooling operation, and MLP represents the two-layer multi-layer perceptron.
[0145] In the spatial attention module, the input feature map first undergoes max-pooling and average-pooling operations based on channels respectively. Then, the pooled features are concatenated along the channel dimension. Next, a 7×7 convolution is performed for dimensionality reduction. Finally, after being processed through the Sigmoid activation function, it is multiplied with the input feature map to obtain the final output of this module. The formula is as follows:
[0146] SA(X) = σ(f7[AvgPool(X), MaxPool(X)])·X
[0147] Among them, SA(X) is the feature map obtained after passing through the spatial attention module, and f7 represents the convolution operation with a convolution kernel size of 7×7.
[0148] 8. Input the high-level difference features into the high-level feature fusion module for fusion. As Figure 7 shown, in the high-level feature fusion module, the input feature is first processed by the convolution module, and then the resolution is enhanced through the upsampling module composed of convolutional layers. The upsampled feature is concatenated with another feature in the channel dimension, and then processed by the same convolution module and upsampling module and concatenated with another feature Finally, the obtained feature is adjusted in resolution through the convolution module and processed by the normalization-based attention module, and the final fused feature is output. The formula is as follows:
[0149]
[0150] Among them, U represents the upsampling module, NAM represents the normalization-based attention module, and F h is the output feature map of the high-level feature fusion module.
[0151] 9. Input the features after high- and low-level fusion into the feature merging module for effective interaction to form a more comprehensive feature representation. As Figure 8 shown, in the feature merging module, first, the low-level fused feature passes through the average pooling layer and two 1×1 convolutional layers, and then performs an element-wise multiplication operation with itself, and then passes through the Softmax activation function to obtain the enhanced shallow fused feature. The formula is as follows:
[0152] F l ′ = softmax(f 1×2 (AvgPool(F l )) · F l )
[0153] Among them, f 1×2 represents the two 1×1 convolutional operations, and F l ′ is the enhanced shallow fused feature.
[0154] At the same time, the high-level fused feature is respectively input into the dilated convolution modules with a convolution kernel size of 3×3 and dilation rates of 1, 3, and 5. The features output by each module are concatenated along the channel dimension, and then feature integration is performed through a 1×1 convolution module to reduce the number of channels and enhance the feature expression ability, and finally the enhanced deep fused feature is obtained. The formula is as follows:
[0155] Fh ′ = f c1 [f c3,1 (F h ), f c3,3 (F h ), f c3,5 (F h )]
[0156] Among them, f c3,1 , f c3,3 , f c3,5 respectively represent atrous convolution modules with a convolution kernel size of 3×3 and atrous rates of 1, 3, and 5, f c1 represents a 1×1 convolution module, and F h ′ is the enhanced deep fusion feature.
[0157] Finally, after processing the enhanced low-level fusion feature through a 1×1 convolution, it is matrix-multiplied with the enhanced high-level fusion feature. Then, the result of the matrix multiplication is normalized through the Softmax activation function and added to the high-level fusion feature. The formula is as follows:
[0158] F l ″ = softmax(f1(F l ′) × F h ′) + F h
[0159] Among them, F l ″ is the processed low-level fusion feature.
[0160] Similarly, the same operation process is performed on the enhanced high-level fusion feature. Finally, the features obtained from the above two steps are concatenated in the channel dimension and integrated through a 1×1 convolution module, and then input into the normalization-based attention module to generate the final merged feature, which is then passed through the upsampling module and the classification layer to obtain the final change detection result. The formula is as follows:
[0161] F c = C(U(NAM(f c1 [F l ″, F h ″])))
[0162] Among them, F h ″ is the processed high-level fusion feature, C represents the classification layer, and F c is the final change detection result.
[0163] 10. The loss function of the overall network is set as: Loss = -[X·logY + (1 - X)·log(1 - Y)], where X represents the predicted change detection result, and Y represents the ground truth label of the given sample. Calculate the loss function, and optimize the model parameters of the change detection framework according to the loss function and the backpropagation process. After training is completed, a trained change detection framework is obtained; use the trained change detection framework to discriminate the input sample and output the change detection result map.
[0164] As Figure 9 Figure 5 is the experimental result of the MDGF change detection network of the present invention on an open-source high-resolution dataset. It can be seen that the changed areas are well detected. The detection effect of the present invention can be further illustrated by a comparative experiment. On the LEVIR-CD dataset, the method of the present invention and other existing methods FC-EF, FC-Siam-Di, FC-Siam-Conc, DTCDSCN, SNUNet, BIT, ChangeFormer, RDPNet are compared. Calculate the accuracy Precision, recall Recall, F1-score F1-index, intersection over union IoU, and overall accuracy Overall Accuracy respectively. Among them, the larger the accuracy Precision, the higher the proportion of correct results among all the results predicted as positive; the larger the recall Recall, the higher the proportion of all positive results correctly predicted; the larger the F1-score F1-index, the better the comprehensive evaluation of the results. Table 1 shows the index values of the change detection results of different methods:
[0165] Table 1 Comparison of MDGF and various methods on the LEVIR-CD dataset
[0166]
[0167] It can be seen that the method of the present invention achieves the best accuracy on this dataset. Figure 10 The visual detection results of the above various methods are given. It can be seen that the performance of the method of the present invention is better than other high-resolution image building change detection methods. The method proposed by the present invention can accurately locate the position of the changed area and has an advantage over other methods in adapting to complex backgrounds.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A high-resolution change detection method based on multi-level difference feature grouping fusion, characterized by: The method comprises the following steps: S1: Preprocess the high-resolution images used as training sets to obtain training samples; S2: Input the training samples into the multi-level difference feature grouping fusion network for training; S3: After training is completed, change detection is performed on the test samples to obtain results.
2. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 1 is characterized by: In step S1, the two-phase high-resolution images used as training sets are preprocessed. First, image registration is performed to ensure that the two-phase images are aligned. Then, atmospheric correction, radiation correction, and geometric correction are performed to eliminate atmospheric effects and radiation differences under different sensors and shooting conditions, and to ensure the geometric accuracy of the images. Then, the images were cropped with a non-overlapping window of 256×256 and data augmentation operations were performed, including random flipping, random rescaling, Gaussian blurring, and random color jittering, where the random rescaling ratio ranged from (0.8-1.2); finally, the images were normalized to ensure that the images were comparable at the same scale.
3. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 2 is characterized by: In step S2, the training samples are input into the multi-stage DiffFormer to extract rich multi-scale features. DiffFormer consists of repeated continuous blocks. In the initial block, the two-phase features are first normalized by layer, and then the difference operation is performed to calculate the difference operator. The normalized features and the difference operator of the corresponding phase are input into the differential attention mechanism (DW-MSA) for processing. The formula is as follows: in, is the two-phase input feature, LN is layer normalization, f l is the difference operator, Represents the characteristics of DW-MSA output; Then, the processed two-phase features are differentiated again and input into DW-MSA again with the corresponding phase features for processing to obtain the residual, which is then added to the input features. The formula is as follows: in, is the processed two-phase feature, m l is the difference operator, represents the features after addition; Finally, the added features are subjected to a layer normalization operation and then passed through a multi-layer perceptron, during which the Gaussian error linear unit is used as the nonlinear activation function and added to the added features. The formula is as follows: Among them, MLP is a multi-layer perceptron. Represents the features of the initial block output; The features output by the initial block are input into the subsequent block and similar operations are performed. The difference is that the second attention operation adopts the global differential attention mechanism (GDW-MSA), which introduces a global query component to make up for the shortcomings of local attention without increasing the amount of calculation. The formula is as follows: in, represents the features of the initial block output, represents the features of the subsequent block DW-MSA output, f l+1 、m l+1 are the difference operators of DW-MSA and GDW-MSA in the subsequent blocks, q g1 is the global query component, represents the features after GDW-MSA addition, Represents the features of the subsequent block output.
4. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 3 is characterized by: In DW-MSA, the input features are first converted into query values through a linear layer, and the difference operator is converted into key values and attribute values. Taking the first attention operation of the initial block as an example, the formula is as follows: K=f l W k ,V=f l W v in, represents the input feature, f l represents the difference operator, Q, K, V represent the query value, key value, and attribute value obtained, and W q , W k , W v Represent the projection weights of query value, key value, and attribute value respectively; Then, after the input is divided into h heads in the channel dimension, non-overlapping windows are used to calculate the window attention, and the calculated attention is spliced in the channel dimension to obtain the final output result of DW-MSA. The formula is as follows: Among them, Q n , K n 、V n Indicates the query value, key value, and attribute value of the nth header. n represents the attention of the nth head, softmax() represents the activation function that maps values to the interval [0,1], d represents the dimension of each head, [] represents the concatenation operation on the channel dimension, and A represents the attention value obtained by DW-MSA; Different from DW-MSA, GDW-MSA introduces a global query component. When calculating the attention of each head, the attention weight guided by the global query is added, as follows: Among them, q g1 represents the global query component, represents the attention of the nth head of GDW-MSA; In GDW-MSA, the global query component is pre-calculated by the global query generator at each stage and shared among all blocks to interact with local key-value pairs. First, the height and width of the input feature map are transformed into the height and width of the local window through repeated feature matching modules, and then the global query component with the same resolution size as the local query is obtained through shaping and repetition. The feature matching module mainly includes convolutional layers and compression activation modules, during which the GELU nonlinear activation function is applied, and residual connections are combined to enhance the transmission efficiency of features. The maximum pooling layer is also used for downsampling to achieve gradual dimensionality reduction of features. The formula is as follows: FM(X)=MaxPool(f1(SE(GELU(f3(X)))))+X) Where X is the input feature map, FM(X) is the feature map obtained after the feature matching module, MaxPool() represents the maximum pooling operation, SE represents the compressed activation module, f1 and f3 represent the convolution operations with the convolution kernel size of 1×1 and 3×3 respectively; In the compression activation module, compression, excitation, and multiplication feature fusion operations are used to adjust the importance of each channel, thereby enhancing the representation ability of the module and reducing the computational complexity. The formula is as follows: SE(X)=res(σ(f l (GELU(f l (res(AvgPool(X)))))))·X Among them, X is the input feature map, SE(X) is the feature map obtained after the compression activation module, AvgPool() represents the average pooling operation, and f l represents a linear layer, σ() represents a Sigmoid activation function, and res() represents a reshaping operation.
5. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 4 is characterized by: The feature maps before and after the first stage DiffFormer extraction are input into SCAM respectively. In this module, the feature map first passes through two 3×3 convolutional layers to achieve preliminary feature conversion, and then the feature is enhanced through the multi-scale expansion deformable convolution block. The enhanced feature map is concatenated with the input feature map in the channel dimension, and then passes through a 3×3 convolutional layer to generate residual features. Finally, the residual feature is added to the original input feature map to achieve feature optimization. The formula is as follows: SCAM(X)=f3[MDDC(f3(ReLU(f3(X)))),X]+X Among them, X is the input feature map, SCAM(X) is the feature map obtained after the spatial change perception module, and MDDC is the multi-scale dilated deformable convolution block; In the multi-scale dilated deformable convolution block, the input feature map is processed by a 1×1 convolution layer and three dilated deformable convolution blocks with a convolution kernel size of 3×3 and a dilated deformable convolution block with a dilation rate of 3, 6, and 12 respectively. The feature maps processed by each branch are concatenated along the channel dimension to form the output of the module. The formula is as follows: Among them, X is the input feature map, MDDC(X) is the feature map obtained after the multi-scale dilated deformable convolution block, They represent dilated deformable convolution blocks with a convolution kernel size of 3×3 and dilation rates of 3, 6, and 12 respectively; In each dilated deformable convolution block, two layers of 3×3 dilated convolution layers and ReLU activation function are first used to expand the receptive field. Then, the transformed feature map is added to the input feature map to ensure that the original information details are not lost. Finally, after a 3×3 convolution layer, a deformable convolution layer is introduced to dynamically adjust the position and shape of the convolution kernel, so as to more accurately adapt to the complex and irregular contour features of the changing area. Taking the dilation rate of 3 as an example, the formula is as follows: Among them, X is the input feature map, f 3,3 represents a dilated convolution operation with a kernel size of 3×3 and a dilation rate of 3, f d Represents a deformable convolution operation.
6. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 5 is characterized by: The multi-scale differential feature map is obtained by performing differential operation on the two-phase multi-scale features extracted by the multi-stage DiffFormer and processed by SCAM. The formula is as follows: Among them, T1 and T2 represent two different phases. They represent the features extracted by DiffFormer in the i-th stage at different time phases, It represents the feature map obtained by the difference of the two phase features after the i-th stage DiffFormer. They represent the features processed by SCAM at different phases. It represents the feature map obtained by the difference of the two phase features after SCAM processing in the i-th stage; The low-level difference features Input to the low-level feature fusion module for final adaptive fusion. In the low-level feature fusion module, first The convolutional module consists of a convolutional layer, a batch normalization layer, and a ReLU activation function and is projected and combined with To add, the formula is as follows: Among them, F is the added feature map, f c Represents the convolution module; the added features are then enhanced through the channel and spatial attention modules respectively, and the enhanced features are added, and then the weight map is generated through the Sigmoid activation function. The formula is as follows: w=σ(CA(F)+SA(F)) Among them, F is the added feature map, w is the weight map, σ() represents the Sigmoid activation function, CA represents the channel attention module, and SA represents the spatial attention module; the weight map is respectively compared with the input feature Perform a multiplication operation, and then add the obtained features to the input features Add them together and finally generate the final fusion feature through 1×1 convolution. The formula is as follows: Among them, w is the weight map, f1 represents 1×1 convolution, F l It is the output feature map of the low-level feature fusion module; In the channel attention module, the input feature map first undergoes the maximum pooling and average pooling operations based on length and width respectively. Then, the pooled features are input into two layers of multi-layer perceptrons respectively and added. Then, the weighted coefficients are processed by the Sigmoid activation function to obtain the weighted coefficients. Finally, the weighted coefficients are multiplied element by element with the input feature map to generate the final output of the module. The formula is as follows: CA(X) = σ(MLP(AvgPool(X)) + MLP(MaxPool(X)))·X, where CA(X) is the feature map obtained after the channel attention module, AvgPool(X) and MaxPool(X) represent the feature maps obtained after the average pooling operation and the maximum pooling operation, respectively, and MLP represents a two-layer multilayer perceptron; In the spatial attention module, the input feature map first undergoes channel-based maximum pooling and average pooling operations, then the pooled features are concatenated along the channel dimension, followed by a 7×7 convolution for dimensionality reduction, and finally, after being processed by the Sigmoid activation function, it is multiplied with the input feature map to obtain the final output of the module. The formula is as follows: SA(X)=σ(f7[AvgPool(X),MaxPool(X)])·X Among them, SA(X) is the feature map obtained after the spatial attention module, and f7 represents the convolution operation with a convolution kernel size of 7×7.
7. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 6 is characterized by: High-level difference features Input to the high-level feature fusion module for final fusion. In the high-level feature fusion module, the input feature First, it is processed by the convolution module, and then the resolution is improved by the upsampling module composed of convolution layers. The upsampled features are combined with another feature The concatenation is performed on the channel dimension, and then processed by the same convolution module and upsampling module and then combined with another feature Finally, the obtained features are spliced and the resolution is adjusted through the convolution module. After being processed by the normalized attention module, the final fusion features are output. The formula is as follows: Among them, U represents the upsampling module, NAM represents the normalized attention module, and F h It is the output feature map of the high-level feature fusion module.
8. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 7 is characterized by: The high-level and low-level fused features are input into the feature merging module for effective interaction to form a more comprehensive feature representation. In the feature merging module, the low-level fused features are first passed through the average pooling layer and two layers of 1×1 convolution, and then multiplied element-by-element with themselves, and then processed by the Softmax activation function to obtain the enhanced shallow fusion features. The formula is as follows: F l ′=softmax(f 1×2 (AvgPool(F l ))·F l ) Among them, f 1×2 represents two layers of 1×1 convolution operations, F l ′ is the enhanced shallow fusion feature; At the same time, the high-level fusion features are input into the dilated convolution modules with a convolution kernel size of 3×3 and dilation rates of 1, 3, and 5. The features output by each module are concatenated along the channel dimension, and then integrated through a 1×1 convolution module to reduce the number of channels and enhance the expressiveness of the features. Finally, the enhanced deep fusion features are obtained. The formula is as follows: F h ′=f c1 [f c3,1 (F h ),f c3,3 (F h ),f c3,5 (F h )] Among them, f c3,1 、f c3,3 、f c3,5 They represent the dilated convolution modules with a kernel size of 3×3 and dilation rates of 1, 3, and 5, respectively. c1 represents a 1×1 convolutional module, F h ′ is the enhanced deep fusion feature; Finally, the enhanced low-level fusion features are processed by 1×1 convolution and matrix multiplied with the enhanced high-level fusion features. Then, the result of the matrix multiplication is normalized by the Softmax activation function and added to the high-level fusion features. The formula is as follows: ″″ F l =softmax(f1(F l )×F h )+F h Among them, F l ″ is the processed low-level fusion feature; Similarly, the same operation process is performed on the enhanced high-level fusion features. Finally, the features obtained by the above two steps are concatenated in the channel dimension and integrated through a 1×1 convolution module. Then, they are input into the normalized attention module to generate the final merged features. After passing through the upsampling module and the classification layer, the final change detection result is obtained. The formula is as follows: ″″ F c =C(U(NAM(f c1 [F l ,F h ]))) Among them, F h ″ is the processed high-level fusion feature, C represents the classification layer, and F c The final change detection result.
9. The high-resolution change detection method of multi-level difference feature grouping fusion according to claim 8 is characterized by: The loss function of the overall network is set as follows: Loss=-[X·logY+(1-X)·log(1-Y)] Wherein, X represents the predicted change detection result, Y represents the ground truth label of a given sample, and the loss function is calculated. The model parameters of the change detection framework are optimized according to the loss function and the back propagation process. After the training is completed, a trained change detection framework is obtained. The input sample is discriminated by the trained change detection framework, and a change detection result graph is output.
Citation Information
Patent Citations
Remote sensing image change detection method based on adaptive Transform and deformable convolution
CN119418204A
Evaluation-based speaker change detection evaluation metrics
US20240135934A1