A remote sensing image change detection method and system based on a perception weighted key-value framework

CN122416277BActive Publication Date: 2026-09-22NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610873752.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-22
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0004]本发明的目的是克服现有技术的不足,为更好的有效解决现有的遥感图像变化检测技术中传统卷积神经网络难以捕获全局空间关系及视觉Transformer在处理高分辨率图像时计算与存储开销极高的问题,提供了一种基于感知加权键值框架的遥感图像变化检测方法及系统,其实现了具有采用感知加权键值框架进行遥感图像变化检测的功能,且通过图像内空间特征聚合与图像间时间特征聚合能显式解耦并处理空间对齐与时序差异判别,不仅能在空间维度上整合多尺度特征并缓解因尺度差异导致的特征错位,还能在时序维度上采用交叉注意力机制进行双时相特征的自适应交互与比较,显著提升了在复杂场景下的变化检测精度

Benefits of technology

(1)、本发明的一种基于感知加权键值框架的遥感图像变化检测方法及系统,首先对输入的双时遥感图像进行卷积下采样并获得双时初始特征图,接着对双时初始特征图进行四向特征平移并获得双时平移后特征图,再对双时初始特征图和双时平移后特征图进行空间混合并获得双时空间混合后特征图,随后采用RWKV编码器对双时空间混合后特征图进行双向WKV注意力的空间归纳偏置并获得双时空间归纳后特征图,再对双时空间归纳后特征图进行通道依赖显示与特征响应重新标定从而获得双时通道调制后特征图,然后对双时通道调制后特征图进行下采样并分别获得四个不同尺度的双时分层特征图,再对双时分层特征图进行图像内空间特征聚合与图像间时间特征聚合并生成变化感知特征图,最后对变化感知特征图采用解码器进行逐步上采样并生成变化检测结果图;有效的实现了该基于感知加权键值框架的遥感图像变化检测方法及系统具有采用基于感知加权键值RWKV线性复杂度框架作为核心特征提取器进行训练并行化与推理线性化统一的功能,不仅克服了传统视觉Transformer在高分辨率遥感图像上的二次计算复杂度瓶颈,还使得能在较低的参数量和计算开销下高效建模图像中的长程上下文依赖,这为精准识别变化区域奠定了坚实基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416277B_ABST
    Figure CN122416277B_ABST
Patent Text Reader

Abstract

The application discloses a kind of remote sensing image change detection method and system based on perception weighted key value framework, first the convolution downsampling of input double-time remote sensing image is carried out and double-time initial feature map is obtained, then four-way feature translation is carried out to double-time initial feature map and double-time translation after feature map is obtained, spatial mixing is carried out to double-time initial feature map and double-time translation after feature map and double-time spatial mixing after feature map is obtained;The application realizes the function of using perception weighted key value framework to carry out remote sensing image change detection, and through image space feature aggregation and inter-image time feature aggregation, spatial alignment and time sequence difference discrimination can be explicitly decoupled and processed, not only can multi-scale features be integrated in spatial dimension and the feature misplacement caused by scale difference be relieved, but also adaptive interaction and comparison of double-time phase features can be carried out in time sequence dimension using cross attention mechanism, which significantly improves the change detection accuracy in complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image detection technology, specifically to a method and system for detecting changes in remote sensing images based on a perceptual weighted key value framework. Background Technology

[0002] Remote sensing image change detection is a technique for identifying scene differences in remote sensing images across multiple time periods, and it is of great significance in fields such as environmental monitoring, urban planning, and disaster assessment. With the widespread adoption of high-resolution remote sensing imagery, this technology faces new challenges: on the one hand, high-resolution images require models to have the ability to capture global context; on the other hand, scenarios such as rapid post-disaster response using drones place strict requirements on lightweight and low-latency models.

[0003] Currently, most existing remote sensing image change detection methods use convolutional neural networks (CNNs) for change detection. While these methods offer high computational efficiency, they are limited by the local receptive field of convolution operations, making it difficult to capture the global spatial relationships required for complex changes. To overcome this limitation, most existing remote sensing image change detection methods introduce the Transformer architecture. The Transformer architecture uses a self-attention mechanism to explicitly model long-distance dependencies and improve accuracy. However, the secondary complexity of self-attention leads to high computational and storage overhead when processing high-resolution remote sensing images, severely limiting its application in resource-constrained scenarios. Therefore, it is necessary to design a remote sensing image change detection method and system based on a perceptual weighted key-value framework. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and to better and more effectively solve the problems of traditional convolutional neural networks' inability to capture global spatial relationships and the extremely high computational and storage overhead of visual Transformers when processing high-resolution images in existing remote sensing image change detection technologies. This invention provides a remote sensing image change detection method and system based on a perceptual weighted key-value framework. It achieves the function of remote sensing image change detection using a perceptual weighted key-value framework, and through intra-image spatial feature aggregation and inter-image temporal feature aggregation, it can explicitly decouple and handle spatial alignment and temporal difference discrimination. It can not only integrate multi-scale features in the spatial dimension and alleviate feature misalignment caused by scale differences, but also use a cross-attention mechanism in the temporal dimension for adaptive interaction and comparison of dual-temporal features, significantly improving the change detection accuracy in complex scenes.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A remote sensing image change detection method based on a perceptual weighted key-value framework includes the following steps: Step A: Perform convolutional downsampling on the input dual-time remote sensing image to obtain the dual-time initial feature map; Step B involves performing a four-way feature translation on the initial dual-time feature map to obtain the dual-time translated feature map, and then spatially blending the initial dual-time feature map and the dual-time translated feature map to obtain the dual-time spatially blended feature map. Step C: Use an RWKV encoder to apply bidirectional WKV attention spatial induction bias to the bi-temporal-space hybrid feature map and obtain the bi-temporal-space induction feature map. Step D: Perform channel-dependent display and feature response recalibration on the dual-temporal-space inductive feature map to obtain the dual-temporal-channel modulated feature map; Step E: Downsample the modulated feature map of the dual-time channels and obtain four dual-time layered feature maps of different scales respectively; Step F involves performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-time hierarchical feature map to generate a change-aware feature map. Step G involves using a decoder to progressively upsample the change-aware feature map and generate a change detection result map.

[0006] The aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework includes step A, which involves convolutional downsampling of the input dual-time remote sensing image to obtain a dual-time initial feature map, wherein the dual-time remote sensing image includes the first-time remote sensing image. Second time remote sensing image The dual-time initial feature map includes a first time initial feature map. Second initial feature map .

[0007] The aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework includes step B, which involves performing a four-directional feature translation on the initial dual-time feature map to obtain a dual-time translated feature map, and then spatially blending the initial dual-time feature map and the dual-time translated feature map to obtain a dual-time spatially blended feature map. The specific steps are as follows. Step B1 involves performing a four-way feature translation on the initial feature map of the two time periods to obtain the feature map after the two time periods are translated. Specifically, this involves translating the initial feature map of the first time period... Second initial feature map The channel dimensions are divided into four groups, and each group is spatially translated along the up, down, left and right directions to obtain the feature map after dual-time translation, as shown in formula (1). (1) in, This is the feature map after double time translation. , , and For spatial indexing, For splicing operations, The input feature map Includes the first initial feature map Second initial feature map The input feature map , The height of the feature map, The width of the feature map. The number of channels in the feature map. Slice the channel; Step B2 involves spatially blending the initial feature map and the translated feature map to obtain the spatially blended feature map, as shown in formula (2). (2) in, This is a feature map after dual temporal-space mixing. It is a four-dimensional displacement function. The key component quantities are those for spatial mixing processes, and the key component quantities are those for spatial mixing processes. , These are the acceptance vector, key vector, and value vector, respectively. It is a gated vector, and the gated vector The gate vector Used to balance the initial feature map and the feature map after the bi-time translation.

[0008] In the aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework, step C involves using an RWKV encoder to apply a bidirectional WKV attention spatial inductive bias to the bitemporal-space mixed feature map and obtaining the bitemporal-space inductive feature map. The specific steps are as follows. Step C1: Construct feature units based on the feature maps after dual temporal-space mixing, as shown in formula (3). (3) in, The key component quantities after four-dimensional displacement. The projection matrix to be learned; Step C2 involves using feature units to perform bidirectional recursion of WKV attention and obtaining the bidirectional WKV attention recursion result, as shown in formula (4). (4) in, This is the result of bidirectional WKV attention recursion. For feature unit index, The length of the feature unit. For the summation index, , , and Key vectors Sum value vector The index is and The corresponding values, and These are the attenuation parameter and the bias parameter, respectively; Step C3: Based on the bidirectional WKV attention recursion result, the feature unit outputs a bi-temporal inductive feature map, as shown in formula (5). (5) in, This is the feature map after bi-temporal induction.

[0009] In the aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework, step D involves performing channel-dependent display and feature response recalibration on the bitemporal-space inductively derived feature map to obtain a bitemporal-channel modulated feature map. The specific steps are as follows: Step D1 involves using global average pooling to extract global spatial information from the feature maps after bitemporal induction and aggregating it into channel descriptors. The specific steps are as follows. (6) in, For channel descriptors, and These are the summation indices for the height and width dimensions, respectively. This is the index of the current feature channel; Step D2, based on the channel descriptor Generate channel statistical vectors, as shown in formula (7). (7) in, This is a channel statistics vector; Step D3 involves using a bottleneck structure consisting of two fully connected layers to process the channel statistical vectors. Channel modulation is performed and the channel modulation weight vector is obtained, as shown in formula (8). (8) in, and These are the weight matrices for a reduced-dimensional fully connected network and the weight matrices for a higher-dimensional fully connected network, respectively. For channel modulation weight vectors, It is the sigmoid function; Step D4, adjust the channel modulation weight vector. Channel multiplication is applied to the feature map after bitemporal induction and the feature response is recalibrated to obtain the feature map after bitemporal channel modulation, as shown in formula (9). (9) in, This is a feature map after dual-time-channel modulation.

[0010] In the aforementioned remote sensing image change detection method based on a perceptual weighted key value framework, step E involves downsampling the modulated feature map of the dual-time channels to obtain four different scaled dual-time layered feature maps. Specifically, this involves downsampling the modulated feature map of the dual-time channels... The input features for the next layer are obtained through downsampling. Then, by repeating step BE, we can obtain the dual-time hierarchical feature map. .

[0011] The aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework, step F, involves performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-temporal layered feature maps to generate a change-aware feature map. The specific steps are as follows: Step F1 involves performing intra-image spatial feature aggregation on the dual-temporal hierarchical feature map to obtain a dual-temporal spatial enhancement feature map. The specific steps are as follows: Step F11, process the dual-time hierarchical feature map Each feature map in All use upsampling to resolution And spliced ​​along the channel dimension into a single sheet ; Step F12, using residual channel mixing on a single sheet Refine and obtain the refined tensor: (10) in, To refine the tensor, This is a channel blending function; Step F13 involves splitting the refined tensor and downsampling it back to the original four scales to generate dual temporal-space enhanced feature maps. ; Step F2 involves aggregating temporal features between images from the dual spatiotemporal enhanced feature maps and generating a change-aware feature map. The specific steps are as follows: Step F21, enhance the dual temporal-space feature maps The channel attention weights are used to calculate the features of one time phase and apply them to another time phase to obtain the dual-time channel attention weight values, as shown in formula (11). (11) in, These are the attention weights for the two time channels. For global average pooling, It is a multilayer perceptron; Step F22, enhance the dual temporal-space feature maps and dual-time channel attention weights Perform cross-refinement to obtain a spatial feature map after dual-temporal cross-refinement, wherein the dual-temporal spatial enhancement feature map Includes first spatiotemporal augmentation features Second spatiotemporal enhancement features The dual-time channel attention weight value Includes first-time channel attention weights Second time channel attention weight value The spatial feature map after dual-time cross refinement It includes the spatial feature map after the first time-cross thinning and the spatial feature map after the second time-cross thinning, as shown in formula (12). ; (12) in, This is the spatial feature map after the first time-cross refinement. This is the spatial feature map after the second time-cross refinement; Step F23, refine the spatial feature map based on the dual-time-crossing method. The bi-temporal attention map is calculated as shown in formula (13). (13) in, This is a dual spatiotemporal attention map. For convolution operations, For average pooling, For max pooling; Step F24: Refine the spatial feature map after dual-time crossover. and bi-temporal attention map Cross-application is performed to generate a change-aware feature map, wherein the dual temporal-space attention map Includes the first spatiotemporal attention map Second spatiotemporal attention map Specifically, as shown in formula (14), (14) in, For change-sensing feature maps.

[0012] In the aforementioned remote sensing image change detection method based on a perceptual weighted key-value framework, step G involves progressively upsampling the change-aware feature map using a decoder to generate a change detection result map. Specifically, the decoder is a lightweight U-Net-style decoder, and the specific steps are as follows. Step G1: Use a decoder to process the change-aware feature map. The process involves layer-by-layer decoding and the generation of intermediate decoding features, as shown in formula (15). (15) in, and The first Layer decoding features and the first Layer decoding features, For feature splicing operations, For upsampling operation, Step G2, using Convolution extracts the highest resolution decoded features from the intermediate decoded features. The change detection result map is obtained by projecting it onto the prediction space and activating it with the Sigmoid function, as shown in formula (16). (16) in, This is a graph showing the results of the change detection.

[0013] A remote sensing image change detection system based on a perceptual weighted key-value framework includes an initial feature acquisition module, a feature translation module, a feature space induction module, a feature modulation module, a hierarchical feature acquisition module, a feature aggregation module, and a feature decoding module. The initial feature acquisition module performs convolutional downsampling on the input dual-temporal remote sensing image to obtain a dual-temporal initial feature map. The feature translation module performs four-way feature translation on the dual-temporal initial feature map to obtain a dual-temporal translated feature map, and then spatially mixes the dual-temporal initial feature map and the dual-temporal translated feature map to obtain a dual-temporal spatially mixed feature map. The feature space induction module uses an RWKV encoder to perform dual-temporal spatial mixing... The post-feature map undergoes spatial induction biasing with bidirectional WKV attention to obtain a bitemporal spatial induction post-feature map; the feature modulation module is used to perform channel-dependent display and feature response recalibration on the bitemporal spatial induction post-feature map to obtain a bitemporal channel modulated post-feature map; the hierarchical feature acquisition module is used to downsample the bitemporal channel modulated post-feature map and obtain four bitemporal hierarchical feature maps at different scales; the feature aggregation module is used to perform intra-image spatial feature aggregation and inter-image temporal feature aggregation on the bitemporal hierarchical feature map to generate a change-aware feature map; the feature decoding module uses a decoder to progressively upsample the change-aware feature map and generate a change detection result map.

[0014] The beneficial effects of this invention are: (1) A remote sensing image change detection method and system based on a perceptual weighted key value framework of the present invention firstly performs convolutional downsampling on the input dual-time remote sensing image to obtain a dual-time initial feature map, then performs four-way feature translation on the dual-time initial feature map to obtain a dual-time translated feature map, then performs spatial mixing on the dual-time initial feature map and the dual-time translated feature map to obtain a dual-time spatially mixed feature map, then uses an RWKV encoder to perform bidirectional WKV attention spatial induction bias on the dual-time spatially mixed feature map to obtain a dual-time spatially induction feature map, then performs channel-dependent display and feature response recalibration on the dual-time spatially induction feature map to obtain a dual-time channel modulated feature map, and then downsamples the dual-time channel modulated feature map to obtain four different scales respectively. The method and system generate a change-aware feature map by performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-time layered feature map. Finally, the change-aware feature map is progressively upsampled using a decoder to generate a change detection result map. This effectively realizes that the remote sensing image change detection method and system based on the perceptual weighted key value framework has the function of unifying parallel training and linear inference by using the perceptual weighted key value RWKV linear complexity framework as the core feature extractor. It not only overcomes the bottleneck of secondary computational complexity of traditional visual Transformer on high-resolution remote sensing images, but also enables efficient modeling of long-range contextual dependencies in images with low parameter quantity and computational overhead. This lays a solid foundation for accurate identification of changed regions.

[0015] (2) This invention can explicitly decouple and process spatial alignment and temporal difference discrimination by aggregating spatial features within an image and aggregating temporal features between images. It can not only integrate multi-scale features in the spatial dimension and alleviate feature misalignment caused by scale differences, but also use a cross-attention mechanism in the temporal dimension to perform adaptive interaction and comparison of dual-temporal features, suppress non-change interference such as illumination and seasonal differences, enhance the sensitivity to real semantic changes, and significantly improve the change detection accuracy in complex scenes.

[0016] (3) The present invention can achieve consistent leading performance in both optical and synthetic aperture radar image SAR change detection tasks. It is not only highly insensitive to imaging mechanisms and noise patterns, but also focuses on extracting the essential characteristics of changes in remote sensing images themselves. Moreover, it does not require specific adjustments or augmentations for images from different sensors. This demonstrates that the remote sensing image change detection method and system have strong generalization ability and modal robustness. The present invention has the potential for practical application in dealing with multi-source heterogeneous remote sensing data. Attached Figure Description

[0017] Figure 1 This is an overall flowchart of a remote sensing image change detection method based on a perceptual weighted key value framework according to the present invention. Figure 2 This is a schematic diagram illustrating the working principle of a remote sensing image change detection system based on a perceptual weighted key value framework according to the present invention. Figure 3 This is a qualitative result diagram on the LEVIR-CD dataset in an embodiment of the present invention; Figure 4 This is a qualitative result diagram on the SAR-CD dataset in an embodiment of the present invention. Detailed Implementation

[0018] The present invention will now be further described with reference to the accompanying drawings.

[0019] like Figure 1 As shown, the present invention provides a remote sensing image change detection method based on a perceptual weighted key-value framework, comprising the following steps: Step A: Convolutional downsampling is performed on the input dual-time remote sensing image to obtain a dual-time initial feature map, wherein the dual-time remote sensing image includes the first-time remote sensing image. Second time remote sensing image The dual-time initial feature map includes a first time initial feature map. Second initial feature map .

[0020] Step B involves performing a four-way feature translation on the initial bi-temporal feature map to obtain the bi-temporal translated feature map. Then, spatial blending is performed on the initial bi-temporal feature map and the bi-temporal translated feature map to obtain the bi-temporal spatially blended feature map. The specific steps are as follows. Step B1 involves performing a four-way feature translation on the initial feature map of the two time periods to obtain the feature map after the two time periods are translated. Specifically, this involves translating the initial feature map of the first time period... Second initial feature map The channel dimensions are divided into four groups, and each group is spatially translated along the up, down, left and right directions to obtain the feature map after dual-time translation, as shown in formula (1). (1) in, This is the feature map after double time translation. , , and For spatial indexing, For splicing operations, The input feature map Includes the first initial feature map Second initial feature map The input feature map , The height of the feature map, The width of the feature map. The number of channels in the feature map. Slice the channel; Step B2 involves spatially blending the initial feature map and the translated feature map to obtain the spatially blended feature map, as shown in formula (2). (2) in, This is a feature map after dual temporal-space mixing. It is a four-dimensional displacement function. The key component quantities are those for spatial mixing processes, and the key component quantities are those for spatial mixing processes. , These are the acceptance vector, key vector, and value vector, respectively. It is a gated vector, and the gated vector The gate vector Used to balance the initial feature map and the feature map after the bi-time translation.

[0021] Step C involves using an RWKV encoder to apply a bidirectional WKV attention spatial inductive bias to the bi-temporal-space mixed feature map and obtaining the bi-temporal-space inductive feature map. The specific steps are as follows. Step C1: Construct feature units based on the feature maps after dual temporal-space mixing, as shown in formula (3). (3) in, The key component quantities after four-dimensional displacement. The projection matrix to be learned; Step C2 involves using feature units to perform bidirectional recursion of WKV attention and obtaining the bidirectional WKV attention recursion result, as shown in formula (4). (4) in, This is the result of bidirectional WKV attention recursion. For feature unit index, The length of the feature unit. For the summation index, , , and Key vectors Sum value vector The index is and The corresponding values, and These are the attenuation parameter and the bias parameter, respectively; Step C3: Based on the bidirectional WKV attention recursion result, the feature unit outputs a bi-temporal inductive feature map, as shown in formula (5). (5) in, This is the feature map after bi-temporal induction.

[0022] Step D involves performing channel-dependent display and feature response recalibration on the bitemporal-space inductive feature map to obtain the bitemporal-channel modulated feature map. The specific steps are as follows. Step D1 involves using global average pooling to extract global spatial information from the feature maps after bitemporal induction and aggregating it into channel descriptors. The specific steps are as follows. (6) in, For channel descriptors, and These are the summation indices for the height and width dimensions, respectively. This is the index of the current feature channel; Step D2, based on the channel descriptor Generate channel statistical vectors, as shown in formula (7). (7) in, This is a channel statistics vector; Step D3 involves using a bottleneck structure consisting of two fully connected layers to process the channel statistical vectors. Channel modulation is performed and the channel modulation weight vector is obtained, as shown in formula (8). (8) in, and These are the weight matrices for a reduced-dimensional fully connected network and the weight matrices for a higher-dimensional fully connected network, respectively. For channel modulation weight vectors, It is the sigmoid function; Step D4, adjust the channel modulation weight vector. Channel multiplication is applied to the feature map after bitemporal induction and the feature response is recalibrated to obtain the feature map after bitemporal channel modulation, as shown in formula (9). (9) in, This is a feature map after dual-time-channel modulation.

[0023] Step E involves downsampling the modulated feature map from the dual-time channels to obtain four different scaled dual-time layered feature maps. Specifically, this involves downsampling the modulated feature map from the dual-time channels. The input features for the next layer are obtained through downsampling. Then, by repeating step BE, we can obtain the dual-time hierarchical feature map. .

[0024] Step F involves performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-temporal hierarchical feature maps to generate a change-aware feature map. The specific steps are as follows. Step F1 involves performing intra-image spatial feature aggregation on the dual-temporal hierarchical feature map to obtain a dual-temporal spatial enhancement feature map. The specific steps are as follows: Step F11, process the dual-time hierarchical feature map Each feature map in All use upsampling to resolution And spliced ​​along the channel dimension into a single sheet ; Step F12, using residual channel mixing on a single sheet Refine and obtain the refined tensor: (10) in, To refine the tensor, This is a channel blending function; Step F13 involves splitting the refined tensor and downsampling it back to the original four scales to generate dual temporal-space enhanced feature maps. ; Step F2 involves aggregating temporal features between images from the dual spatiotemporal enhanced feature maps and generating a change-aware feature map. The specific steps are as follows: Step F21, enhance the dual temporal-space feature maps The channel attention weights are used to calculate the features of one time phase and apply them to another time phase to obtain the dual-time channel attention weight values, as shown in formula (11). (11) in, These are the attention weights for the two time channels. For global average pooling, It is a multilayer perceptron; Step F22, enhance the dual temporal-space feature maps and dual-time channel attention weights Perform cross-refinement to obtain a spatial feature map after dual-temporal cross-refinement, wherein the dual-temporal spatial enhancement feature map Includes first spatiotemporal augmentation features Second spatiotemporal enhancement features The dual-time channel attention weight value Includes first-time channel attention weights Second time channel attention weight value The spatial feature map after dual-time cross refinement It includes the spatial feature map after the first time-cross thinning and the spatial feature map after the second time-cross thinning, as shown in formula (12). ; (12) in, This is the spatial feature map after the first time-cross refinement. This is the spatial feature map after the second time-cross refinement; Step F23, refine the spatial feature map based on the dual-time-crossing method. The bi-temporal attention map is calculated as shown in formula (13). (13) in, This is a dual spatiotemporal attention map. For convolution operations, For average pooling, For max pooling; Step F24: Refine the spatial feature map after dual-time crossover. and bi-temporal attention map Cross-application is performed to generate a change-aware feature map, wherein the dual temporal-space attention map Includes the first spatiotemporal attention map Second spatiotemporal attention map Specifically, as shown in formula (14), (14) in, For change-sensing feature maps.

[0025] Step G involves progressively upsampling the change-aware feature map using a decoder to generate a change detection result map. Specifically, the decoder is a lightweight U-Net-style decoder, and the specific steps are as follows. Step G1: Use a decoder to process the change-aware feature map. The process involves layer-by-layer decoding and the generation of intermediate decoding features, as shown in formula (15). (15) in, and The first Layer decoding features and the first Layer decoding features, For feature splicing operations, For upsampling operation, Step G2, using Convolution extracts the highest resolution decoded features from the intermediate decoded features. The change detection result map is obtained by projecting it onto the prediction space and activating it with the Sigmoid function, as shown in formula (16). (16) in, This is a graph showing the results of the change detection.

[0026] like Figure 2 As shown, a remote sensing image change detection system based on a perceptual weighted key-value framework includes an initial feature acquisition module, a feature translation module, a feature space induction module, a feature modulation module, a hierarchical feature acquisition module, a feature aggregation module, and a feature decoding module. The initial feature acquisition module performs convolutional downsampling on the input dual-time remote sensing image to obtain a dual-time initial feature map. The feature translation module performs four-way feature translation on the dual-time initial feature map to obtain a dual-time translated feature map, and then spatially mixes the dual-time initial feature map and the dual-time translated feature map to obtain a dual-time spatially mixed feature map. The feature space induction module uses an RWKV encoder to perform dual-time spatial... The hybrid feature map is spatially inductively biased with bidirectional WKV attention to obtain a bitemporal spatially inductive feature map; the feature modulation module is used to perform channel-dependent display and feature response recalibration on the bitemporal spatially inductive feature map to obtain a bitemporal channel modulated feature map; the hierarchical feature acquisition module is used to downsample the bitemporal channel modulated feature map to obtain four bitemporal hierarchical feature maps at different scales; the feature aggregation module is used to perform intra-image spatial feature aggregation and inter-image temporal feature aggregation on the bitemporal hierarchical feature map to generate a change-aware feature map; the feature decoding module uses a decoder to progressively upsample the change-aware feature map and generate a change detection result map.

[0027] To illustrate the effectiveness of the present invention, a specific embodiment of the remote sensing image change detection method and system of the present invention is described below.

[0028] To verify the rationality and effectiveness of the method of this invention, this embodiment conducts simulation experiments on the publicly available optical remote sensing image change detection datasets LEVIR-CD and WHU-CD, as well as the synthetic aperture radar image change detection dataset SAR-CD. The effectiveness of the algorithms is evaluated by comparing the intersection-over-union ratio (IoU), F1 score, precision (P), and recall (R) of each algorithm, which are defined as follows: ; ; ; ; Among them, TP, FP, and FN represent the number of true positives, false positives, and false negatives, respectively; all indicators are calculated at the pixel level; these standards provide a comprehensive evaluation of detection accuracy and positioning quality.

[0029] To achieve a trade-off between accuracy and efficiency, this embodiment uses three model instances of different sizes: a minimum T-type, a small S-type, and a regular B-type. They differ in embedding dimension and encoder depth. The specific configuration is as follows: (1) The embedding dimension of the T-type in the method of the present invention is [32,48,96,160], and the encoder depth is [2,2,4,2].

[0030] (2) The embedding dimension of the S-type in the method of the present invention is [32,64,128,192], and the encoder depth is [3,3,6,3].

[0031] (3) The embedding dimension of the B type of the method of the present invention is [48,72,144,240], and the encoder depth is [3,3,6,3].

[0032] The comparison results of different algorithms in this embodiment on the LEVIR-CD dataset and the WHU-CD dataset are shown in Table 1.

[0033] Table 1. Comparison results of different algorithms on the LEVIR-CD and WHU-CD datasets;

[0034] All indicators are expressed as percentages, with the highest value indicated in bold and the second highest value indicated by an underline.

[0035] like Figure 3 As shown, the color coding of the prediction output is as follows: white represents true positive (TP), black represents true negative (TN), green represents false positive (FP), and red represents false negative (FN).

[0036] The comparison results of different algorithms on the SAR-CD dataset in this embodiment are shown in Table 2.

[0037] Table 2. Comparison results of different algorithms on the SAR-CD dataset;

[0038] All metrics are expressed as percentages, with the highest values ​​highlighted in bold.

[0039] like Figure 4 As shown, the color coding of the prediction output is as follows: white represents true positive (TP), black represents true negative (TN), green represents false positive (FP), and red represents false negative (FN).

[0040] The performance and technical effects of the ChangeRWKV model proposed in this invention will be described below with reference to the embodiments and experimental results.

[0041] As shown in Table 1, the B-type method of this invention achieves new state-of-the-art performance on the widely used LEVIR-CD optical dataset, with an Intersection over Union (IoU) of 85.46% and an F1 score of 92.16%, significantly outperforming existing state-of-the-art methods. More importantly, the lightweight T-type model of this invention, containing only 4.66M parameters and 9.40 G FLOPs, achieves an IoU of 84.92%, surpassing most existing methods. On the WHU-CD dataset, the medium-sized S-type model further achieves an IoU of 90.06% and an F1 score of 94.77%, fully demonstrating the effectiveness and scalability of the proposed architecture in high-resolution optical imagery.

[0042] As shown in Table 2, the differential learning capability of the model was validated on the SAR-CD dataset, which contains synthetic variations and significant speckle noise. The method of this invention exhibits excellent generalization performance on SAR imagery, with the base model achieving an IoU of 97.18% and an F1 score of 98.57%, setting a new benchmark. Even the lightweight T-mode and small S-mode versions outperform previous state-of-the-art methods with a significantly reduced number of parameters, demonstrating that the model of this invention can learn robust and mode-independent variation representations without requiring specific tuning for SAR data.

[0043] In addition, such as Figure 3 The visualization comparison results further validate the superiority of the method presented in this invention. Compared with other methods, this invention generates a clearer and more complete change detection mask, especially excelling in complex scenes. This invention accurately characterizes the boundaries of small, irregular buildings in dense urban environments and maintains high fidelity for image edge targets, fully demonstrating the robustness of the hierarchical encoder and fusion module. Simultaneously, this invention effectively suppresses spurious predictions caused by background clutter and exhibits strong tolerance to common annotation noise, thus obtaining clean and reliable change detection results. This robustness is also applicable to different sensor modalities.

[0044] like Figure 4 As shown, even in the presence of significant speckle noise and non-semantic variations in SAR data, the method of this invention can still generate clear and accurate detection results, demonstrating the good generalization ability of the framework of this invention under heterogeneous conditions. This invention can learn variation patterns independent of sensor type.

[0045] In summary, the method of the present invention demonstrates efficient, robust, and scalable change detection performance in both optical and SAR images, fully verifying the technical effectiveness and application potential of the remote sensing image change detection method and system of the present invention.

[0046] Therefore, the remote sensing image change detection method and system of this invention breaks through the quadratic complexity of traditional ViT in terms of method complexity, achieves near-linear time scalability, and significantly improves scalability and edge device deployment capability under high-resolution imagery. In terms of change detection performance, it can effectively suppress spurious changes caused by illumination, registration errors and noise, and maintain clear boundaries and regional integrity. At the same time, it has significant improvements over existing methods in key indicators such as cross-union ratio and F1 score, making its change detection performance superior.

[0047] In summary, the remote sensing image change detection method and system based on a perceptual weighted key-value framework of the present invention first performs convolutional downsampling on the input dual-time remote sensing image to obtain a dual-time initial feature map. Then, it performs four-way feature translation on the dual-time initial feature map to obtain a dual-time translated feature map. Next, it spatially mixes the dual-time initial feature map and the dual-time translated feature map to obtain a dual-time spatially mixed feature map. Subsequently, it uses an RWKV encoder to apply bidirectional WKV attention spatial induction bias to the dual-time spatially mixed feature map to obtain a dual-time spatially induction feature map. Finally, it performs channel-dependent display and feature extraction on the dual-time spatially induction feature map. The system responds to recalibration to obtain a dual-time-channel modulated feature map. Then, it downsamples the dual-time-channel modulated feature map to obtain four different-scale dual-time-layered feature maps. Next, it performs intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-time-layered feature maps to generate a change-aware feature map. Finally, it uses a decoder to progressively upsample the change-aware feature map and generate a change detection result map. This effectively realizes the remote sensing image change detection method and system based on a perceptual weighted key value framework. It employs a perceptual weighted key value RWKV linear complexity framework as the core feature extractor for parallelized training and linearized inference. This invention not only overcomes the bottleneck of secondary computational complexity in high-resolution remote sensing images by traditional visual Transformers, but also enables efficient modeling of long-range contextual dependencies in images with lower parameter count and computational overhead, laying a solid foundation for accurate identification of changing regions. Furthermore, by explicitly decoupling and handling spatial alignment and temporal difference discrimination through intra-image spatial feature aggregation and inter-image temporal feature aggregation, this invention not only integrates multi-scale features in the spatial dimension and alleviates feature misalignment caused by scale differences, but also employs a cross-attention mechanism in the temporal dimension for adaptive interaction and comparison of bi-temporal features, suppressing... Non-changing disturbances such as illumination and seasonal differences enhance the sensitivity to real semantic changes, significantly improving the accuracy of change detection in complex scenes. This invention achieves consistently leading performance in both optical and synthetic aperture radar (SAR) image change detection tasks. It is not only highly insensitive to imaging mechanisms and noise patterns, but also focuses on extracting the essential characteristics of changes in remote sensing images themselves, without requiring specific adjustments or augmentations for different sensor images. This demonstrates that the remote sensing image change detection method and system have strong generalization ability and modal robustness. This invention has the potential for practical application in dealing with multi-source heterogeneous remote sensing data.

[0048] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting changes in remote sensing images based on a perceptual weighted key-value framework, characterized in that: Includes the following steps, Step A: Perform convolutional downsampling on the input dual-time remote sensing image to obtain the dual-time initial feature map; Step B involves performing a four-way feature translation on the initial bi-temporal feature map to obtain the bi-temporal translated feature map. Then, spatial blending is performed on the initial bi-temporal feature map and the bi-temporal translated feature map to obtain the bi-temporal spatially blended feature map. The specific steps are as follows. Step B1 involves performing a four-way feature translation on the initial feature map of the two time periods to obtain the feature map after the two time periods are translated. Specifically, this involves translating the initial feature map of the first time period... Second initial feature map The channel dimensions are divided into four groups, and each group is spatially translated along the up, down, left and right directions to obtain the feature map after dual-time translation, as shown in formula (1). (1) in, This is the feature map after dual time translation. , , and For spatial indexing, For splicing operations, The input feature map, the input feature map Includes the first initial feature map Second initial feature map The input feature map , The height of the feature map, The width of the feature map. The number of channels in the feature map. Slice the channel; Step B2 involves spatially blending the initial feature map and the translated feature map to obtain the spatially blended feature map, as shown in formula (2). (2) in, This is a feature map after dual temporal-space mixing. It is a four-dimensional displacement function. The key component quantities are those for spatial mixing processes, and the key component quantities are those for spatial mixing processes. , These are the acceptance vector, key vector, and value vector, respectively. It is a gated vector, and the gated vector The gate vector Used to balance the initial feature map and the translated feature map in both time periods; Step C involves using an RWKV encoder to apply a bidirectional WKV attention spatial inductive bias to the bi-temporal-space mixed feature map and obtaining the bi-temporal-space inductive feature map. The specific steps are as follows. Step C1: Construct feature units based on the feature maps after dual temporal-space mixing, as shown in formula (3). (3) in, The key component quantities after four-dimensional displacement. The projection matrix to be learned; Step C2 involves using feature units to perform bidirectional recursion of WKV attention and obtaining the bidirectional WKV attention recursion result, as shown in formula (4). (4) in, This is the result of bidirectional WKV attention recursion. For feature unit index, The length of the feature unit. For the summation index, , , and Key vectors Sum value vector The index is and The corresponding values, and These are the attenuation parameter and the bias parameter, respectively; Step C3: Based on the bidirectional WKV attention recursion result, the feature unit outputs a bi-temporal inductive feature map, as shown in formula (5). (5) in, This is a feature map after bi-temporal induction; Step D: Perform channel-dependent display and feature response recalibration on the dual-temporal-space inductive feature map to obtain the dual-temporal-channel modulated feature map; Step E: Downsample the modulated feature map of the dual-time channels and obtain four dual-time layered feature maps of different scales respectively; Step F involves performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-time hierarchical feature map to generate a change-aware feature map. Step G involves using a decoder to progressively upsample the change-aware feature map and generate a change detection result map.

2. The remote sensing image change detection method based on a perceptual weighted key value framework according to claim 1, characterized in that: Step A: Convolutional downsampling is performed on the input dual-time remote sensing image to obtain a dual-time initial feature map, wherein the dual-time remote sensing image includes the first-time remote sensing image. Second time remote sensing image The dual-time initial feature map includes a first time initial feature map. Second initial feature map .

3. The remote sensing image change detection method based on a perceptual weighted key value framework according to claim 2, characterized in that: Step D involves performing channel-dependent display and feature response recalibration on the bitemporal-space inductive feature map to obtain the bitemporal-channel modulated feature map. The specific steps are as follows. Step D1 involves using global average pooling to extract global spatial information from the feature maps after bitemporal induction and aggregating it into channel descriptors. The specific steps are as follows. (6) in, For channel descriptors, and These are the summation indices for the height and width dimensions, respectively. This is the index of the current feature channel; Step D2, based on the channel descriptor Generate channel statistical vectors, as shown in formula (7). (7) in, This is a channel statistics vector; Step D3 involves using a bottleneck structure consisting of two fully connected layers to process the channel statistical vectors. Channel modulation is performed and the channel modulation weight vector is obtained, as shown in formula (8). (8) in, and These are the weight matrices for a reduced-dimensional fully connected network and the weight matrices for a higher-dimensional fully connected network, respectively. For channel modulation weight vectors, It is the sigmoid function; Step D4, adjust the channel modulation weight vector. Channel multiplication is applied to the feature map after bitemporal induction and the feature response is recalibrated to obtain the feature map after bitemporal channel modulation, as shown in formula (9). (9) in, This is a feature map after dual-time-channel modulation.

4. The remote sensing image change detection method based on a perceptual weighted key-value framework according to claim 3, characterized in that: Step E involves downsampling the modulated feature map from the dual-time channels to obtain four different scaled dual-time layered feature maps. Specifically, this involves downsampling the modulated feature map from the dual-time channels. The input features for the next layer are obtained through downsampling. Then, by repeating step BE, we can obtain the dual-time hierarchical feature map. .

5. The remote sensing image change detection method based on a perceptual weighted key value framework according to claim 4, characterized in that: Step F involves performing intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-temporal hierarchical feature maps to generate a change-aware feature map. The specific steps are as follows. Step F1 involves performing intra-image spatial feature aggregation on the dual-temporal hierarchical feature map to obtain a dual-temporal spatial enhancement feature map. The specific steps are as follows: Step F11, for dual-time hierarchical feature maps Each feature map in All use upsampling to resolution And spliced ​​along the channel dimension into a single sheet ; Step F12, using residual channel mixing on a single sheet Refine and obtain the refined tensor: (10) in, To refine the tensor, This is a channel blending function; Step F13 involves splitting the refined tensor and downsampling it back to the original four scales to generate dual temporal-space enhanced feature maps. ; Step F2 involves aggregating temporal features between images from the dual spatiotemporal enhanced feature maps and generating a change-aware feature map. The specific steps are as follows: Step F21, enhance the dual temporal-space feature maps The channel attention weights are used to calculate the features of one time phase and apply them to another time phase to obtain the dual-time channel attention weight values, as shown in formula (11). (11) in, This represents the attention weight values ​​for both time channels. For global average pooling, It is a multilayer perceptron; Step F22, enhance the dual temporal-space feature maps and dual-time channel attention weights Perform cross-refinement to obtain a spatial feature map after dual-temporal cross-refinement, wherein the dual-temporal spatial enhancement feature map Includes first spatiotemporal augmentation features Second spatiotemporal enhancement features The dual-time channel attention weight value Includes first-time channel attention weights Second time channel attention weight value The spatial feature map after dual-time cross refinement It includes the spatial feature map after the first time-cross thinning and the spatial feature map after the second time-cross thinning, as shown in formula (12). ; (12) in, This is the spatial feature map after the first time-cross refinement. This is the spatial feature map after the second time-cross refinement; Step F23, refine the spatial feature map based on the dual-time-crossing method. The bi-temporal attention map is calculated as shown in formula (13). (13) in, This is a dual spatiotemporal attention map. For convolution operations, For average pooling, For max pooling; Step F24: Refine the spatial feature map after dual-time crossover. and bi-temporal attention map Cross-application is performed to generate a change-aware feature map, wherein the dual temporal-space attention map Includes the first spatiotemporal attention map Second spatiotemporal attention map Specifically, as shown in formula (14), (14) in, For change-sensing feature maps.

6. The remote sensing image change detection method based on a perceptual weighted key-value framework according to claim 5, characterized in that: Step G involves progressively upsampling the change-aware feature map using a decoder to generate a change detection result map. Specifically, the decoder is a lightweight U-Net-style decoder, and the specific steps are as follows. Step G1: Use a decoder to process the change-aware feature map. The process involves layer-by-layer decoding and the generation of intermediate decoding features, as shown in formula (15). (15) in, and The first Layer decoding features and the first Layer decoding features, For feature splicing operations, For upsampling operation, Step G2, using Convolution extracts the highest resolution decoded features from the intermediate decoded features. The change detection result map is obtained by projecting it onto the prediction space and activating it with the Sigmoid function, as shown in formula (16). (16) in, This is a graph showing the results of the change detection.

7. A remote sensing image change detection system based on a perceptual weighted key-value framework, wherein the specific detection process of the remote sensing image change detection system is based on the remote sensing image change detection method according to any one of claims 1-6, characterized in that: It includes an initial feature acquisition module, a feature translation module, a feature space induction module, a feature modulation module, a hierarchical feature acquisition module, a feature aggregation module, and a feature decoding module. The initial feature acquisition module is used to perform convolutional downsampling on the input dual-time remote sensing image and obtain a dual-time initial feature map. The feature translation module is used to perform four-way feature translation on the initial feature map of the two time periods and obtain the feature map after the two time periods translation. Then, the initial feature map of the two time periods and the feature map after the two time periods translation are spatially mixed to obtain the feature map after the two time periods spatially mixed. The feature space induction module is used to perform bidirectional WKV attention spatial induction bias on the feature map after bi-temporal space mixing using an RWKV encoder and obtain the feature map after bi-temporal space induction. The feature modulation module is used to perform channel-dependent display and feature response recalibration on the dual-temporal-space inductive feature map to obtain the dual-temporal-channel modulated feature map. The hierarchical feature acquisition module is used to downsample the dual-time channel modulated feature map and obtain four dual-time hierarchical feature maps of different scales respectively; The feature aggregation module is used to perform intra-image spatial feature aggregation and inter-image temporal feature aggregation on the dual-time hierarchical feature map and generate a change-aware feature map. The feature decoding module uses a decoder to progressively upsample the change-aware feature map and generate a change detection result map.

Citation Information

Patent Citations

  • Remote sensing image change detection method and device based on RWKV model

    CN120808134A

  • Modal loss-oriented lightweight self-insight fusion RWKV crack segmentation method and system

    CN122023802A