Onboard compression alignment detection method, system, and storage medium

CN122821391APending Publication Date: 2026-09-25GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611269508.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

若直接存储完整历史影像,随着观测区域、时间跨度和影像分辨率的增加,会带来显著的星上存储压力;若将历史影像压缩为面向视觉质量的码流,又可能无法充分保留对齐和变化判别所需的任务相关信息

Benefits of technology

(1)存储效率高:本发明仅保存历史参考影像的熵编码瓶颈潜变量,无需长期存储完整历史影像,适用于多区域、多时相、长周期的星上监测任务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821391A_ABST
    Figure CN122821391A_ABST
Patent Text Reader

Abstract

The application discloses a kind of on-board compression alignment detection method, system and storage medium, the method is first by sharing encoder to historical reference image is encoded with multiple levels of features, the deepest bottleneck feature is compressed into bottleneck latent variable by quantization and entropy encoding, and is stored in on-board memory;The same operation is performed to the current observation image to obtain the current compressed bottleneck latent variable.Subsequently, in the compressed latent space, based on the history and current latent variable, the global geometric transformation parameter is predicted, and the current latent variable is roughly aligned.Then the latent variable after rough alignment and the history latent variable are input into the shared decoder, and the dense residual displacement field is predicted in the multi-level upsampling process, and the local fine deformation correction is performed on the current feature.Finally, the highest resolution reference feature and the aligned current feature are input into the temporal symmetry fusion module to obtain the fusion feature, and the change detection is performed based on this, and the change probability map is output.The embodiment of the application can realize on-board low-latency change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an on-board compression alignment detection method, system, and storage medium, belonging to the fields of intelligent interpretation of remote sensing images, on-board computing, image compression, image registration, and change detection technology. Background Technology

[0002] Remote sensing image change detection is an important means of monitoring dynamic changes on the Earth's surface, and it is widely used in disaster emergency response, urban expansion monitoring, and ecological environment assessment. Traditional remote sensing change detection typically employs a serial workflow of "on-board imaging—image downlink—ground preprocessing—change detection": after the satellite acquires high-resolution observation images, the complete image or large-scale raw data needs to be downlinked to the ground station, and then radiometric correction, geometric correction, orthorectification, image registration, and change detection model inference are performed sequentially. This workflow can make full use of the powerful computing power of the ground and external auxiliary data, but the end-to-end link is relatively long, and the response time often cannot meet the minute-level or near real-time application requirements of sudden events such as earthquakes, floods, and landslides.

[0003] Existing deep learning-based change detection methods are mostly based on Siamese networks, U-Net-type encoder-decoder structures, Transformers, cross-temporal attention, or multi-scale feature fusion structures. These methods typically assume that the input dual-temporal images have been accurately registered and orthorectified, and they mainly focus on identifying semantically changed regions, paying insufficient attention to translation, rotation, scale changes, perspective distortion, and local nonlinear misalignments that are common in raw satellite observations.

[0004] Existing unregistered change detection schemes can be broadly categorized into two types: the first type is a cascaded scheme combining external registration and change detection, such as first performing registration using manual feature methods like SIFT or ORB, and then inputting the registered dual-temporal images into a change detection network; the second type is a registration-integrated change detection scheme, such as adding perspective transformation prediction, residual migration prediction, or local deformation correction branches to the change detection network. While these schemes alleviate the unregistered problem to some extent, most are geared towards ground-based processing environments and do not simultaneously consider onboard historical image storage, computational complexity, peak memory usage, latency, and end-to-end mission objectives.

[0005] Furthermore, in long-term onboard monitoring, change detection requires the retention of historical reference observations. Directly storing complete historical images would place significant onboard storage pressure as the observation area, time span, and image resolution increase; compressing historical images into a visually quality-oriented bitstream might not adequately preserve the mission-related information needed for alignment and change discrimination. Therefore, how to achieve efficient compressed storage of historical reference data and rapid alignment detection of current observation images under conditions of limited onboard storage, computing power, and communication bandwidth has become an urgent technical problem to be solved. Summary of the Invention

[0006] In view of this, the present invention provides an on-board compressed alignment detection method, system and storage medium, which achieves compact preservation of historical information, robust processing of unregistered information and detection of low-latency changes on the satellite by using a shared encoder-decoder architecture, single bottleneck latent variable storage, global coarse alignment of compressed latent space, progressive local fine alignment at the decoding end and fusion of temporal symmetry features.

[0007] The first objective of this invention is to provide an on-board compression alignment detection method.

[0008] The second objective of this invention is to provide an on-board compression alignment detection system.

[0009] A third objective of this invention is to provide a storage medium.

[0010] The first objective of this invention can be achieved by adopting the following technical solution: An on-board compression alignment detection method, the method comprising: S1: Obtain historical reference images, and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features; S2: Input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. S3: Obtain the current observed image to obtain the current compression bottleneck latent variables; S4: Read the historical compression bottleneck latent variable from the on-board memory, predict the global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment on the current compression bottleneck latent variable according to the global geometric transformation parameters; S5: Input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, predict the dense residual displacement field level by level according to the current feature and reference feature of each level, and perform local deformation correction on the current feature. S6: Input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; S7: Perform change detection based on the fusion features and output a change probability map.

[0011] Further, step S4 includes: S41: Construct a matching-aware representation, wherein the matching-aware representation is the channel concatenation result of the historical compression bottleneck latent variable, the current compression bottleneck latent variable, the absolute difference features of the two, and the element-wise product features of the two. S42: Input the matching perception representation into the global transformation predictor to predict multiple corner offsets, and obtain the homography matrix by solving through direct linear transformation; S43: Based on differentiable space sampling operation, spatial transformation is performed only on the current compression bottleneck latent variable according to the homography matrix to complete coarse alignment.

[0012] Furthermore, step S4 also includes: S44: Construct a mask for the unchanged region using the changed labels; S45: In the absence of supervision by real geometric transformation parameters, the unchanged region mask is used to constrain the reference latent variables within the unchanged region to maintain feature consistency with the coarsely aligned current latent variables, and to eliminate the influence of invalid sampling regions.

[0013] Further, step S5 includes: S51: After upsampling at each level of decoding, a local matching representation is constructed, which includes the reference feature at the current scale, the current feature, the absolute difference feature between the two, and the channel concatenation result of the element-wise product feature of the two. S52: Based on the local matching representation, predict the two-dimensional dense residual displacement field at the current scale, while using a scaling factor to limit the displacement amplitude and using the tanh function to constrain the displacement range; S53: Based on the differentiable spatial sampling operation, spatial correction is performed only on the current feature according to the two-dimensional dense residual displacement field to complete the fine alignment.

[0014] Furthermore, the fine alignment process uses an invariant region feature consistency loss, a displacement smoothing regularization term, and a displacement amplitude regularization term.

[0015] Further, step S6 includes: S61: Perform summation, absolute difference, and element-wise product operations on the highest resolution reference feature and the aligned current feature; S62: Concatenate the computation results along the channel dimension and perform feature mapping through 1×1 convolution to obtain fused features.

[0016] Furthermore, the shared decoder does not obtain shallow skip features from historical images, but only regenerates multi-scale task features step by step based on the historical compression bottleneck latent variables and the coarsely aligned current compression bottleneck latent variables, and simultaneously completes feature alignment during the regeneration process.

[0017] Furthermore, it also includes phased training strategies, including: Compressed pre-training phase: Only the shared encoder, quantization and entropy encoding module, shared decoder, and training reconstruction head are activated to initialize model parameters; Change detection warm-up phase: Remove the reconstruction head, freeze the shared encoder and entropy model, and train the temporal symmetric fusion module and change detection head on the registered dual-temporal samples; Alignment warm-up phase: Activate the global coarse alignment module and the local fine alignment module, freeze the change detection branch, and optimize the alignment module using the consistency loss of the unchanged region; End-to-end fine-tuning stage: Jointly optimize all components of the inference stage, including the shared encoder, quantization and entropy coding module, global coarse alignment module, shared decoder, local fine alignment module, temporal symmetric fusion module, and detection head. The total loss includes change detection loss, bit rate constraint, global alignment loss, and local alignment loss, but does not include reconstruction loss.

[0018] The second objective of this invention can be achieved by adopting the following technical solution: An on-board compression alignment detection system, the system comprising: The first acquisition unit is used to acquire historical reference images and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features. The generation unit is used to input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. The second acquisition unit is used to acquire the currently observed image in order to obtain the current compression bottleneck latent variables; The alignment unit is used to read the historical compression bottleneck latent variable from the on-board memory, predict global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment of the current compression bottleneck latent variable according to the global geometric transformation parameters. The correction unit is used to input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, the dense residual displacement field is predicted level by level according to the current feature and reference feature of each level, and local deformation correction is performed on the current feature. The fusion unit is used to input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; The detection unit is used to perform change detection based on the fused features and output a change probability map.

[0019] The third objective of this invention can be achieved by adopting the following technical solution: A storage medium storing a program that, when executed by a processor, implements the above-described on-board compression alignment detection method.

[0020] The present invention has the following advantages over the prior art: (1) High storage efficiency: This invention only saves the entropy coding bottleneck latent variables of historical reference images, without the need to store complete historical images for a long time, and is suitable for on-board monitoring missions with multiple regions, multiple time phases and long cycles.

[0021] (2) Short end-to-end link: This invention does not require the complete image to be down to the ground for registration. It can directly perform end-to-end change detection on the unregistered dual-temporal image on the satellite, shortening the emergency response link.

[0022] (3) Low computational redundancy: The shared encoding-decoding structure enables compression, alignment and detection to reuse the same feature path, avoiding the repeated calculation of encoding, decoding and feature extraction in traditional cascaded schemes.

[0023] (4) Strong robustness of unregistered alignment: global coarse alignment corrects the main geometric deviations in the compressed latent space, and local fine alignment corrects residual displacements at multiple scales at the decoding end, which can simultaneously handle global perspective differences and local nonlinear misalignments.

[0024] (5) Friendly to real change areas: Use the mask of the unchanged area to constrain the alignment loss, avoid forcibly aligning real change areas such as building demolition and reconstruction, road changes, etc., thereby reducing false changes and missed detections.

[0025] (6) Strong temporal stability: The temporal symmetric fusion structure reduces the impact of input order on change detection output at the mechanism level, and is suitable for data processing with different observation orders in actual tasks.

[0026] (7) Suitable for on-board deployment: The reconstruction head is only used in the training stage and removed in the inference stage. The final model only outputs the change detection results, which effectively reduces the number of parameters, memory usage and end-to-end latency.

[0027] (8) Sufficient experimental verification: The embodiments show that the present invention has achieved high F1 and IoU values ​​on multiple datasets and has low end-to-end latency on embedded platforms such as Jetson TX2. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0029] Figure 1 This is a flowchart of the on-board compression alignment detection method according to Embodiment 1 of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0031] Example 1: like Figure 1 As shown, this embodiment provides an on-board compression alignment detection method, the method comprising: S1: Obtain historical reference images, and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features.

[0032] S2: Input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory.

[0033] S3: Obtain the current observed image to obtain the current compression bottleneck latent variable.

[0034] In this embodiment, the shared encoder, quantization module, and entropy coding module are integrated into the shared coding compression module, which encodes the input image into multi-scale features and compresses the deepest bottleneck features into a compact latent variable bitstream. The on-board storage module is used to store the entropy-coded bottleneck latent variables corresponding to historical reference images, but does not store complete historical images or shallow multi-scale features. Specifically, the image acquisition module acquires remote sensing images of the same area at different time phases, including the historical reference image x_t1 and the currently observed image x_t2. The shared encoder E performs multi-level feature encoding on the historical reference image x_t1 to obtain multi-scale features Z_t1^1, Z_t1^2, Z_t1^3, and Z_t1^4 from shallow to deep. Only the deepest bottleneck feature Z_t1^4 is input into the quantization and entropy coding module Q / EC to obtain the compressed bottleneck latent variable of the historical reference image, i.e., the historical compressed bottleneck latent variable Z~_t1^4, and it is written to the on-board memory. The complete historical image and shallow features Z_t1^1, Z_t1^2, and Z_t1^3 are not saved. When the current observation image x_t2 arrives, the bottleneck feature Z_t2^4 of the current image is generated in real time through the same shared encoder, and then quantized to obtain the current compressed bottleneck latent variable Z~_t2^4.

[0035] It is worth noting that the shared encoder processes historical reference images and current observation images using shared weights, thereby ensuring that the features of the two time phases are in the same representation space. The shared encoder progressively reduces the spatial resolution and improves the semantic abstraction capability to obtain multi-scale features. For historical reference images, the on-board storage strategy in this embodiment is: Memory←Z~_t1^4. This strategy significantly reduces the data storage and transmission overhead of historical images, while the bottleneck latent variables still retain high-level semantic information useful for global geometric alignment and change discrimination.

[0036] It should be noted that while shallow features contain rich spatial details, their large size makes them unsuitable for long-term storage in onboard memory. This embodiment only quantizes and entropy-encodes the deepest bottleneck features to obtain a compact bitstream. Z_t^l = E_l(x_t), Z_t^l = E_l(Z_t^{l-1}),l=2,3,4; Z~_t^4 = Q / EC(Z_t^4) Where x_t represents the remote sensing image at time t, E_l represents the l-th level of the shared encoder, l=1,2,3,4, Z_t^l represents the feature output by the l-th level encoder, and Z~_t^4 represents the compressed bottleneck latent variable after quantization and entropy encoding.

[0037] S4: Read the historical compression bottleneck latent variable from the on-board memory, predict the global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment on the current compression bottleneck latent variable according to the global geometric transformation parameters.

[0038] In this embodiment, step S4 is executed by the global coarse alignment module, which includes a matching-aware representation construction unit, a global transformation predictor, and a differentiable space sampling unit. This module predicts homography or equivalent global geometric transformations within the compressed latent space and performs coarse alignment on the current latent variable. Specifically, it reads the historical compressed bottleneck latent variable Z_t1^4 from the onboard memory, constructs a matching-aware representation with the current compressed bottleneck latent variable Z_t2^4, estimates the global transformation parameters using the global transformation predictor, and coarsely aligns the current latent variable to the reference coordinate system using differentiable space sampling.

[0039] To address the global geometric deviations in dual-temporal images, such as translation, rotation, scaling, shearing, and perspective distortion, this invention achieves global coarse alignment by compressing the bottleneck latent space. This design offers two advantages: first, it has low latent space resolution and low computational cost, making it suitable for onboard deployment scenarios; second, the bottleneck features contain strong semantic information, resulting in greater robustness to illumination changes and local texture differences.

[0040] Further, step S4 includes: S41: Construct a matching-aware representation, wherein the matching-aware representation is the channel concatenation result of the historical compression bottleneck latent variable, the current compression bottleneck latent variable, the absolute difference features of the two, and the element-wise product features of the two.

[0041] This step is performed by the match-aware representation construction unit. The match-aware representation is as follows: m_g =Concat(Z~_t1^4, Z~_t2^4, |Z~_t1^4 - Z~_t2^4|, Z~_t1^4 ⊙ Z~_t2^4) Where m_g represents the matching-aware representation, and Concat represents channel concatenation.

[0042] S42: Input the matching perception representation into the global transformation predictor to predict multiple corner offsets, and obtain the homography matrix by solving through direct linear transformation.

[0043] This step is as follows: theta_g = GTP(m_g), H_g = DLT(dst, dst + theta_g) Where theta_g represents the predicted global geometric transformation parameters, GTP represents the global transformation predictor, H_g represents the homography matrix, DLT represents the direct linear transformation, dst represents the target point coordinates in the reference coordinate system, and dst + theta_g represents the source coordinates after mapping by the global transformation parameters. This embodiment predicts the offsets of four corner points.

[0044] S43: Based on differentiable space sampling operation, spatial transformation is performed only on the current compression bottleneck latent variable according to the homography matrix to complete coarse alignment.

[0045] This step is performed by the differentiable space sampling unit, as shown in the following equation: Zbar_t2^4 = W(Z~_t2^4; H_g) Where Zbar_t2^4 represents the current compression bottleneck latent variable after coarse alignment, and W represents the differentiable space sampling operation.

[0046] As an optional implementation, step S4 includes: First, the historical reference latent variables (dimensions B×C×H×W) and the current latent variables (dimensions B×C×H×W) are combined to construct a four-dimensional matching feature with dimensions B×4C×H×W that is rich in temporal change information.

[0047] Subsequently, this feature is input into a global transform predictor; the global transform predictor consists of a 1×1 convolutional layer (the number of channels is mapped to a preset hidden layer dimension) and a 3×3 convolutional layer (the number of channels is mapped to a preset hidden layer dimension) connected in sequence. Each of them is followed by group normalization (GN) and SiLU activation function. Then, the spatial dimension is compressed to 1×1 through a two-dimensional adaptive average pooling layer (AdaptiveAvgPool2d), resulting in a global representation with a dimension of B×out_hidden×1×1. Then, it passes through two fully connected layers (the last fully connected layer outputs 2 channels), supplemented by SiLU activation, and finally outputs a corner offset vector with a dimension of B×8. After Tanh activation and maximum corner offset constraint (MaxCornerOffset), this vector is transformed into a corner offset of B×4×2 through tensor reshape.

[0048] Meanwhile, the system predefines the standard corner coordinates (-1,-1), (1,-1), (-1,1), and (1,1) as the target point dst. The standard corner coordinates are added point-by-point to the predicted corner offsets to obtain the source point src. A 3×3 homography transformation matrix Hg is then obtained through direct linear transformation (DLT). Finally, this matrix is ​​used to perform differentiable space sampling on the current latent variable, outputting the coarsely aligned current latent variable Zbar_t2^4.

[0049] During model training, since actual on-board observations typically do not provide true geometric transformation parameters, this embodiment does not directly supervise the global transformation parameters.

[0050] Furthermore, step S4 also includes: S44: Construct a mask for the unchanged region using the changed labels.

[0051] S45: In the absence of supervision by real geometric transformation parameters, the unchanged region mask is used to constrain the reference latent variables within the unchanged region to maintain feature consistency with the coarsely aligned current latent variables, and to eliminate the influence of invalid sampling regions.

[0052] In this embodiment, the change label refers to a pixel-by-pixel binary map annotated from the dual-temporal remote sensing image, used to distinguish changed pixels from unchanged pixels. An unchanged region mask is constructed based on this label to indirectly supervise the geometric alignment of the latent space. When performing homography transformation on the current latent variable, some output pixels may map outside the valid range of the original image, forming meaningless invalid regions. Therefore, the system generates a binary valid sampling mask with the same size as the latent variable, marking positions whose mapped coordinates fall within the original image range as valid, and marking the rest as invalid. When calculating the feature consistency loss of the unchanged region, constraints are only applied to pixels that are both within the unchanged region and valid, to avoid noise gradients from invalid regions interfering with the training stability of the global transform predictor.

[0053] S5: Input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, predict the dense residual displacement field level by level according to the current feature and reference feature of each level, and perform local deformation correction on the current feature.

[0054] In this embodiment, the shared alignment decoding module includes a shared decoder and a local fine alignment module, used to regenerate multi-scale task features stepwise from the bottleneck latent variables, and perform progressive local alignment on the current temporal features during each decoding stage. Specifically, the reference compressed latent variables and the coarsely aligned current latent variables are input into the shared decoder to generate reference features and current features at multiple stages; after each decoding stage, the local fine alignment module predicts the dense residual displacement field based on the dual-temporal features, absolute difference features, and element-wise product features, and performs local deformation correction only on the current features.

[0055] It should be noted that a single global transformation is insufficient to fully describe the local nonlinear deformations in on-board unorthorectified images. Therefore, this embodiment embeds a local fine alignment module after each upsampling stage of the shared decoder. Each local fine alignment module receives the newly generated reference features and the current features, constructs a local matching representation, predicts the two-dimensional dense residual displacement field at that scale, and performs local deformation correction only on the current features.

[0056] The shared decoder in this embodiment differs from the traditional U-Net architecture. Since historical shallow features are not stored, this decoder abandons shallow skip connections in historical images and instead recovers multi-scale task features step-by-step based on compressed bottleneck latent variables. This decoder not only performs feature recovery but also collaborates with the local fine alignment module during the step-by-step upsampling process to achieve progressive spatial correction of the current features.

[0057] h_t1^4 = Z~_t1^4, hbar_t2^4 = Zbar_t2^4 u_t1^{l-1} = D_l(h_t1^l), u_t2^{l-1} = D_l(hbar_t2^l), hbar_t2^{l-1} = LFA_l(u_t1^{l-1}, u_t2^{l-1}) Specifically, the historical compressed latent variable Z_t1^4 after entropy decoding and the current latent variable Zbar_t2^4 after coarse alignment are denoted as the decoder inputs h_t1^4 and hbar_t2^4, respectively. In the l-th level decoding stage, the deep features are first mapped to a higher resolution through the upsampling layer D_l to obtain the intermediate features u_t1^{l-1} and u_t2^{l-1}. Subsequently, the local fine alignment module LFA_l uses u_t1^{l-1} as a reference to correct u_t2^{l-1} and outputs the finely aligned current feature hbar_t2^{l-1}. This shared alignment-decoding mechanism effectively avoids the feature extraction redundancy caused by introducing independent registration branches and change detection branches, ensuring end-to-end collaborative optimization of compressed representation learning, feature recovery, spatial alignment, and change detection within a single network pipeline.

[0058] Further, step S5 includes: S51: After upsampling at each level of decoding, a local matching representation is constructed. The local matching representation includes the channel concatenation result of the reference feature at the current scale, the current feature, the absolute difference feature between the two, and the element-wise product feature of the two.

[0059] This step is as follows: m_l = Concat(h_t1^l, h_t2^l, |h_t1^l - h_t2^l|, h_t1^l ⊙ h_t2^l) Where m_l represents the local matching representation, h_t1^l represents the l-th level reference feature, and h_t2^l represents the l-th level current feature.

[0060] S52: Based on the local matching representation, predict the two-dimensional dense residual displacement field at the current scale, while using a scaling factor to limit the displacement amplitude and using the tanh function to constrain the displacement range.

[0061] This step is as follows: Delta^l = s_l * tanh(P_l(m_l)) Where Delta^l represents the dense residual displacement field predicted at level l, s_l represents the scaling factor at level l, which limits the maximum displacement amplitude at the current scale, tanh represents limiting the output to a preset range, such as [-1,1], and P_l represents the lightweight prediction head at level l.

[0062] S53: Based on the differentiable spatial sampling operation, spatial correction is performed only on the current feature according to the two-dimensional dense residual displacement field to complete the fine alignment.

[0063] This step is as follows: hbar_t2^l = W(h_t2^l; Delta^l) Where hbar_t2^l represents the l-th level current feature after fine alignment.

[0064] The local fine alignment module only performs spatial correction on the current temporal features, while the historical reference features are always fixed in the reference coordinate system.

[0065] The aligned current features obtained by the local fine alignment module at each decoding scale are used for feature consistency supervision at that scale to constrain the alignment quality. On the other hand, they are used as input features for the next decoding stage to participate in geometric correction at a higher resolution level. Thus, a latent space geometric alignment process is formed on the basis of global coarse alignment, which is corrected step by step from deep to shallow and from coarse to fine.

[0066] As an optional implementation, step S5 includes: 1. Feature matching construction: Based on the reference features with dimensions B×C×Hl×Wl and the current features, a multimodal matching feature m_l with dimensions B×4C×Hl×Wl is generated. This feature retains the original semantics, difference information and similarity information, providing rich contextual basis for subsequent displacement prediction.

[0067] 2. Migration Prediction Network: The generated matching features m_l are fed into the migration prediction network to regress the dense displacement field. This network consists of a series of convolutional layers and activation functions: a 1×1 convolutional layer (output channels = C channels) connected in sequence, a 3×3 depthwise separable convolutional layer (each channel convolves independently without interference, output channels = C channels), a 1×1 convolutional layer (output channels = C channels), and a 3×3 convolutional layer (output channels = 2). Finally, after passing through the Tanh activation function and max-pooling, the output is a dense displacement field Delta^l (corresponding to the dx, dy offset in the spatial domain) with dimensions B×2×Hl×Wl. The 1×1 convolutional layer and the 3×3 depthwise separable convolutional layer are followed by group normalization (GN) and SILU activation functions.

[0068] 3. Displacement Deformation and Output: After obtaining the dense displacement field Delta^l, the module uses a differentiable deformation operation to perform a sub-pixel-level spatial transformation of the current feature according to the predicted offset. This operation only performs local geometric correction on the current feature, keeping the historical reference features unchanged. The final output is the locally finely aligned current feature (dimensions B×C×Hl×Wl), thus completing the geometric consistency correction at this level.

[0069] Furthermore, the fine alignment process uses an invariant region feature consistency loss, a displacement smoothing regularization term, and a displacement amplitude regularization term.

[0070] It is worth noting that this embodiment uses the consistency loss of unchanged region features, displacement smoothing regularization term and displacement amplitude regularization term to constrain local fine alignment, so that the alignment process is as smooth and reasonable as possible, and avoids forcibly matching the real changed region to the reference feature.

[0071] S6: Input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature.

[0072] In this embodiment, the temporal symmetric fusion module performs summation, absolute difference, element-wise multiplication, and channel mapping on the reference feature and the aligned current feature to obtain a dual-temporal fusion feature that is insensitive to the input order. Specifically, the highest resolution reference feature and the finally aligned current feature are input into the temporal symmetric fusion module to construct a fusion feature that includes common structural information, change response information, and feature-related information.

[0073] Further, step S6 includes: S61: Perform summation, absolute difference, and element-wise multiplication operations on the highest resolution reference feature and the aligned current feature.

[0074] S62: Concatenate the computation results along the channel dimension and perform feature mapping through 1×1 convolution to obtain fused features.

[0075] Steps S61 to S62 are as follows: f_0 = phi(Concat[h_t1^0 + hbar_t2^0, |h_t1^0 - hbar_t2^0|, h_t1^0 ⊙hbar_t2^0]) Where f_0 represents the fused feature, phi() represents the 1×1 convolutional feature map, h_t1^0 represents the highest resolution reference feature shared by the decoder output, and hbar_t2^0 represents the current feature after final alignment. The summation term describes the common structure of the two phases, the absolute difference term highlights the change response, and the element-wise product term characterizes feature consistency and correlation.

[0076] It is worth noting that the core of change detection is to determine whether a change exists between two temporal phases, and the output should not depend on the input order of the images. Therefore, this embodiment employs a temporally symmetric fusion strategy. The final reference features and the aligned current features are combined using summation, absolute difference, and element-wise multiplication to construct a fused representation, which is then mapped using a 1×1 convolution to obtain the detection input. Since these three operations are inherently symmetric, the impact of the input temporal order on the output can be reduced at the structural level.

[0077] S7: Perform change detection based on the fusion features and output a change probability map.

[0078] In this embodiment, the change detection output module is used to output a change probability map, a binary change map, a change region vector, or alarm information based on the fused features. Specifically, the change detection head performs convolution mapping and sigmoid activation on the fused features to obtain a pixel-level change probability map; it can be further thresholded to generate a binary change map or extract the change region; in other embodiments, a change region mask can be further generated based on the threshold as supervision information used in the training phase.

[0079] y_hat = sigmoid(H_CD(f_0)) Where y_hat represents the predicted change probability map obtained by mapping the response value output by the change detection head to the (0,1) interval after the Sigmoid activation function; H_CD represents the change detection head, which extracts change-sensitive features through convolution operation.

[0080] In some embodiments, the optional reconstruction initialization module is used only in the compressed pre-training phase to stabilize the initialization process of the shared encoder, entropy model, and shared decoder. In subsequent end-to-end training and on-board inference phases, this module is removed, does not participate in computation, and is not included in the final deployment output. Specifically, in the on-board inference phase, only the compressed latent variable storage, shared encoder, entropy decoder, global coarse alignment, shared decoder, local fine alignment, temporal symmetric fusion, and change detection head are retained; the reconstruction head is removed, and the final deployed model directly outputs the change detection results.

[0081] Furthermore, phased training strategies include: The compressed pre-training phase activates only the shared encoder, quantization and entropy coding modules, shared decoder, and training reconstruction head, initializing the model parameters. This phase optimizes the bitrate constraint and self-reconstruction loss, ensuring sufficient space and semantic information is retained for the bottleneck latent variables, and initializes the decoder's feature regeneration capability.

[0082] Change detection warm-up phase: Remove the reconstruction head, freeze the shared encoder and entropy model, and train the temporal symmetric fusion module and change detection head on the registered dual-temporal samples to enable the model to learn stable semantic change representations first.

[0083] Alignment warm-up phase: Using synthetic or naturally unregistered samples, activate the global coarse alignment module and the local fine alignment module, freeze the change detection branch, and optimize the alignment module using the consistency loss of unchanged regions.

[0084] End-to-end fine-tuning stage: Jointly optimize all components of the inference stage, including the shared encoder, quantization and entropy coding module, global coarse alignment module, shared decoder, local fine alignment module, temporal symmetric fusion module, and detection head. The total loss includes change detection loss, bit rate constraint, global alignment loss, and local alignment loss, but does not include reconstruction loss.

[0085] in, Total loss; Loss detection for changes; For bitrate constraints; This is a global coarse alignment loss, similar to the following loss for consistency of unchanged region features; The local fine alignment loss includes the consistency loss of features in unchanged regions, the displacement smoothing regularization term, and the displacement amplitude regularization term, which are used to constrain the rationality and stability of the dense residual displacement field predicted by the decoder step by step. , , The weighting coefficients for each loss component are used to balance detection accuracy, compression efficiency, and geometric alignment quality, adapting to on-board storage and computation constraints. This strategy shifts compressed representation from being geared towards vision reconstruction to being geared towards bitrate efficiency, spatial alignment, and change discrimination, making it more suitable for on-board mission-driven compression and change detection.

[0086] It is worth noting that the global coarse alignment loss and the consistency loss of unchanged region features at each level of the decoding end adopt the same supervision paradigm. Both use the unchanged region mask generated by the changed label to constrain the consistency of the bi-temporal features within the region. The difference between the two is that the global coarse alignment loss acts on the compressed bottleneck latent variable space and is combined with the global homography transformation, while the consistency loss at the decoding end acts on the multi-level decoding features and is combined with the dense residual displacement field. The two constitute a coarse-to-fine alignment supervision.

[0087] The bitrate constraint mentioned above is as follows: in, This represents the probability distribution estimated by the entropy model.

[0088] The change detection loss mentioned above is as follows: in, Indicates the change detection loss. This represents the binary cross-entropy loss. This represents the Dice loss, calculated based on the Dice coefficient. It effectively alleviates the class imbalance problem and improves the segmentation accuracy of changing regions (foreground). A graph representing the predicted probability of change. The label represents the actual changes, which are manually labeled as masking.

[0089] The aforementioned local fine alignment loss is calculated as follows: in, This represents the final total loss for local fine alignment. This represents the loss weight of the l-th layer, used to balance the importance of different scales. and These represent hyperparameters that control the intensity of smoothing regularization and displacement amplitude regularization, respectively.

[0090] The consistency loss of the unchanged region features mentioned above is as follows: in, This represents the local alignment loss of the l-th layer. A mask representing the unchanged region. Represents the distance calculation function. Represents the characteristic normalization function, This represents a constant and is used to prevent the denominator from being zero.

[0091] The displacement smoothing regularization term mentioned above is calculated based on the L1 norm, as shown in the following equation: in, Indicates the first l Layer smoothing loss, Indicates the first l Dense residual displacement field predicted by the layer, , These refer to the spatial gradients of the displacement field in the X direction (horizontal) and Y direction (vertical), respectively.

[0092] The displacement amplitude regularization term mentioned above is calculated based on the L2 norm, as shown in the following formula: in, Indicates the first l Displacement amplitude loss of the layer.

[0093] It should be noted that during the training phase, a reconstruction branch can be optionally introduced as auxiliary supervision to improve the feature extraction and regeneration capabilities of the encoder and decoder. The reconstruction loss uses the L1 norm to calculate the error between the reconstructed image and the original image.

[0094] In some embodiments, when deployed on a spaceborne edge computing device, this embodiment can operate according to the following process: After the historical reference image arrives, the shared encoder generates bottleneck latent variables and writes them into the on-board memory; after the current image arrives, the current image is encoded in real time, and the historical latent variables of the corresponding region are read; global coarse alignment is performed in the compressed latent space; multi-scale features are regenerated through the shared decoder, and local fine alignment is completed step by step; after temporal symmetric fusion, the change map is output by the detection head; finally, only the change map, alarm information, or region of interest are downloaded, reducing the bandwidth occupied by the full image download. Since the reconstruction head has been removed during the on-board inference stage, the final deployment model of this invention only outputs the change detection results, without generating additional reconstruction output, and without introducing the parameter, computation, and memory consumption brought by the reconstruction branch.

[0095] In some embodiments, to train and validate the robustness of the model under unregistered satellite observation conditions, random geometric perturbations can be applied to the current temporal images in the registered dataset, while keeping the historical reference images and change labels unchanged in the reference coordinate system. Perturbation types include translation, rotation, scaling, shearing, and perspective corner jitter, and examples of perturbation parameters are shown in Table 1.

[0096] Table 1 Examples of perturbation parameters constructed from unregistered samples

[0097] The perturbation parameters mentioned above are only used to construct training or test samples and are not included in the input during the inference phase. The model does not need to use real geometric transformation parameters, ground control points, or external registration results during the inference phase.

[0098] This embodiment validates the accuracy, robustness, storage efficiency, and embedded operating efficiency on change detection benchmarks such as LEVIR-CD, WHU-CD, and SYSU-CD. Examples of change detection performance under moderate unregistration conditions are shown in Table 2.

[0099] This embodiment maintains good robustness under different misregistration intensities. For example, under heavy perturbation conditions, the F1 score reaches 89.7%, 90.8%, and 83.6% on LEVIR-CD, WHU-CD, and SYSU-CD, respectively, demonstrating the complementary effect of global coarse alignment in the compressed latent space and progressive local fine alignment at the decoding end.

[0100] Table 2. Examples of the change detection performance of the present invention under moderate unregistration conditions.

[0101] As shown in Table 3, in terms of system efficiency, at a bitrate of 3.5 bpp, the historical reference representation is approximately 0.12 MB with a 512×512 input and approximately 0.03 MB with a 256×256 input; the model parameter size is approximately 7.56 M, the FLOPs are approximately 74.50 G / 18.63 G, and the peak video memory is approximately 704 MB / 236 MB. On the Jetson TX2 platform, the end-to-end processing latency for 256×256 dual-temporal images is approximately 118.2 ms, and the throughput is approximately 0.55 Mpixel / s.

[0102] Table 3 Examples of on-board efficiency and embedded deployment effects of the present invention

[0103] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium.

[0104] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the order of execution of the described steps may be changed. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0105] Example 2: This embodiment provides an on-board compression alignment detection system, the system comprising: The first acquisition unit is used to acquire historical reference images and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features. The generation unit is used to input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. The second acquisition unit is used to acquire the currently observed image in order to obtain the current compression bottleneck latent variables; The alignment unit is used to read the historical compression bottleneck latent variable from the on-board memory, predict global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment of the current compression bottleneck latent variable according to the global geometric transformation parameters. The correction unit is used to input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, the dense residual displacement field is predicted level by level according to the current feature and reference feature of each level, and local deformation correction is performed on the current feature. The fusion unit is used to input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; The detection unit is used to perform change detection based on the fused features and output a change probability map.

[0106] Example 3: This embodiment provides a storage medium, which is a computer-readable storage medium, storing a computer program. When the computer program is executed by a processor, it implements the on-board compression alignment detection method of Embodiment 1 above, as follows: S1: Obtain historical reference images, and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features; S2: Input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. S3: Obtain the current observed image to obtain the current compression bottleneck latent variables; S4: Read the historical compression bottleneck latent variable from the on-board memory, predict the global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment on the current compression bottleneck latent variable according to the global geometric transformation parameters; S5: Input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, predict the dense residual displacement field level by level according to the current feature and reference feature of each level, and perform local deformation correction on the current feature. S6: Input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; S7: Perform change detection based on the fusion features and output a change probability map.

[0107] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0108] In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this embodiment, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0109] The computer-readable storage medium described above can be used to write computer programs for executing this embodiment in one or more programming languages ​​or combinations thereof. These programming languages ​​include object-oriented programming languages—such as Java, Python, and C++—and conventional procedural programming languages—such as C or similar programming languages. The program can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0110] In summary, this invention achieves compact preservation of historical information, robust handling of unregistered data, and detection of low-latency changes on satellite by using a shared encoder-decoder architecture, single-bottleneck latent variable storage, global coarse alignment of compressed latent space, progressive local fine alignment at the decoding end, and fusion of temporal symmetric features.

[0111] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A method for on-board compression alignment detection, characterized in that, include: S1: Obtain historical reference images, and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features; S2: Input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. S3: Obtain the current observed image to obtain the current compression bottleneck latent variables; S4: Read the historical compression bottleneck latent variable from the on-board memory, predict the global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment on the current compression bottleneck latent variable according to the global geometric transformation parameters; S5: Input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, predict the dense residual displacement field level by level according to the current feature and reference feature of each level, and perform local deformation correction on the current feature. S6: Input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; S7: Perform change detection based on the fusion features and output a change probability map.

2. The on-board compression alignment detection method according to claim 1, characterized in that, Step S4 includes: S41: Construct a matching-aware representation, wherein the matching-aware representation is the channel concatenation result of the historical compression bottleneck latent variable, the current compression bottleneck latent variable, the absolute difference features of the two, and the element-wise product features of the two. S42: Input the matching perception representation into the global transformation predictor to predict multiple corner offsets, and obtain the homography matrix by solving through direct linear transformation; S43: Based on differentiable space sampling operation, spatial transformation is performed only on the current compression bottleneck latent variable according to the homography matrix to complete coarse alignment.

3. The on-board compression alignment detection method according to claim 2, characterized in that, Step S4 also includes: S44: Construct a mask for the unchanged region using the changed labels; S45: In the absence of supervision by real geometric transformation parameters, the unchanged region mask is used to constrain the reference latent variables within the unchanged region to maintain feature consistency with the coarsely aligned current latent variables, and to eliminate the influence of invalid sampling regions.

4. The on-board compression alignment detection method according to claim 1, characterized in that, Step S5 includes: S51: After upsampling at each level of decoding, a local matching representation is constructed, which includes the reference feature at the current scale, the current feature, the absolute difference feature between the two, and the channel concatenation result of the element-wise product feature of the two. S52: Based on the local matching representation, predict the two-dimensional dense residual displacement field at the current scale, while using a scaling factor to limit the displacement amplitude and using the tanh function to constrain the displacement range; S53: Based on the differentiable spatial sampling operation, spatial correction is performed only on the current feature according to the two-dimensional dense residual displacement field to complete the fine alignment.

5. The on-board compression alignment detection method according to claim 4, characterized in that, The fine alignment process uses the consistency loss of unchanged region features, the displacement smoothing regularization term, and the displacement amplitude regularization term.

6. The on-board compression alignment detection method according to claim 1, characterized in that, Step S6 includes: S61: Perform summation, absolute difference, and element-wise product operations on the highest resolution reference feature and the aligned current feature; S62: Concatenate the computation results along the channel dimension and perform feature mapping through 1×1 convolution to obtain fused features.

7. The on-board compression alignment detection method according to claim 1, characterized in that, The shared decoder does not obtain shallow jump features from historical images. Instead, it regenerates multi-scale task features step by step based solely on the historical compression bottleneck latent variables and the coarsely aligned current compression bottleneck latent variables, and simultaneously completes feature alignment during the regeneration process.

8. The on-board compression alignment detection method according to claim 1, characterized in that, It also includes phased training strategies, including: Compressed pre-training phase: Only the shared encoder, quantization and entropy encoding module, shared decoder, and training reconstruction head are activated to initialize model parameters; Change detection warm-up phase: Remove the reconstruction head, freeze the shared encoder and entropy model, and train the temporal symmetric fusion module and change detection head on the registered dual-temporal samples; Alignment warm-up phase: Activate the global coarse alignment module and the local fine alignment module, freeze the change detection branch, and optimize the alignment module using the consistency loss of the unchanged region; End-to-end fine-tuning stage: Jointly optimize all components of the inference stage, including the shared encoder, quantization and entropy coding module, global coarse alignment module, shared decoder, local fine alignment module, temporal symmetric fusion module, and detection head. The total loss includes change detection loss, bit rate constraint, global alignment loss, and local alignment loss, but does not include reconstruction loss.

9. An on-board compression alignment detection system, characterized in that, include: The first acquisition unit is used to acquire historical reference images and perform multi-level feature encoding on the historical reference images through a shared encoder to obtain multi-scale features. The generation unit is used to input the deepest bottleneck feature in the multi-scale features into the quantization and entropy coding module to generate historical compressed bottleneck latent variables, which are stored in the on-board memory. The second acquisition unit is used to acquire the currently observed image in order to obtain the current compression bottleneck latent variables; The alignment unit is used to read the historical compression bottleneck latent variable from the on-board memory, predict global geometric transformation parameters based on the historical compression bottleneck latent variable and the current compression bottleneck latent variable within the compression latent space, and perform coarse alignment of the current compression bottleneck latent variable according to the global geometric transformation parameters. The correction unit is used to input the coarsely aligned current compression bottleneck latent variable and the historical compression bottleneck latent variable into the shared decoder. During the multi-level upsampling process of the shared decoder, the dense residual displacement field is predicted level by level according to the current feature and reference feature of each level, and local deformation correction is performed on the current feature. The fusion unit is used to input the highest resolution reference feature and the finally aligned current feature into the temporal symmetric fusion module to obtain the fused feature; The detection unit is used to perform change detection based on the fused features and output a change probability map.

10. A readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the on-board compression alignment detection method according to any one of claims 1 to 8.