Reference enhancement compression method for satellite remote sensing images

CN122824898APending Publication Date: 2026-09-25NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611274932.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]然而,现有技术在处理具有复杂频域特性的遥感图像时仍面临挑战

Benefits of technology

[0021]本发明的有益效果是:本发明通过双域检索机制和门控融合策略,能够引入并稳健利用外部相关先验信息,提升极低侧信息开销下的熵参数预测精度,降低遥感图像压缩码率并改善关键地物重建质量。具体效果将在下文结合实施例进行描述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824898A_ABST
    Figure CN122824898A_ABST
Patent Text Reader

Abstract

The application discloses a reference enhancement compression method and system for satellite remote sensing images, comprising: obtaining the current features and potential representation of the image to be compressed, and performing dual-domain similarity evaluation and fusion with the offline cached features in the shared reference library, so as to select a target reference image. The encoding end writes the index of the target reference image into the compression code stream, the decoding end reads the reference image synchronously according to the index, and extracts the reference conditional features aligned with the potential space. The reference conditional features are gate-fused with the basic hyper-prior conditions to generate enhanced hyper-prior conditions, which are used to guide the prediction of the entropy model parameters, and the entropy coding and decoding of the image are completed. Through the dual-domain retrieval mechanism and the gate-fusion strategy, the application can introduce and robustly use external related prior information, improve the prediction accuracy of the entropy parameters under the condition of extremely low side information overhead, reduce the compression code rate of the remote sensing image, and improve the reconstruction quality of key ground objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image compression technology, and in particular to a reference enhancement compression method for satellite remote sensing images. Background Technology

[0002] With the rapid development of satellite Earth observation technology, the data volume of high-resolution remote sensing images is growing exponentially. These images are typically characterized by their huge area, complex and repetitive ground textures, and continuous acquisition across multiple time phases. Given the limitations of limited satellite-to-ground transmission bandwidth and cloud storage resources, efficiently compressing massive amounts of satellite remote sensing images while preserving key ground features such as road and building edges at extremely low bit rates is a crucial technical aspect of the remote sensing data link.

[0003] Existing image compression methods are mainly divided into traditional coding standards that rely on manually designed rules (such as JPEG2000 and VVC) and learning-based image compression methods based on deep neural networks. In recent years, learning-based image compression methods have achieved excellent performance in natural image compression by automatically learning the latent representation and probability distribution of images. Most current learning-based methods adopt a single-image compression architecture, that is, the model estimates the latent distribution only by relying on the internal pixel redundancy of the current image to be compressed. In order to further improve the compression ratio, some studies have attempted to introduce external reference images, match similar scenes through semantic feature networks, and perform weighted fusion of the feature tensors of the reference images with the features of the current image in the spatial or channel dimensions, in order to provide additional probabilistic priors for the entropy model.

[0004] However, existing technologies still face challenges when processing remote sensing images with complex frequency domain characteristics. On the one hand, relying solely on macroscopic semantic matching is insufficient to guarantee that reference information has actual entropy reduction value in the microscopic compressed feature space, easily introducing invalid references. On the other hand, remote sensing images are greatly affected by imaging conditions, and blind feature fusion can easily lead to high-frequency noise or local structural differences directly interfering with the parameter prediction of the entropy model, resulting in inaccurate probability estimation.

[0005] Therefore, there is an urgent need for a remote sensing image compression scheme that can more accurately assess reference value and achieve robust utilization of prior information. Summary of the Invention

[0006] Purpose of the invention: To provide a reference enhancement compression method for satellite remote sensing images, in order to solve the above-mentioned problems existing in the prior art.

[0007] Technical solution: To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A reference enhancement compression method for satellite remote sensing images includes: Acquire the remote sensing image to be compressed, and extract the current image features, latent representation, and basic advanced prior conditions of the remote sensing image to be compressed; A dual-domain similarity evaluation is performed based on the current image features and the offline cache features of candidate images in the pre-built shared reference image library, and the evaluation results are fused to select the target reference image. Write the index of the target reference image into the compressed bitstream; parse the compressed bitstream at the decoding end to obtain the index, and read the corresponding target reference image from the shared reference image library according to the index; Extract reference condition features from the target reference image, wherein the reference condition features are aligned with the compressed latent space in which the latent representation resides; The reference condition features are gated and fused with the basic super-prior conditions to generate enhanced super-prior conditions; the entropy model parameters are predicted based on the enhanced super-prior conditions. The latent representation is entropy encoded and decoded using the entropy model parameters to obtain the reconstructed remote sensing image.

[0008] In one embodiment, the current image features include current semantic features and current latent spatial features, and the offline cached features include cached semantic features and cached latent spatial features; The dual-domain similarity evaluation based on the current image features and the offline cached features of candidate images in a pre-built shared reference library includes: A semantic similarity assessment is performed based on the current semantic features and the cached semantic features to obtain the semantic similarity. Based on the current latent space features and the cached latent space features, a compressed latent space similarity evaluation is performed to obtain the compressed latent space similarity.

[0009] In one embodiment, a compressed latent space similarity evaluation is performed based on the current latent space features and the cached latent space features to obtain the compressed latent space similarity, including: Extract the channel variance of the current latent spatial features and the cached latent spatial features; Calculate the channel-level statistical correlation between the current potential space features and the cached potential space features, and estimate the channel correlation coefficient based on the channel-level statistical correlation; Based on the channel correlation coefficient, the proportion of residual conditional variance after introducing reference conditions is determined, and the proportion of residual conditional variance is converted into channel potential bitrate savings. Using the variance of each channel in the current latent space as weight, the potential bitrate savings corresponding to each channel are weighted, summed, and normalized to obtain the compression latent space similarity.

[0010] In one embodiment, the step of fusing the evaluation results and selecting the target reference image specifically includes: Extract the global pooling vector of the current latent spatial features, and concatenate the global pooling vector with the current semantic features to obtain an image content summary; The image content summary is input into a pre-trained weight prediction network, which outputs adaptive dual-domain fusion weights. The semantic similarity and the compressed latent space similarity are weighted and fused using the adaptive dual-domain fusion weights to obtain a comprehensive evaluation score; The target reference image is selected by sorting the candidate images in the shared reference image library according to their comprehensive evaluation scores.

[0011] In one embodiment, extracting the reference condition features of the target reference image includes: The target reference image is input into a feature extraction layer that shares weights with the backbone network to extract intermediate reference features. The intermediate reference features are input into a pre-trained adaptation layer for channel dimensionality reduction and local feature mapping, and the reference condition features aligned with the compressed latent space are output.

[0012] In one embodiment, the gating fusion of the reference condition features and the basic prior conditions includes: Extract the global spatial features of the basic prior conditions and the reference conditions to generate channel-level gating; By concatenating the basic prior conditions and the reference condition features, a spatial-level gating system is generated. The reference condition features are jointly modulated based on the channel-level gating and the space-level gating.

[0013] In one embodiment, after jointly modulating the reference condition features, a frequency domain segmented modulation step is further included: Frequency domain transformation is performed on the basic prior conditions and the modulated reference condition features respectively, dividing them into multiple frequency band features; Combined with a preset frequency band confidence modulation factor, frequency band gating weights corresponding to each frequency band feature are generated independently; The reference condition features of the corresponding frequency band are frequency-domain modulated using the frequency band gate weights and fused with the basic prior conditions of the corresponding frequency band to obtain frequency band fusion features; Perform an inverse frequency domain transformation on the frequency band fusion features to map them back to the spatial domain, thereby generating the enhanced prior conditions.

[0014] In one embodiment, when multiple target reference images are selected, the gated fusion further includes a multi-reference complementary fusion step: Calculate the correlation matrix among the reference condition features corresponding to each of the target reference images; Based on the correlation matrix, the uniqueness score corresponding to each target reference image is determined, and the uniqueness score is normalized to obtain the complementarity weight. The modulation intensities of the channel-level gating and the spatial-level gating corresponding to each of the target reference images are weighted and adjusted using the complementary weights to generate the enhanced prior conditions.

[0015] In one embodiment, the parameters of the enhanced prior condition prediction entropy model include: The enhanced prior conditions are input into the pre-trained basic prediction network to obtain the basic probability parameters; The enhanced prior conditions are input into a pre-trained entropy parameter adaptation layer to obtain a domain correction term for the characteristics of satellite remote sensing images. The basic probability parameters are superimposed and corrected with the domain correction term to obtain the entropy model parameters.

[0016] In one embodiment, the network modules used in the method are obtained through pre-training via the following steps: Pre-training was performed using a general image dataset to obtain a backbone network with basic feature extraction and image reconstruction capabilities; Domain adaptation fine-tuning is performed using a remote sensing image dataset. While keeping the backbone network parameters frozen, an adaptation layer for extracting the reference condition features, a gated fusion network for generating the enhanced prior conditions, and an entropy parameter adaptation layer for predicting the entropy model parameters are jointly trained.

[0017] In one embodiment, the shared reference library and the offline cache features are pre-built through the following steps: Multiple remote sensing reference images are acquired to construct the shared reference image library, and a unique index is assigned to each of the remote sensing reference images; The global semantic features of each remote sensing reference image are extracted using a pre-trained semantic feature extraction network and used as the cached semantic features for offline caching. The compressed latent space output of each remote sensing reference image is extracted using the intermediate layer of the backbone network and used as the cached latent space features for offline caching. The channel mean, channel variance, and global pooling vector of the cached latent space features are calculated and cached offline.

[0018] In one embodiment, writing the index of the target reference image into the compressed bitstream includes: A fixed encoding length is determined based on the total number of candidate images in the shared reference image library and the preset number of target reference images K; wherein, the number of target reference images K is shared in advance by the encoding end and the decoding end. The index is encoded using the fixed encoding length, and the encoded index is written into the header of the compressed bitstream to achieve side information synchronization at the decoding end.

[0019] In one embodiment, the adaptation layer employs a bottleneck structure, and the step of inputting the intermediate reference features into the pre-trained adaptation layer for channel dimensionality reduction and local feature mapping includes: The intermediate reference features are compressed through a channel reduction layer; A local feature adaptation layer is used to perform remote sensing domain local texture mapping on the compressed features; The channel dimension is increased by using a channel recovery layer to output the reference condition features.

[0020] In one embodiment, prior to acquiring the remote sensing image to be compressed, the method further includes: Based on geographical distribution or land cover category attributes, the shared reference map library is divided into multiple sub-map libraries; After acquiring the remote sensing image to be compressed, extract the geographic location information or land feature attribute labels from the remote sensing image to be compressed; The corresponding sub-map library is activated based on the geographic location information or the feature attribute label, so as to perform the dual-domain similarity assessment within the activated sub-map library.

[0021] The beneficial effects of this invention are as follows: By employing a dual-domain retrieval mechanism and a gated fusion strategy, this invention can introduce and robustly utilize relevant prior information from external sources, thereby improving the prediction accuracy of entropy parameters with extremely low side information overhead, reducing the compression bitrate of remote sensing images, and improving the reconstruction quality of key ground features. Specific effects will be described below in conjunction with embodiments. Attached Figure Description

[0022] Figure 1 This is a flowchart of the present invention.

[0023] Figure 2 This is a flowchart of the dual-domain similarity evaluation method of the present invention.

[0024] Figure 3 This is a flowchart of the present invention for calculating the similarity of compressed latent spaces.

[0025] Figure 4 This is a flowchart of the present invention for fusing evaluation results and selecting a target reference image.

[0026] Figure 5This is a flowchart of the present invention for extracting reference condition features of the target reference image. Detailed Implementation

[0027] Combination Figure 1 Description of Example 1: This example provides a reference enhancement compression method for satellite remote sensing images, mainly including the following steps: Step 101: Obtain the remote sensing image to be compressed, and extract the current image features, latent representation, and basic prior conditions of the remote sensing image to be compressed.

[0028] After receiving satellite remote sensing images, they are preprocessed, such as normalized, cropped, or padded, to meet the input size requirements of the compression network, resulting in preprocessed images.

[0029] The preprocessed image is input into the basic learning-based compression model, and the latent representation is extracted by the analysis transformation network. At the same time, the super-prior analysis network is used to generate super-prior latent variables.

[0030] After quantizing and encoding the variable, the decoder can obtain the corresponding quantized prior variable, which is then processed by the prior synthesis network to output the basic prior conditions.

[0031] During this process, current image features are also extracted simultaneously for subsequent retrieval, providing a data foundation for introducing external reference information.

[0032] Step 102: Perform dual-domain similarity evaluation based on the current image features and the offline cache features of candidate images in the pre-built shared reference image library, and fuse the evaluation results to select the target reference image.

[0033] Studies have found that satellite remote sensing images commonly exhibit external correlations, such as the recurrence of similar ground features and continuous acquisition across multiple time phases. To fully utilize these naturally occurring external references, this step evaluates the suitability of external candidate images for compression from two dimensions: the semantic domain and the compression potential spatial domain.

[0034] Specifically, the extracted current image features are compared with features from a shared reference image library that has been pre-built offline and cached.

[0035] The first dimension is semantic similarity, which measures the degree of correlation between the two in terms of macro-level content such as land cover categories and scene layout, to avoid introducing completely irrelevant interfering images.

[0036] The second dimension is compressed latent space similarity, which can deeply compress the feature space inside the model and is used to predict the reference image to reduce the uncertainty of the latent representation of the current image.

[0037] The evaluation results from the two dimensions are combined and weighted to calculate a comprehensive retrieval score. The target reference image that is most helpful to the current compression task is then selected based on the score ranking.

[0038] It should be noted that semantic similarity can be achieved using methods such as cosine similarity, and normalization methods can also be adopted using existing methods, which will not be elaborated on in this article.

[0039] Step 103: Write the index of the target reference image into the compressed bitstream.

[0040] To ensure consistent use of the external reference image at both the encoding and decoding ends, after completing the above dual-domain retrieval, the encoding end directly writes the identifier index of the selected target reference image into the header of the final output compressed bitstream.

[0041] Since only fixed-length index identifiers are transmitted instead of the pixel data of the reference image itself, this side information transmission mechanism generates very low additional bit rate overhead, making it suitable for compression scenarios of large-format satellite remote sensing images.

[0042] Step 104: At the decoding end, the compressed bitstream is parsed to obtain the reference image identifier index, and the corresponding target reference image is read from the shared reference image library according to the reference image identifier index.

[0043] During the decoding phase, the decoding device first parses the header information of the received compressed bitstream and extracts the reference image identifier index carried within. Since both the encoding and decoding ends deploy a consistent shared reference image library, the decoding end only needs to retrieve the same high-quality reference image directly from the local image library based on this index, avoiding the need to re-perform complex reference retrieval calculations based on the original image and ensuring strict synchronization of the encoding and decoding operations.

[0044] Step 105: Extract the reference condition features of the target reference image, which are aligned with the compressed latent space in which the latent representation is located.

[0045] After obtaining the target reference image, a reference coding network corresponding to the transformation process of the master compression model is used to extract its features. The reference conditional features output by this extraction process are mapped through an adaptation layer to maintain consistency with the master latent representation of the image to be compressed in terms of channel dimension and spatial scale.

[0046] Through spatial alignment mechanisms, external reference information can be seamlessly injected into the original compressed feature space, avoiding damage to the general rate-distortion performance of the basic compressed model due to feature dimension mismatch.

[0047] Step 106: Gated fusion of the reference condition features and the basic prior conditions is performed to generate enhanced prior conditions.

[0048] Considering that there may be slight differences in imaging time, sensor conditions, or local texture details between the external reference image and the image to be compressed, unfiltered reference information may interfere with subsequent parameter predictions.

[0049] Therefore, this step introduces an adaptive gating mechanism to selectively inject reference information. By evaluating the relationship between the fundamental prior conditions and the characteristics of the reference conditions, the utilization intensity of the reference information in different channels, spatial locations, and even frequency components is dynamically controlled.

[0050] The reference condition features, filtered and modulated by the gating mechanism, are superimposed on the basic prior conditions to form enhanced prior conditions containing high-quality external prior information.

[0051] Step 107: Based on the enhanced prior conditions, predict the entropy model parameters.

[0052] The enhanced prior conditions generated above are input into the basic entropy parameter prediction network and the related domain adaptation module.

[0053] Guided by enhanced prior conditions, the network can more accurately capture the probability distribution characteristics of the latent representation of remote sensing images, and thus output the corrected entropy model parameters.

[0054] Entropy model parameters typically include data such as mean and scale parameters used to describe the characteristics of Gaussian distribution. The improvement in its prediction accuracy determines the bit rate space that can be saved in the subsequent entropy coding process.

[0055] Step 108: Use the entropy model parameters to entropy encode and decode the latent representation to obtain the reconstructed latent representation, and process the reconstructed latent representation through a synthetic transform network to obtain the reconstructed remote sensing image.

[0056] The encoding end performs entropy encoding on the quantized latent representation based on the accurate mean and scale parameters obtained from the prediction, generating the final compressed bitstream body.

[0057] Correspondingly, the decoding end uses the exact same enhanced prior conditions and predicted entropy model parameters to entropy decode the received bitstream, recover the quantized latent representation, and input it into the synthetic transformation network to finally reconstruct a satellite remote sensing image with higher visual quality and more complete preservation of key ground feature details.

[0058] To address the issue that existing methods are prone to introducing invalid reference images, this solution establishes a predictive latent spatial gain metric based on latent spatial statistical features and combines it with image content summarization to dynamically predict dual-domain fusion weights. This transforms the retrieval target from a single visual similarity to a compression gain-prioritized approach, eliminating interfering images that are similar in macroscopic scenes but have large differences in statistical distribution, thereby improving the matching accuracy of reference information.

[0059] Example 2: Further explanation of the offline construction of the shared reference library and the dynamic activation mechanism of the sub-libraries.

[0060] In one possible implementation, the shared reference library and the offline cache features are pre-built through the following steps: Step 201: Obtain multiple remote sensing reference images to construct the shared reference image library, and assign a unique index to each of the remote sensing reference images.

[0061] During the offline preparation phase, the system collects historical satellite imagery data covering different times, regions, and various land cover types to form a basic dataset. After uniformly cropping and normalizing the pixel values ​​of the image data in the basic dataset, it is stored in the database.

[0062] To facilitate lightweight side information transmission and fast addressing during online compression, each image in the database is assigned a fixed numerical number as a unique index, thereby establishing a complete shared reference image library.

[0063] Step 202: Use a pre-trained semantic feature extraction network to extract global semantic features of each remote sensing reference image, which are then used as the cached semantic features for offline caching.

[0064] For each reference image in the image library, it is fed into a pre-trained semantic feature extraction model. This feature extraction model can be implemented using a residual network, visual Transformer, or other semantic encoding network pre-trained or fine-tuned on a remote sensing image dataset, and the classification head structure can be removed.

[0065] After multi-layer feature extraction, the input image is compressed into a one-dimensional feature vector using global average pooling, such as a 2048-dimensional or other preset-dimensional floating-point vector. This vector can represent the abstract information of the image in terms of land cover categories and macro-scene layout. The system persistently stores this as a cached semantic feature on the local medium.

[0066] Residual networks, convolution, and pooling operations can be implemented using existing technologies.

[0067] Step 203: Extract the compressed latent space output of each remote sensing reference image using the intermediate layer of the backbone network as the cached latent space features for offline caching, and calculate and cache the channel mean, channel variance, and global pooling vector of the cached latent space features offline.

[0068] Simultaneously with semantic feature extraction, the reference image is input into the analysis and transformation network of the basic learning-based compression model. The output tensor of this analysis and transformation network when it propagates forward to a specific intermediate layer is extracted; this output tensor serves as the cached latent space feature. To support the conditional entropy reduction prediction calculation in the subsequent online retrieval stage, the system further performs statistical operations based on this tensor.

[0069] Specifically, for each channel of the feature tensor, the channel mean μc and channel variance σ are calculated in the spatial dimension. 2 c The channel standard deviation σ is obtained from the channel variance. c This generates the channel mean and channel variance. Simultaneously, global average pooling and global standard deviation pooling are performed on the entire feature tensor, and the two are concatenated into a compact global pooling vector.

[0070] It should be noted that the offline calculation process described above only needs to be executed up to the intermediate layer of the analysis transform network; it does not require completing the full encoding and decoding process. The processing time for a single image is typically in the millisecond range. Furthermore, the mean, variance, and correlation coefficient (discussed later) are all basic statistical formulas.

[0071] The cached data consists mainly of low-dimensional statistical vectors rather than the original high-dimensional feature maps. For example, for the 192 channels of a specific level output, the cached data per image only requires a few hundred bytes, making it possible to build a reference library containing tens of thousands of images on conventional storage devices.

[0072] Step 204: Divide the shared reference map library into multiple sub-map libraries according to the geographical distribution or land feature category attributes.

[0073] To improve the matching efficiency of online retrieval, the system reads the metadata attached to the reference image.

[0074] Based on the latitude and longitude coordinate range recorded in the metadata, images covering similar geographical areas are grouped into the same spatial sub-library. Alternatively, based on scene classification labels, images containing similar terrain features can be grouped into the same attribute sub-library.

[0075] This logical partitioning management allows the vast global image library to be divided into multiple independent or partially overlapping retrieval domains.

[0076] Step 205: After acquiring the remote sensing image to be compressed, extract the geographic location information or land feature attribute tags of the remote sensing image to be compressed.

[0077] When entering the real-time online compression process, the system parses the file header or associated auxiliary description file of the remote sensing image to be compressed, and attempts to read the spatial positioning data or pre-labeled scene type data at the time of acquisition.

[0078] Step 206: Activate the corresponding sub-map library based on the geographic location information or the land feature attribute label, so as to perform the dual-domain similarity assessment within the activated sub-map library.

[0079] In this step, the system performs a conditional decision branch based on the metadata extraction results from the previous step to determine the final search scope.

[0080] In the first scenario, when the system successfully extracts valid geographic location information or feature attribute tags, it uses this information as a query keyword to match and activate the corresponding sub-library in the map library management module.

[0081] Subsequent semantic feature matching and latent spatial feature matching operations will be strictly limited to the activated sub-graph library.

[0082] Local retrieval strategies can filter out interfering images that are physically far away or irrelevant in the scene, reducing the number of operations for online similarity calculation and lowering the overall processing latency of the system.

[0083] In the second scenario, when the image to be compressed lacks metadata support, making it impossible to extract valid geographic location information or feature attribute labels, the system maintains the default global search mode. In this case, the dual-domain similarity evaluation module traverses all candidate images in the entire shared reference image library, scoring and ranking them based on the image's content features to ensure that usable reference images can still be found even in the absence of prior information.

[0084] Example 3: Further explanation of the specific process of dual-domain similarity assessment and fusion.

[0085] In one possible implementation, the current image features include current semantic features and current latent spatial features, and the offline cache features include cached semantic features and cached latent spatial features; like Figure 2 and Figure 3 As shown, the dual-domain similarity evaluation based on the current image features and the offline cached features of candidate images in the pre-built shared reference image library includes the following steps: Step 301: Perform semantic similarity evaluation based on the current semantic features and the cached semantic features to obtain semantic similarity.

[0086] After extracting the current semantic features of the image to be compressed, the operation module compares them with the cached semantic features of each candidate image in the shared reference library.

[0087] In practice, the cosine similarity algorithm is used to calculate the directional consistency between two high-dimensional semantic feature vectors. To ensure a unified weighted calculation with the score of the latent spatial dimension, the calculated original cosine values ​​are linearly translated and scaled, and normalized to a value range of 0 to 1 to generate the final semantic similarity score. This score reflects the degree of matching between the images and macroscopic land cover categories.

[0088] Step 302: Extract the channel variance of the current latent spatial features and the cached latent spatial features.

[0089] Read the current latent spatial feature tensor of the image to be compressed and the cached latent spatial feature tensor of the candidate image. For each independent channel in the feature tensor, read or calculate its distribution variance in the spatial dimension. The channel variance data is used to characterize the dispersion of each channel and its relative coding contribution, and is a basic statistic for subsequent evaluation of the compression gain of the reference image.

[0090] Step 303: Based on the channel statistics of the current potential spatial features and the cached potential spatial features, calculate the channel-level statistical distance and convert the channel-level statistical distance into an equivalent channel correlation coefficient.

[0091] Calculate the channel-level statistical correlation between the current potential space features and the cached potential space features, and estimate the channel correlation coefficient based on the channel-level statistical correlation; To measure the effectiveness of the reference image in the compressed latent space, the system evaluates the statistical matching degree between the two at the channel level. For the c-th feature channel, based on the channel mean μX,c and channel standard deviation σX,c of the current latent space features, and the cached channel mean μR,c and cached channel standard deviation σR,c of the candidate image, the channel-level statistical distance dc is calculated; then, the channel-level statistical distance dc is mapped to an equivalent channel correlation coefficient ρc with a value between 0 and 1. This equivalent channel correlation coefficient is used for subsequent retrieval scoring and does not represent a direct calculation of the true joint statistics from the marginal statistics.

[0092] Step 304: Determine the remaining conditional variance ratio after introducing reference conditions based on the channel correlation coefficient, and convert the remaining conditional variance ratio into channel potential bitrate savings.

[0093] After obtaining the equivalent channel correlation coefficients, the system constructs an estimate of the compression gains that the reference information may bring based on the Gaussian source approximation. For the c-th feature channel, if its equivalent channel correlation coefficient is ρc, then the proportion of the residual conditional variance after introducing the reference condition can be approximately expressed as 1-ρc². The smaller this proportion, the stronger the effect of the reference feature on reducing the uncertainty of the current potential representation.

[0094] Subsequently, the system uses the logarithmic relationship between source entropy and variance to map the residual conditional variance ratio into a specific bit-saving estimate.

[0095] Specifically, the estimated potential rate savings for the c-th feature channel is ΔRc = -0.5 × log2(max(1-ρc²,εlog)); Where ΔRc is the estimated potential rate savings of the c-th feature channel, log2 represents the logarithmic operation to the base 2, ρc is the equivalent channel correlation coefficient of the c-th feature channel, and εlog is a positive stability constant used to prevent zero overflow of the logarithmic operation.

[0096] Step 305: Using the variance of each channel of the current latent space feature as weight, perform a weighted summation and normalization of the potential bitrate savings corresponding to each channel to obtain the compression latent space similarity.

[0097] Since feature channels with larger variances typically contain more information to be encoded, their bitrate savings are more valuable. Therefore, a weighted summation operation is performed on the estimated potential bitrate savings ΔRc for all channels, using the variance of each channel of the current latent space features as the weight coefficient wc. After summation, the combined value is mapped to the interval between 0 and 1 using an extremum normalization method (min-max method) within the currently activated candidate set, generating a compressed latent space similarity with uniform dimensions. This similarity is used to estimate the potential of the reference image to reduce the uncertainty of the current image's latent representation.

[0098] Specifically, the weight coefficient wc = σ²X,c / (Σcσ²X,c + εw); where wc is the weight coefficient of the c-th channel, X represents the current image to be compressed, c represents the channel index, σ²X,c is the latent spatial variance of the current image to be compressed X in the c-th channel, and εw is a positive stability constant used for weight normalization.

[0099] like Figure 4 As shown, according to one aspect of this application, the step of fusing the evaluation results and selecting the target reference image specifically includes: Step 306: Extract the global pooling vector of the current latent spatial features, and concatenate the global pooling vector with the current semantic features to obtain an image content summary.

[0100] To enable the dual-domain fusion process to adapt to the content characteristics of different images, the system constructs summary information to describe the overall attributes of the image. The low-dimensional vector `vlatent(X)`, obtained by extracting the current latent spatial features and processing them through global average pooling and global standard deviation pooling, is concatenated with the current semantic feature vector along the channel dimension to generate an image content summary vector `fsummary` containing semantic and underlying statistical information. This image content summary reflects the semantic content and latent spatial statistical complexity of the current image.

[0101] Step 307: Input the image content summary into the pre-trained weight prediction network and output adaptive dual-domain fusion weights.

[0102] The image content summary vector is input into a lightweight multilayer perceptron network. This network analyzes the texture complexity and semantic richness of the image through two fully connected layers and nonlinear activation operations, and dynamically outputs a weight between 0 and 1.

[0103] α=sigmoid(W2·ReLU(W1·fsummary+b1)+b2); Where α is the adaptive dual-domain fusion weight, sigmoid is the activation function that maps the output to the interval between 0 and 1, ReLU is the linear rectified function, W1 and W2 are the weight parameters of the first and second layers, respectively, fsummary is the image content summary vector, and b1 and b2 are the bias terms of the first and second layers, respectively.

[0104] As an alternative, when computational resources are limited or when processing batches of remote sensing images with high homogeneity, the system can bypass the aforementioned weight prediction network and directly preset a fixed hyperparameter as the dual-domain fusion weight. For example, the weight can be set to a constant 0.4, giving a fixed proportion of emphasis to the evaluation results of the latent spatial domain, thereby reducing the complexity of online computation.

[0105] In another embodiment, the above weight prediction network can also be summarized as α=sigmoid(MLP(fsummary(X))).

[0106] Step 308: The semantic similarity and the compressed latent space similarity are weighted and fused using the adaptive dual-domain fusion weights to obtain a comprehensive evaluation score.

[0107] After obtaining the fusion weights, the multiplier will combine the semantic similarity Ss em Multiplying by the fusion weight α, this compresses the latent space similarity S. latent Multiply this by the difference between 1 and the fusion weight. Then add the two products together to output the final comprehensive evaluation score S. fuse That is, Sfuse =α×Ss em +(1-α)S latent .

[0108] This score takes into account both the macroscopic consistency of image content and the potential for microscopic compression gain.

[0109] Step 309: Sort the candidate images in the shared reference image library according to their comprehensive evaluation scores, and select the target reference image.

[0110] After evaluating all candidate images in the image library, the sorting module arranges them in descending order based on their comprehensive evaluation scores. A predetermined number of image sequences are then selected from the top of the sequence, forming the Top-K reference images, which are used as the target reference image set to assist in the compression of the current image.

[0111] Example 4: Further explanation of the implementation process of extremely low overhead bitstream synchronization and reference feature space alignment.

[0112] In one possible implementation, writing the index of the target reference image into the compressed bitstream includes the following steps: Step 401: Determine the fixed encoding length based on the total number of candidate images in the shared reference image library and the preset number of target reference images K.

[0113] After determining the target reference image, the system needs to establish a mechanism to transmit the selected image information to the decoding end. The number of target reference images, K, can be used as a system parameter shared in advance by both the encoding and decoding ends; when K needs to adapt to changes in the image, the encoding end first writes K or the length of the index field in the bitstream header, and then writes the corresponding reference image index.

[0114] To avoid transmitting massive amounts of image pixel data, the computation module uses index-side information. The system reads the total number of candidate images registered in the image library management module and, combined with the number of target reference images actually selected in the current retrieval phase, calculates the number of binary bits required to transmit these indices. This bit calculation aims to achieve unambiguous address mapping with minimal data payload.

[0115] For a reference image library of size N and a preset selection of K reference images, the binary encoding length Bindex is: Bindex = K × ceil(log2(N)); Where K is the number of target reference images selected, ceil is the rounding up function, and N is the total number of candidate images in the shared reference library.

[0116] Step 402: Encode the index using the fixed encoding length and write the encoded index into the header of the compressed bitstream to achieve side information synchronization at the decoding end.

[0117] After obtaining the fixed-length binary string, the bitstream multiplexer converts the digital index of the target reference image into a binary code of the corresponding length.

[0118] The encoded data segment is inserted at the beginning of the output compressed bitstream. The pre-write strategy enables the decoding device to prioritize parsing the reference index when receiving the data stream, and then synchronously extract the same reference image from the local shared image library.

[0119] Since the amount of index data is usually in the tens of bits range, the additional transmission overhead it brings is relatively low compared to the tens of thousands of bits of the main image bitstream.

[0120] Furthermore, the input required for gated fusion is determined by the basic prior conditions, reference condition features, and target reference image index available at the decoding end, without the need for additional transmission of semantic similarity and compressed latent space similarity.

[0121] Therefore, the header of the compressed bitstream only needs to carry necessary synchronization information such as the reference image index, without adding extra overhead for similarity-side information due to gating fusion.

[0122] After the decoding end parses the reference image index and reads the corresponding reference image, it generates gating weights through a reference feature extraction network and a gating fusion network consistent with the encoding end, so that the gating operations at both the encoding and decoding ends are consistent.

[0123] Combination Figure 5 The process is described below.

[0124] Step 403: Input the target reference image into the feature extraction layer that shares weights with the backbone network to extract intermediate reference features.

[0125] To avoid introducing an independent reference network that would lead to inconsistent feature space distribution, the feature extraction module reuses the front-end network structure of the basic learning-based compressed model.

[0126] Specifically, the target reference image is input into the first few layers of the backbone analysis and transformation network. During the model training and inference phases, the network weights of these feature extraction layers are set to a frozen state, maintaining absolute consistency with the backbone network processing the image to be compressed. After forward computation through this part of the network with shared weights, intermediate reference features with basic physical semantics are output.

[0127] Step 404: Input the intermediate reference features into the pre-trained adaptation layer for channel dimensionality reduction, local feature mapping and channel recovery, and map them to the same channel dimension and spatial scale as the basic super-prior conditions through 1×1 convolution or super-prior fusion adaptation layer, and output the reference condition features.

[0128] While intermediate reference features retain the underlying structure of the general image, they still require domain transformation to account for the specific properties of remote sensing images. The processing unit feeds these intermediate reference features into an adaptation layer with a bottleneck structure. This adaptation layer, through a series of linear transformations and nonlinear activation operations, injects prior information from the remote sensing domain without altering the original feature space distribution, ultimately outputting reference condition features with the same spatial dimension as the main encoded data stream.

[0129] The reference condition features output by the adaptation layer can be expressed as: Fadapter(x)=x+Conv1×1up(φ(Conv3×3(φ(Conv1×1down(x))))); Where x is the intermediate reference feature of the input, Conv1×1down is the channel dimensionality reduction layer, Conv3×3 is the local feature adaptation layer, Conv1×1up is the channel recovery layer, φ is the nonlinear activation function, and the plus sign indicates the residual connection operation.

[0130] Step 405: Channel compression is performed on the intermediate reference features through a channel dimensionality reduction layer to obtain channel compressed features.

[0131] In the first stage of executing the aforementioned bottleneck structure, the convolutional unit uses a 1×1 kernel-sized 2D convolution to perform a linear projection across channels onto the input intermediate reference features. This operation aims to reduce the channel depth of the feature map, for example, compressing the original 192 feature channels to 48 channels. This not only reduces the computational complexity of subsequent local feature mapping but also forces the network to extract the most representative reference information.

[0132] Step 406: Use the local feature adaptation layer to perform remote sensing domain local texture mapping on the compressed channel features to obtain local mapped features.

[0133] After channel compression, the spatial processing module performs a 3x3 two-dimensional convolution operation on the low-dimensional feature map. This convolution operation within the local receptive field is specifically designed to capture the unique fine line structures, small object edges, and repetitive texture patterns found in satellite remote sensing images. Through this mapping layer, the reference features more closely match the local statistical characteristics of the remote sensing image to be compressed in terms of spatial topology.

[0134] Step 407: Upgrade the channel dimension of the local mapping features through the channel recovery layer to output the reference condition features.

[0135] In the final stage of the bottleneck structure, the dimension restoration module again uses a 1×1 2D convolution to expand the locally mapped features through channels. The number of compressed channels is restored to the same dimension as the intermediate input reference features, for example, upscaling it back to 192 channels. The upscaled features are then added element-wise with the original input features, thus retaining the original general feature base while overlaying adaptation information from the remote sensing domain. The output is the final aligned reference conditional feature, providing standardized data input for subsequent gated fusion.

[0136] Furthermore, in the face of limited transmission bandwidth and model generalization bottlenecks, this solution writes fixed-length index identifiers into the bitstream header and combines a two-stage lightweight fine-tuning strategy of freezing backbone features. This ensures strict synchronization between the encoding and decoding ends with only extremely low side information overhead, and enables the model to smoothly adapt to satellite overhead view and fine line structure characteristics, thereby improving the overall rate-distortion performance of the system.

[0137] Example 5: Further explanation of the specific implementation process of the multidimensional adaptive gating fusion mechanism.

[0138] In one possible implementation, the gating fusion of the reference condition features and the basic prior conditions includes the following steps: Step 501: Extract the global spatial features of the basic prior conditions and the reference conditions to generate channel-level gating.

[0139] In order to control the injection intensity of reference information at the channel dimension, the computation module extracts the global distribution information of basic prior conditions and reference condition features at the spatial level.

[0140] Specifically, global average pooling is performed on the two feature tensors to generate corresponding one-dimensional feature vectors.

[0141] The two feature vectors are concatenated along the channel direction; in an optional implementation, the reference identifier obtained by the target reference image index mapping can also be embedded and concatenated together.

[0142] The concatenated composite vector is input into a multilayer perceptron, which outputs a channel-level gated vector with values ​​ranging from 0 to 1 through a nonlinear mapping. Each element of this gated vector is used to adjust the reference information retention ratio of the corresponding feature channel.

[0143] Step 502: Concatenate the basic prior conditions and the reference condition features to generate spatial-level gating.

[0144] At the spatial level, the processing unit generates spatial-level gated inputs based on the basic prior conditions and reference conditions available at the decoding end.

[0145] The basic prior conditions and reference conditions are stacked and concatenated along the channel dimension. The concatenated 3D tensor is input into a 3×3 convolutional layer. After convolution and activation function processing, a single-channel spatial gating matrix is ​​output. Each element of this matrix corresponds to a spatial coordinate position in the feature map and is used to control the injection weight of reference information at that position.

[0146] Step 503: Based on the channel-level gating and the spatial-level gating, the reference condition features are jointly modulated to obtain the modulated reference condition features; when only one reference image is selected and frequency domain segmented modulation is not performed, the modulated reference condition features are added element-wise to the basic prior conditions to generate the enhanced prior conditions.

[0147] After obtaining the channel-level gating vector and the spatial-level gating matrix, the multiplier uses a matrix broadcasting mechanism to multiply the two element-wise, generating a three-dimensional joint gating tensor that has modulation capabilities in both the channel and spatial dimensions.

[0148] The three-dimensional joint gating tensor is further multiplied element-wise with the input reference condition features to achieve preliminary screening and modulation of the reference features, suppressing feature components that are not beneficial to the current compression task.

[0149] When multiple target reference images are selected, the gated fusion further includes a multi-reference complementary fusion step: Step 504: Calculate the correlation matrix between the reference condition features corresponding to each target reference image.

[0150] When the retrieval system returns multiple target reference images, to avoid the repeated injection of homogeneous information, the system initiates a multi-reference complementary evaluation process. For the K reference condition features extracted from the K target reference images, the computation module calculates the inner product between each pair of features and performs norm normalization, thereby constructing a correlation matrix of size K by K. The off-diagonal elements of this matrix quantify the overlap between any two reference images in the feature space.

[0151] Step 505: Determine the uniqueness score corresponding to each target reference image based on the correlation matrix, and normalize the uniqueness score to obtain the complementarity weight.

[0152] Based on the aforementioned correlation matrix, the system evaluates the information independence of each reference image relative to the other images. When the number of target reference images K is greater than 1, for the i-th reference image, the average of the absolute values ​​of its correlation with all other images is calculated, and the uniqueness score of the i-th image is obtained by subtracting the average from 1; when K equals 1, the complementarity weight is 1, and the correlation matrix calculation is skipped.

[0153] u_i=1-(1 / (K-1))*Σj≠i|Mij|; Where ui is the uniqueness score of the i-th reference image, K is the total number of target reference images, Σj≠i represents the summation of all indices j not equal to i, and |Mij| is the absolute value of the element in the i-th row and j-th column of the correlation matrix.

[0154] The system uses an exponential normalization function with a temperature parameter to convert the uniqueness scores of each image into complementary weights.

[0155] u*i=exp(ui / τ) / Σkexp(uk / τ); Where u*i represents the complementarity weights of the i-th reference image, exp is the natural exponential operation, τ is the temperature parameter used to control the sharpness of the weight distribution, and Σk represents the summation of the exponential terms over all K images. The temperature parameter can be set according to the training or validation results to adjust the suppression strength of redundant references.

[0156] Step 506: The modulation intensity of the channel-level gating and the spatial-level gating corresponding to each of the target reference images is weighted and adjusted using the complementary weights to obtain the weighted modulated reference condition features corresponding to each of the target reference images. The weighted modulated reference condition features are summed to obtain the multi-reference fusion reference features. When frequency domain split-band modulation is not performed, the multi-reference fusion reference features are added to the basic advanced prior conditions to generate the enhanced advanced prior conditions.

[0157] After obtaining the complementary weights for each reference image, the system multiplies these scalar weights by the 3D joint gating tensor of the corresponding image generated in step 503, thereby achieving a secondary scaling of the gating intensity. Subsequently, the weighted gating features of all reference images are summed to obtain multi-reference fusion reference features, which serve as inputs for subsequent direct fusion or frequency domain split-band modulation.

[0158] Step 507: Perform frequency domain transformation on the basic prior conditions and the modulated reference condition features or multi-reference fusion reference features respectively, and divide them into multiple frequency band features in the frequency domain.

[0159] To adapt to the differentiated requirements of different frequency components in remote sensing images for reference information, the system implements more refined gating operations in the frequency domain. Specifically, for the basic prior conditions and the aforementioned modulated reference condition features or multi-reference fusion reference features, a two-dimensional discrete cosine transform is independently performed along the spatial dimension to obtain the corresponding frequency domain features; then, the frequency domain coefficients are divided into multiple frequency bands such as low frequency, mid frequency, and high frequency according to a preset frequency index.

[0160] In one implementation, the frequency domain segmented modulation step may specifically include the following steps: Step 508: Combine the preset frequency band confidence modulation factor to independently generate the frequency band gating weights corresponding to each frequency band in the frequency domain.

[0161] For each frequency band, the computation module concatenates the basic prior frequency domain features, reference condition frequency domain features, and a preset frequency band confidence modulation factor of that frequency band, and inputs them into an independent convolutional layer with a kernel size of 1x1 to generate a gated weight matrix specific to that frequency band.

[0162] Specifically, it includes the following steps: Step 1, Latent frequency band decomposition: For the prior features H∈ Cy×Hy×Wy and reference condition feature Fref∈ Cy×Hy×Wy, perform 2D-DCT on each channel along the spatial dimension to obtain frequency domain features, and divide them into B frequency bands according to frequency. In this embodiment, B can be 3, corresponding to low frequency, mid frequency, and high frequency respectively, and b is the b-th frequency band: D(Hc)=DCT2D(Hc), D(Fref,c)=DCT2D(Fref,c), c=1,...,Cy; Frequency band division can be achieved using a sawtooth scanning frequency index. Let the total frequency coefficient be Hy×Wy, and it can be divided into three bands: low, medium, and high according to a preset ratio: Hc(b)=Maskb⊙D(Hc), b∈{1(low),2(medium),3(high)}; Similarly, for the reference feature, Fref,c(b) = Maskb⊙D(Fref,c) is obtained.

[0163] Step 2: Generation of independent gating weights for each frequency band: Gating weights are generated independently for each frequency band b.

[0164] Step 3, Frequency Domain Fusion Output: H′(b)=H(b)+Gb⊙Fref(b), b=1,...,B; Where Gb=σ(Conv1×1b([H(b),Fref(b),sb·1])), sb is the frequency band-level confidence modulation factor, and sb is the low-frequency band level confidence modulation factor. l It can be initialized to a higher value, high frequency band s h It can be initialized to a lower value; The frequency band confidence modulation factor can be generated based on the learnable frequency band parameters or on the global statistical information of the basic prior conditions and reference conditions. The above values ​​are only initialization examples.

[0165] Step 4, Frequency Domain Recombination and Inverse Transformation: After combining the fusion results of each frequency band according to the original frequency index position, perform a two-dimensional inverse discrete cosine transform to obtain the enhanced prior condition H′.

[0166] In practical applications, the reliability modulation factor of the low-frequency band can be initialized to a higher value to encourage the use of stable low-frequency structures of similar ground features; while the modulation factor of the high-frequency band can be initialized to a lower value to encourage the system to conservatively inject high-frequency textures that are susceptible to noise interference.

[0167] Step 509: Modulate the reference conditional frequency domain features of the corresponding frequency band using the frequency band gating weights, and fuse them with the basic prior frequency domain features of the corresponding frequency band to obtain the frequency band fusion features.

[0168] After obtaining the gate weights for each frequency band, the multiplier performs element-wise multiplication of the weight matrix with the reference conditional frequency domain features of the corresponding frequency band in the frequency domain. The reference components modulated in this frequency domain are then added element-wise with the basic prior frequency domain features of the corresponding frequency band, thereby independently injecting reference information in each frequency band and generating multiple frequency band fusion features.

[0169] Step 510: Combine the fusion features of each frequency band according to the frequency index, and perform inverse frequency domain transformation to map back to the spatial domain to generate the enhanced prior conditions.

[0170] The aforementioned multiple frequency band fusion features are combined according to their original frequency index positions to reconstruct a complete frequency domain feature matrix. A two-dimensional inverse discrete cosine transform is then performed on the frequency domain feature matrix to map it back from the frequency domain to the original spatial domain. The final output is an enhanced prior condition that has undergone joint modulation by frequency sensing and multi-reference complementary modulation.

[0171] To address the noise disturbance that may be caused by reference information, this scheme introduces frequency domain bandgap gating and multi-reference complementarity mechanisms. It makes full use of low-frequency structural information, carefully injects high-frequency detail differences, and automatically suppresses the redundant superposition of homogeneous references. This effectively prevents the interference of imaging condition differences on the probability estimation of the entropy model and improves reconstruction defects such as road breaks and building edge blurring.

[0172] Example 6: Further explanation of the domain adaptive entropy parameter prediction and two-stage lightweight training strategy.

[0173] In one possible implementation, the step of predicting entropy model parameters based on the enhanced prior conditions includes the following steps: Step 601: Input the enhanced prior conditions into the pre-trained basic prediction network to obtain the basic probability parameters.

[0174] After acquiring the enhanced hyperprior conditions that incorporate multidimensional external prior information, these are used as input data and passed to the basic entropy parameter prediction network. This basic prediction network has already solidified the feature mapping relationship for a general image distribution during the pre-training stage, which will be described later.

[0175] After forward propagation through multiple layers of convolutional kernels within the network, the system outputs a set of initial statistics to describe the distribution characteristics of the potential representation, specifically including the basic mean and basic scale parameters, which constitute the basic benchmark for probability estimation.

[0176] Step 602: Input the enhanced prior conditions into the pre-trained entropy parameter adaptation layer to obtain a domain correction term for the characteristics of satellite remote sensing images.

[0177] To compensate for the biases of the basic prediction network in handling the unique overhead view and repetitive textures of satellite remote sensing, the system sets up a computational branch that runs parallel to the basic prediction network.

[0178] The same enhanced prior conditions are input into a smaller entropy parameter adaptation layer. This adaptation layer is specifically designed to respond to the local frequency distribution and microstructural differences in remote sensing images, and outputs a set of statistical variables in the form of residuals.

[0179] This variable is the domain correction term, which includes a mean correction term and a scale correction term. It is used to quantitatively compensate for the basic statistics.

[0180] Step 603: The basic probability parameters are superimposed and corrected with the domain correction term to obtain the entropy model parameters.

[0181] After obtaining the two sets of outputs, the adder performs the element-by-element addition operation in the channel and spatial dimensions.

[0182] The final distribution mean is obtained by adding the base mean and the mean correction term, and the preliminary distribution scale is obtained by adding the base scale parameter and the scale correction term.

[0183] To satisfy the mathematical constraint that the scale parameter in the probability density function must be positive, the system performs a nonlinear smoothing activation operation on the initial distribution scale, for example, using the softplus function, and adds a zero overflow prevention constant to output the final entropy model parameters. The aforementioned mean and scale parameters are then provided to the arithmetic encoder for lossless recording of the quantized latent representation.

[0184] Step 604: Pre-train using a general image dataset to obtain a backbone network with basic feature extraction and image reconstruction capabilities.

[0185] The parameters of all the above online inference modules depend on the offline two-stage training strategy. In the first stage of general pre-training, the system loads a high-resolution image dataset containing rich natural scenes. Without introducing a reference branch, the weights of the analysis transform network, synthetic transform network, super-prior network, and basic prediction network are iteratively updated through backpropagation algorithm with the goal of minimizing the weighted sum of encoding bitrate and decoding distortion.

[0186] The overall rate-distortion joint loss function is L = R + λ * D; Where R is the estimated coding rate term, λ is the Lagrange multiplier that controls the tradeoff between compression ratio and image quality, and D is the distortion term that calculates the mean square error between the original image and the reconstructed image.

[0187] Pre-training was performed using a general image dataset (such as the Flickr2K dataset). Training parameters: patch size 256×256, batch size 8, Adam optimizer used, initial learning rate 1×10⁻⁶. -4 Cosine annealing to 1×10 -6 Training for 200 epochs.

[0188] After multiple rounds of iterative learning, the backbone network converged and acquired stable general image compression and feature representation capabilities.

[0189] Step 605: Perform domain adaptation fine-tuning using the remote sensing image dataset. While keeping the backbone network parameters frozen, jointly train the adaptation layer for extracting the reference condition features, the gated fusion network for generating the enhanced prior conditions, and the entropy parameter adaptation layer for predicting the entropy model parameters.

[0190] After entering the second stage of domain adaptation fine-tuning, the system switches to the satellite remote sensing image dataset and locks the backbone network weights trained in the first stage, preventing them from being updated during backpropagation. When constructing training batches, the system randomly retrieves several reference images from the locally stored reference image library for the current image to be compressed, simulating a real online dual-domain retrieval environment.

[0191] While keeping the backbone network parameters frozen, the system utilizes the complete forward propagation data stream including the reference branch to calculate the rate-distortion joint loss and the gradient. For the weight prediction network in the dual-domain similarity fusion, softmax or pairwise ranking loss is used for optimization within the candidate set during the training phase, enabling the gradient to act on the fusion weights; during the inference phase, Top-K discrete selection is performed based on the comprehensive evaluation score. The optimizer updates the weight parameters of the bottleneck adaptation layer in the reference condition extraction module, the weight prediction network in the dual-domain similarity fusion, the multidimensional adaptive gating fusion network, and the entropy parameter adaptation layer. This lightweight fine-tuning strategy not only keeps the training computation cost at a low level but also avoids the destruction of the original general compression capability by small-scale remote sensing datasets, enabling the model to achieve point-to-point adaptation to the spatial distribution and frequency characteristics of remote sensing images while retaining basic generalization performance.

[0192] According to one aspect of this application, part of the processing procedure is as follows: For the retrieved reference image R i This invention employs a reference encoder that shares weights with the master compression model before the L-layer transformation: ; Where: ga(1:L) represents the first L layers of the main analysis transform network; it shares weights with the coding backbone and can be frozen during the remote sensing domain adaptation stage; Aref is a lightweight reference adapter that can be trained; Eref(Ri) outputs reference intermediate features, which are mapped by a 1×1 convolution or a super-prior fusion adapter to obtain reference condition features Ri' aligned with the basic super-prior conditions H.

[0193] The reference adapter uses a Bottleneck structure: Aref(F)=F+Conv1x1up(φ(Conv3x3(φ(Conv1x1down(F))))); Conv1x1 down Used for channel dimensionality reduction, Conv3x3 for local adaptation, Conv1x1 up Used for channel recovery. This structure has a small number of parameters and can adapt to remote sensing image features without compromising the generality of the main compression model.

[0194] For each of the K reference images, we obtain R1', R2', ..., RK', and then map them to the same channel dimension and spatial scale as H through 1x1 convolution or a super-prior fusion adapter.

[0195] For each reference image Ri, channel-level gating and spatial-level gating are generated; the gating is generated based on the basic prior conditions H and reference condition features Ri' available at the decoding end.

[0196] Channel-level gating: Gich=sigmoid(MLPch(concat(GAP(H),GAP(Ri')))); Spatial-level gating: Gisp = sigmoid(Convsp(concat(H,Ri'))); Where H is the basic prior condition, Ri' is the reference condition feature corresponding to the i-th target reference image, and GAP represents global average pooling.

[0197] Joint modulation gating: ; When the reference condition features and the basic prior conditions have high consistency in channel statistics and spatial structure, gating tends to increase reference injection; when the reference condition features and the basic prior conditions differ greatly, gating reduces the intensity of reference injection to avoid compressing unhelpful reference interference entropy models.

[0198] To adapt to the differentiated reference information requirements of different frequency components in satellite remote sensing images, this invention performs a two-dimensional discrete cosine transform along the spatial dimension on the prior condition H and the reference condition feature Ri' or the fused reference feature Fref to obtain frequency domain features: D(H)=DCT(H); D(Fref)=DCT(Fref); According to the frequency index, it is divided into three frequency bands: low frequency, medium frequency, and high frequency. ; get: ; Gating is generated independently for each frequency band: G(b)=sigmoid(Convb([H(b),Fref(b),sb·1])); Where s b This is the frequency band reliability modulation factor. It can be used to adjust the low-frequency band s... l Initialize to a higher value to encourage the use of stable large-scale structural references; set the high-frequency band s h Initialize to a low value to keep the model conservative against texture noise and differences in imaging conditions.

[0199] Frequency band fusion is: H′(b) = H(b) + G(b) ⊙ Fref(b); Finally, the fusion results of each frequency band are combined according to the original frequency index and returned to the spatial domain by inverse DCT: Hfreq=IDCT(Compose({H′(b)}b∈{l,m,h})); This module makes reference injection no longer limited to the spatial or channel domain, but can control the intensity of reference utilization separately for low-frequency ground feature layout, mid-frequency boundary structure and high-frequency texture details.

[0200] When K reference images are retrieved and K is greater than 1, the correlation matrix between the reference condition features is calculated. For any two reference images Ri' and Rj', calculate: Mij =<GAP(Ri'),GAP(Rj')> / (||GAP(Ri')||2·||GAP(Rj')||2+ε); The uniqueness score for each reference image is: ui=1-(1 / (K-1))Σj≠i|Mij|; when K equals 1, the complementarity weight is 1.

[0201] The complementarity weights are obtained through softmax: ûi=exp(ui / τ) / Σk=1Kexp(uk / τ); Where τ is the temperature parameter. The multi-reference fusion reference feature is represented as: Fmulti=Σi=1Kûi·Gi⊙Ri'; when frequency domain segmented modulation is not performed, the enhanced prior condition is Hmulti=H+Fmulti; when frequency domain segmented modulation is performed, Fmulti is used as Fref to enter the frequency domain fusion step.

[0202] This mechanism deweights redundant references that are highly similar to other reference images, while upweighting references that provide unique ground feature structure or frequency information, thereby suppressing duplicate reference injections and improving multi-reference complementarity.

[0203] The basic hyperprior condition H is combined with the modulated reference condition features or multi-reference fused reference features using the direct fusion or frequency domain segmentation fusion methods described above to obtain a unique enhanced hyperprior condition H'. The enhanced hyperprior condition H' is then input into the basic entropy parameter prediction network hep. Obtain the basic probability parameters: ; At the same time, H' inputs the entropy parameter AdapterA ep The remote sensing domain correction term is obtained as follows: The final entropy model parameters are: μ = μ0 + Δμ, σ = softplus(σ0 + Δσ) + εσ; softplus and εσ are used to ensure that the scale parameter σ is positive. The encoder quantizes the latent representation based on μ and σ. Entropy encoding is performed; the decoding end uses the same H', μ, σ pair. Perform entropy decoding.

[0204] It should be noted that the above-mentioned number of network layers, number of channels, number of frequency bands, frequency band division ratio, training dataset and training hyperparameters are all exemplary configurations. Those skilled in the art can adjust them according to the remote sensing image resolution, sensor type, bit rate target and model size, and should not be construed as limiting the scope of protection of this application.

Claims

1. A reference enhancement compression method for satellite remote sensing images, characterized in that, include: Acquire the remote sensing image to be compressed, and extract the current image features, latent representation, and basic advanced prior conditions of the remote sensing image to be compressed; A dual-domain similarity evaluation is performed based on the current image features and the offline cache features of candidate images in the pre-built shared reference image library, and the evaluation results are fused to select the target reference image. Write the index of the target reference image into the compressed bitstream; The compressed bitstream is parsed at the decoding end to obtain the index, and the corresponding target reference image is read from the shared reference image library according to the index; Extract reference condition features from the target reference image, wherein the reference condition features are aligned with the compressed latent space in which the latent representation resides; The reference condition features are gated and fused with the basic hyperprior conditions to generate enhanced hyperprior conditions. Based on the parameters of the enhanced prior condition prediction entropy model; The latent representation is entropy encoded and decoded using the entropy model parameters to obtain the reconstructed remote sensing image.

2. The method according to claim 1, characterized in that, The current image features include current semantic features and current latent spatial features, and the offline cached features include cached semantic features and cached latent spatial features; A dual-domain similarity evaluation is performed based on the current image features and the offline cached features of candidate images in a pre-built shared reference library, including: A semantic similarity assessment is performed based on the current semantic features and the cached semantic features to obtain the semantic similarity. Based on the current latent space features and the cached latent space features, a compressed latent space similarity evaluation is performed to obtain the compressed latent space similarity.

3. The method according to claim 2, characterized in that, Based on the current latent space features and the cached latent space features, a compressed latent space similarity evaluation is performed to obtain the compressed latent space similarity, including: Extract the channel variance of the current latent spatial features and the cached latent spatial features; Calculate the channel-level statistical correlation between the current potential space features and the cached potential space features, and estimate the channel correlation coefficient based on the channel-level statistical correlation; Based on the channel correlation coefficient, the proportion of residual conditional variance after introducing reference conditions is determined, and the proportion of residual conditional variance is converted into channel potential bitrate savings. Using the variance of each channel in the current latent space as weight, the potential bitrate savings corresponding to each channel are weighted, summed, and normalized to obtain the compression latent space similarity.

4. The method according to claim 2, characterized in that, The evaluation results are fused, and a target reference image is selected, including: Extract the global pooling vector of the current latent spatial features, and concatenate the global pooling vector with the current semantic features to obtain an image content summary; The image content summary is input into a pre-trained weight prediction network, which outputs adaptive dual-domain fusion weights. The semantic similarity and the compressed latent space similarity are weighted and fused using the adaptive dual-domain fusion weights to obtain a comprehensive evaluation score; The target reference image is selected by sorting the candidate images in the shared reference image library according to their comprehensive evaluation scores.

5. The method according to claim 1, characterized in that, Extracting reference condition features from the target reference image includes: The target reference image is input into a feature extraction layer that shares weights with the backbone network to extract intermediate reference features. The intermediate reference features are input into a pre-trained adaptation layer for channel dimensionality reduction and local feature mapping, and the reference condition features aligned with the compressed latent space are output.

6. The method according to claim 1, characterized in that, Gating and fusing the reference condition features with the basic prior conditions includes: Extract the global spatial features of the basic prior conditions and the reference conditions to generate channel-level gating; By concatenating the basic prior conditions and the reference condition features, a spatial-level gating system is generated. The reference condition features are jointly modulated based on the channel-level gating and the space-level gating.

7. The method according to claim 6, characterized in that, After jointly modulating the reference condition features, the process further includes a frequency domain segmented modulation step: Frequency domain transformation is performed on the basic prior conditions and the modulated reference condition features respectively, dividing them into multiple frequency band features; Combined with a preset frequency band confidence modulation factor, frequency band gating weights corresponding to each frequency band feature are generated independently; The reference condition features of the corresponding frequency band are frequency-domain modulated using the frequency band gate weights and fused with the basic prior conditions of the corresponding frequency band to obtain frequency band fusion features; Perform an inverse frequency domain transformation on the frequency band fusion features to map them back to the spatial domain, thereby generating the enhanced prior conditions.

8. The method according to claim 6, characterized in that, When multiple target reference images are selected, the gated fusion further includes a multi-reference complementary fusion step: Calculate the correlation matrix among the reference condition features corresponding to each of the target reference images; Based on the correlation matrix, the uniqueness score corresponding to each target reference image is determined, and the uniqueness score is normalized to obtain the complementarity weight, wherein when the number of target reference images is 1, the complementarity weight is 1. The modulation intensities of the channel-level gating and the spatial-level gating corresponding to each of the target reference images are weighted and adjusted using the complementary weights to generate the enhanced prior conditions.

9. The method according to claim 1, characterized in that, Based on the parameters of the enhanced prior condition prediction entropy model, including: The enhanced prior conditions are input into the pre-trained basic prediction network to obtain the basic probability parameters; The enhanced prior conditions are input into a pre-trained entropy parameter adaptation layer to obtain a domain correction term for the characteristics of satellite remote sensing images. The basic probability parameters are superimposed and corrected with the domain correction term to obtain the entropy model parameters.

10. A reference enhancement and compression system for satellite remote sensing images, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the reference enhancement compression method for satellite remote sensing images as described in any one of claims 1 to 9.