A JPEG lossless recompression method based on multi-scale reference

By adopting multi-scale reference conditional encoding and decoding methods in JPEG image compression and combining feature alignment technology, the problem of image quality degradation in traditional JPEG image compression during multiple compression is solved, and more efficient image compression and recovery effects are achieved.

CN119653088BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510181685.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The traditional JPEG image compression method will lead to a decrease in image quality and a decrease in compression efficiency during multiple compression and decompression. The prior art has failed to fully utilize the conditional information of the reference image for effective recompression.

Method used

The JPEG lossless heavy compression method based on multi-scale reference is adopted. By inputting the current JPEG image and the reference JPEG image, downsampling, upsampling, conditional encoding, conditional decoding and feature alignment are performed, feature information of different scales are fused, and the DCT coefficients are more accurate probabilistic modeling is carried out to improve the compression efficiency of entropy coding.

Benefits of technology

It realizes the maintenance of high image quality and improves compression efficiency during multiple compression processes, reduces redundant data, and is suitable for large-scale image data processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653088B_ABST
    Figure CN119653088B_ABST
Patent Text Reader

Abstract

The present invention proposes a JPEG lossless recompression method based on multi-scale reference. The method of the present invention includes: inputting a current JPEG image and a reference JPEG image, performing corresponding processing, finally obtaining current JPEG encoded data, and outputting the current JPEG image. The corresponding processing includes: downsampling, upsampling, conditional encoding, conditional decoding, and feature alignment. By utilizing the multi-scale correlation characteristics of the reference JPEG and the current JPEG, the method of the present invention realizes multi-scale conditional encoding, improving the efficiency and quality of JPEG lossless recompression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image compression technology, and specifically to a JPEG lossless recompression method based on multi-scale reference. Background Art

[0002] The JPEG (Joint Photographic Experts Group) image compression standard is widely used in the fields of digital image processing and storage, and has become one of the most common image compression formats due to its good compression efficiency and low computational complexity. However, with the continuous development of image compression and transmission technologies, the traditional JPEG encoding method has gradually revealed its limitations in terms of compression quality and efficiency in some application scenarios, especially in cases where multiple compressions and decompressions are required. During the traditional JPEG recompression process, the image quality often significantly deteriorates, resulting in the loss of image details and a reduction in compression performance. Therefore, how to improve the quality and compression efficiency of JPEG images during multiple compression processes has become an important research issue in the field of image processing.

[0003] In recent years, image coding methods based on deep learning have gradually emerged. By learning the hyperprior features and complex probability models of images, they can significantly improve the compression effect. Especially with the support of conditional generation models and feature alignment technologies, image recompression can not only achieve a higher compression ratio, but also reduce distortion while retaining more image information. Therefore, researching JPEG recompression methods based on reference images and using the conditional information in the reference image to guide the compression and decompression of the current image has become an effective way to improve the quality of image recompression.

[0004] Most current technologies focus on directly optimizing compression by learning image features, but these methods usually ignore the potential of the reference image during the compression process and fail to make full use of the conditional information of the reference image for effective recompression. Therefore, how to combine the feature information of the reference image and the current image and perform refined recompression through conditional encoding and decoding processes to achieve higher compression quality and efficiency remains a technical problem to be solved urgently. Summary of the Invention

[0005] Object of the Invention: Aiming at the deficiencies of the prior art, the present invention proposes a JPEG lossless recompression method based on multi-scale reference, which is particularly suitable for scenarios that need to process large-scale picture data, such as cloud storage, video surveillance, image transmission, etc.

[0006] The method of the present invention includes: inputting a current JPEG image and a reference JPEG image, performing corresponding processing, finally obtaining current JPEG encoded data, and outputting the current JPEG image. The corresponding processing includes: downsampling, upsampling, conditional encoding, conditional decoding, and feature alignment.

[0007] The downsampling includes: taking the input current JPEG image and reference JPEG image as the first-scale current JPEG image C1 and the first-scale reference JPEG image R1 respectively, where C1 undergoes conditional encoding to obtain the first-scale encoded data B1; downsampling C1 and R1 respectively to obtain the second-scale current JPEG image C2 and the second-scale reference JPEG image R2, where C2 undergoes conditional encoding to obtain the second-scale encoded data B2; downsampling C2 and R2 respectively to obtain the third-scale current JPEG image C3 and the third-scale reference JPEG image R3, where C3 undergoes conditional encoding to obtain the third-scale encoded data B3.

[0008] The upsampling includes: the third-scale encoded data B3 undergoes conditional decoding to obtain the third-scale current JPEG image C3, and the third-scale current JPEG image C3 is upsampled to obtain the upsampled image C3’ of C3; the second-scale encoded data B2 undergoes conditional decoding to obtain the second-scale current JPEG image C2, and the second-scale current JPEG image C2 is upsampled to obtain the upsampled image C2’ of C2; the first-scale encoded data B1 undergoes conditional decoding to obtain the first-scale current JPEG image C1, and the first-scale current JPEG image C1 is used as the final output to obtain the current JPEG image.

[0009] The conditional encoding includes the following steps:

[0010] Step 1-1, extracting the hyperprior feature: parsing the input current JPEG image to obtain the DCT coefficients (discrete cosine transform coefficients) y of three color components, and then extracting the hyperprior feature z through the encoder function of the hyperprior encoder ;

[0011] ;

[0012] Performing entropy encoding on the extracted hyperprior feature z to obtain hyperprior feature encoded data , and the formula is:

[0013] ,

[0014] where is the entropy encoding function;

[0015] Step 1-2, conditional information fusion:

[0016] The goal of conditional information fusion is to combine the hyperprior feature z with the conditional information c to generate a fused feature representation , through the fusion function to complete:

[0017] ;

[0018] When encoding the current JPEG image C3 at the third scale, the third-scale reference JPEG image R3 is used as the conditional information c to guide the encoding; when encoding the current JPEG image C2 at the second scale, the data obtained by aligning the features of the upsampled image C3' and the second-scale reference JPEG image R2 is used as the conditional information c to guide the encoding; when encoding the current JPEG image C1 at the first scale, the data obtained by aligning the features of the upsampled image C2' and the first-scale reference JPEG image R1 is used as the conditional information c to guide the encoding;

[0019] Steps 1-3, conditional probability modeling and entropy coding: Input the fused feature representation into the entropy parameter network , and the entropy parameter network is used to predict the mean and variance of the DCT coefficient y, and the formula is:

[0020] ,

[0021] Then, use the mean and variance to perform probability modeling on the DCT coefficient y, and assume that the DCT coefficient y follows a Gaussian distribution, and the formula is:

[0022] ,

[0023] where represents the probability distribution of the DCT coefficient y under the condition of , and indicates that the probability distribution is a Gaussian distribution with a mean of and a variance of ;

[0024] According to the probability distribution of the DCT coefficient y, use the entropy coding algorithm to encode the DCT coefficient y, and finally obtain the encoded data of the DCT coefficient y ;

[0025] Steps 1-4, data encapsulation and output: Integrate the encoded data of the hyperprior feature and the encoded data of the DCT coefficient y to form the final JPEG compression data packet B:

[0026] ,

[0027] Among them, Encode is the encapsulation function.

[0028] The conditional decoding includes:

[0029] Step 2-1, De-encapsulation and extraction: De-encapsulate the input JPEG compressed data packet B, and the formula is:

[0030] ,

[0031] Among them, Decode is the de-encapsulation function;

[0032] Step 2-2, Hyperprior feature decoding: Entropy-decode the hyperprior feature encoded data , and the decoded hyperprior feature z is expressed as:

[0033] ,

[0034] Among them, represents the entropy decoding function;

[0035] Step 2-3, Conditional information fusion:

[0036] Fuse the decoded hyperprior feature z with the input conditional information c to generate the fused feature :

[0037] ,

[0038] When decoding the third-scale encoded data B3, use the third-scale reference JPEG image R3 as the conditional information c to guide the decoding; when decoding the second-scale encoded data B2, use the data after aligning the upsampled image C3' and the second-scale reference JPEG image R2 features as the conditional information c to guide the decoding; when decoding the first-scale encoded data B1, use the data after aligning the upsampled image C2' and the first-scale reference JPEG image R1 features as the conditional information c to guide the decoding;

[0039] Step 2-4, Conditional probability modeling and entropy decoding: Input the fused feature representation into the entropy parameter network in Step 1-3 to generate the mean and variance , and use the mean and variance to perform probability modeling on the DCT coefficient y, and then according to the probability distribution of the DCT coefficient y, use the entropy decoding algorithm to decode y to obtain the decoded data of the DCT coefficient y, and finally output the current JPEG image.

[0040] The feature alignment includes:

[0041] Step 3-1, feature extraction: Feature extraction is performed on the reference JPEG image and the current JPEG image respectively, and the feature representation of each image is extracted through a neural network. Let the feature representation of the reference JPEG image be , and the feature representation of the current JPEG image be . The process of feature extraction is:

[0042] ,

[0043] ,

[0044] wherein, and respectively represent the reference JPEG image and the current JPEG image, represents the neural network model for extracting image features;

[0045] Step 3-2, feature fusion: Generate a synthetic feature representation :

[0046] ,

[0047] wherein, is a fusion coefficient, and its value range is between 0 and 1. This coefficient controls the importance ratio of the features of the reference image and the current image;

[0048] Align the feature representation to obtain the aligned feature representation :

[0049] ;

[0050] wherein, is a transformation function for aligning the fused features.

[0051] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method.

[0052] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the method are executed.

[0053] Beneficial effects: The present invention realizes reference-based multi-scale conditional coding. By utilizing the correlation characteristics of the reference JPEG and the current JPEG image, fusing feature information at different scales, and performing more accurate probability modeling on DCT coefficients (discrete cosine transform coefficients), the compression efficiency of entropy coding is improved, and redundant data is effectively reduced. Description of the drawings

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 It is the schematic diagram of conditional coding.

[0056] Figure 2 It is the schematic diagram of conditional decoding.

[0057] Figure 3 It is the schematic diagram of feature alignment.

[0058] Figure 4 It is the flowchart for implementing conditional encoding and decoding of the current JPEG image.

[0059] Figure 5 It is the comparison chart of the bit savings rate per pixel between the method of the present invention and the traditional image compression scheme.

[0060] Figure 6 It is the comparison chart of the bits per pixel between the method of the present invention and the traditional image compression scheme. Detailed implementation manners

[0061] Now, the present disclosure will be discussed with reference to several example implementations. It should be understood that these implementations are discussed only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, rather than implying any limitation on the scope of the present disclosure.

[0062] The embodiment of the present invention provides a JPEG lossless recompression method based on multi-scale reference, including: inputting the current JPEG image and the reference JPEG image, performing corresponding processing, and finally obtaining the current JPEG encoded data and outputting the current JPEG image. The corresponding processing includes: downsampling, upsampling, conditional coding, conditional decoding, and feature alignment.

[0063] The downsampling includes:

[0064] The current JPEG image and the reference JPEG image of the input are respectively used as the current JPEG image C1 of the first scale and the reference JPEG image R1 of the first scale, where C1 is conditionally encoded to obtain the encoded data B1 of the first scale; C1 and R1 are respectively downsampled to obtain the current JPEG image C2 of the second scale and the reference JPEG image R2 of the second scale, where C2 is conditionally encoded to obtain the encoded data B2 of the second scale; C2 and R2 are respectively downsampled to obtain the current JPEG image C3 of the third scale and the reference JPEG image R3 of the third scale, where C3 is conditionally encoded to obtain the encoded data B3 of the third scale.

[0065] The upsampling includes: the encoded data B3 of the third scale is conditionally decoded to obtain the current JPEG image C3 of the third scale, and the current JPEG image C3 of the third scale is upsampled to obtain the upsampled image C3' of C3; the encoded data B2 of the second scale is conditionally decoded to obtain the current JPEG image C2 of the second scale, and the current JPEG image C2 of the second scale is upsampled to obtain the upsampled image C2' of C2; the encoded data B1 of the first scale is conditionally decoded to obtain the current JPEG image C1 of the first scale, and the current JPEG image C1 of the first scale is used as the final output to obtain the current JPEG image.

[0066] The conditional encoding includes the following steps: Step 1-1, extracting the hyperprior feature: the input current JPEG image is parsed to obtain the DCT coefficients (discrete cosine transform coefficients) y of three color components, and then the hyperprior encoder's encoder function is used to extract the hyperprior feature z:

[0067] ;

[0068] The extracted hyperprior feature z is entropy encoded to obtain the hyperprior feature encoded data , and the formula is:

[0069] ,

[0070] where is the entropy encoding function;

[0071] Step 1-2, conditional information fusion:

[0072] The goal of conditional information fusion is to combine the hyperprior feature z with the conditional information c to generate a fused feature representation , which is completed through the fusion function :

[0073] ;

[0074] When encoding the current JPEG image C3 at the third scale, the third-scale reference JPEG image R3 is used as the conditional information c to guide the encoding; when encoding the current JPEG image C2 at the second scale, the data after feature alignment of the upsampled image C3' and the second-scale reference JPEG image R2 is used as the conditional information c to guide the encoding; when encoding the current JPEG image C1 at the first scale, the data after feature alignment of the upsampled image C2' and the first-scale reference JPEG image R1 is used as the conditional information c to guide the encoding;

[0075] Steps 1-3, conditional probability modeling and entropy encoding: The fused feature representation is input into the entropy parameter network , and the entropy parameter network is used to predict the mean and variance of the DCT coefficient y. The formula is:

[0076] ,

[0077] Then, the mean and variance are used to perform probability modeling on the DCT coefficient y. It is assumed that the DCT coefficient y follows a Gaussian distribution. The formula is:

[0078] ,

[0079] where represents the probability distribution of the DCT coefficient y under the condition of , and indicates that the probability distribution is a Gaussian distribution with a mean of and a variance of ;

[0080] According to the probability distribution of the DCT coefficient y, the entropy encoding algorithm is used to encode the DCT coefficient y, and finally the encoded data of the DCT coefficient y is obtained;

[0081] Steps 1-4, data encapsulation and output: The encoded data of the hyperprior feature and the encoded data of the DCT coefficient y are integrated to form the final JPEG compression data packet B:

[0082] ,

[0083] where Encode is the encapsulation function.

[0084] The conditional decoding includes:

[0085] Step 2-1, Decapsulation and Extraction: Decapsulate the input JPEG compressed data packet B, with the formula:

[0086] ,

[0087] where Decode is the decapsulation function;

[0088] Step 2-2, Hyperprior Feature Decoding: Entropy-decode the hyperprior feature encoded data The decoded hyperprior feature z is expressed as:

[0089] ,

[0090] where represents the entropy decoding function;

[0091] Step 2-3, Conditional Information Fusion:

[0092] Fuse the decoded hyperprior feature z with the input conditional information c to generate the fused feature :

[0093] ,

[0094] When decoding the third-scale encoded data B3, use the third-scale reference JPEG image R3 as the conditional information c to guide the decoding; when decoding the second-scale encoded data B2, use the data after aligning the features of the upsampled image C3’ and the second-scale reference JPEG image R2 as the conditional information c to guide the decoding; when decoding the first-scale encoded data B1, use the data after aligning the features of the upsampled image C2’ and the first-scale reference JPEG image R1 as the conditional information c to guide the decoding;

[0095] Step 2-4, Conditional Probability Modeling and Entropy Decoding: Input the fused feature representation into the entropy parameter network in Step 1-3 to generate the mean and variance , and use the mean and variance to perform probability modeling on the DCT coefficient y, and then according to the probability distribution of the DCT coefficient y, use the entropy decoding algorithm to decode y to obtain the decoded data of the DCT coefficient y, and finally output the current JPEG image.

[0096] The feature alignment includes:

[0097] Step 3-1, Feature Extraction: Respectively perform feature extraction on the reference JPEG image and the current JPEG image, and extract the feature representation of each image through a neural network. Let the feature representation of the reference JPEG image be The feature of the current JPEG image is represented as The process of feature extraction is as follows:

[0098] ,

[0099] ,

[0100] Among them, and respectively represent the reference JPEG image and the current JPEG image, represents a neural network model for extracting image features;

[0101] Step 3-2, feature fusion: Generate a synthetic feature representation :

[0102] ,

[0103] Among them, is a fusion coefficient, and its value range is between 0 and 1. This coefficient controls the importance ratio of the features of the reference image and the current image;

[0104] Align the feature representation to obtain the aligned feature representation :

[0105] ;

[0106] Among them, is a transformation function for aligning the fused features.

[0107] The present invention also provides an electronic device, including a processor and a memory. The memory stores program codes. When the program codes are executed by the processor, the processor executes the steps of the method.

[0108] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method.

[0109] Figure 1 Illustrates the implementation principle of conditional coding. First, the input current JPEG image is parsed to obtain the DCT coefficients (discrete cosine transform coefficients) y of three color components; then, the hyperprior encoder extracts the hyperprior feature z from the DCT coefficients y, and z represents the deep information of the image; then, z is compressed through quantization and entropy coding to generate the encoded data of the hyperprior feature ; Subsequently, the extracted hyperprior feature z is fused with the input conditional information c to obtain the fused feature During the fusion process, the relative importance of the features of the reference JPEG image and the current JPEG image can be adjusted through weighted fusion; then, the fusion result is input into the entropy parameter network to generate the mean value of the DCT coefficient y and variance , which are used for subsequent probability modeling and compression. Using the generated mean value and variance, probability modeling is performed on the DCT coefficient y, and the DCT coefficient y is compressed through entropy coding technology to obtain the encoded data of the DCT coefficient y ; finally, by integrating the hyperprior feature encoded data and the encoded data of the DCT coefficient y , the final JPEG compression data packet B is formed.

[0110] Figure 2 shows the implementation principle of conditional decoding. First, the input JPEG encoded data B is unpacked to extract the hyperprior feature encoded data and the encoded data of the DCT coefficient y ; then, the hyperprior feature encoded data is decoded using entropy decoding technology, and the hyperprior feature z is restored through the hyperprior decoder. The hyperprior feature z is fused with the input conditional information c to generate a fused feature ; then, the fused feature is input into the entropy parameter network to generate the mean value for decoding and variance ; finally, the encoded data of the DCT coefficient y is entropy decoded using the mean value and variance , and the final current JPEG image is gradually restored.

[0111] Figure 3 shows the implementation principle of feature alignment. First, feature extraction is performed on the reference JPEG image and the current JPEG image respectively to obtain the feature representations of the two images; then, by weighted fusing these two feature representations, a synthetic feature representation is generated. The relative importance of the features of the reference JPEG image and the current JPEG image is adjusted by setting the fusion coefficient during the fusion process; next, the fused feature is aligned using the spatial transformation network to ensure that the features of the reference JPEG image and the current JPEG image are spatially consistent, thereby achieving the matching and consistency of the JPEG images. This feature alignment process provides more accurate image feature support for subsequent conditional encoding and decoding.

[0112] Figure 4The conditional encoding and decoding process of the current JPEG image is shown. First, the input current JPEG image and the reference JPEG image are respectively used as the first-scale current JPEG image C1 and the first-scale reference JPEG image R1; C1 and R1 are respectively downsampled to obtain the second-scale current JPEG image C2 and the second-scale reference JPEG image R2; then C2 and R2 are respectively downsampled to obtain the third-scale current JPEG image C3 and the third-scale reference JPEG image R3; then, using R3 as the conditional information to guide the conditional encoding of C3 to obtain the encoded data B3 of C3, and using R3 as the conditional information to guide the conditional decoding of B3 to recover C3; then, C3 is upsampled to obtain the upsampled image C3', the features of C3' and R2 are aligned, and the aligned data is used to guide the conditional encoding of C2 to obtain the encoded data B2 of C2, and the aligned data is used to guide the conditional decoding of B2 to recover C2; subsequently, C2 is upsampled to obtain the upsampled image C2', the features of C2' and R1 are aligned, and the aligned data is used to guide the conditional encoding of C1 to obtain the encoded data B1 of C1, and the aligned data is used to guide the conditional decoding of B1 to recover C1. C1 is used as the final output current JPEG image, and the encoded data B1, B2, and B3 are integrated to form the current JPEG image encoded data.

[0113] Figure 5 and Figure 6 respectively show the comparison of the method of the present invention (denoted as Ours) with four traditional single-image compression schemes and a similar image compression algorithm in terms of the performance metrics of bits per pixel savings rate and bits per pixel. The several schemes are JPEG, Lepton, JPEG XL, Brunsli, and the frequency domain block matching algorithm Frequency Domain Block Matching (denoted as FDBM, paper source: Hongwei Sha, Ming Lu, and Zhan Ma, "Lossless JPEG Recompression for Similar Images via Frequency Domain Block Matching," IEEE PCS, Feb. 2024), where FDBM is a similar image compression algorithm. The bar chart shows the performance of the six schemes on the Caltech Pedestrian dataset, the UVG 1080P video dataset, and the UVG 4K dataset. It can be seen that compared with the single-image compression algorithm and the similar image compression algorithm FDBM, the present method can achieve lower bits per pixel and higher bits per pixel savings rate.

[0114] The present invention provides a JPEG lossless recompression method based on multi-scale reference, which adopts technologies such as multi-scale reference, superprior feature extraction and compression, conditional encoding and decoding, and feature alignment, and can achieve efficient image compression and restoration. The present invention can adapt to the image compression requirements of different resolutions, and further improves the compression effect and restoration accuracy. Generally speaking, the method of the present invention has high efficiency and wide application potential in the field of image compression and restoration, and is especially suitable for scenarios that require a high compression ratio and have high requirements for image quality, such as image storage and transmission, etc.

Claims

1. A JPEG lossless recompression method based on multi-scale reference, characterized in that: include: Input a current JPEG image and a reference JPEG image, perform corresponding processing, finally obtain current JPEG encoded data, and output the current JPEG image, wherein the corresponding processing includes: downsampling, upsampling, conditional encoding, conditional decoding and feature alignment; The down sampling includes: The input current JPEG image and the reference JPEG image are respectively used as the first-scale current JPEG image C1 and the first-scale reference JPEG image R1, wherein C1 is conditionally encoded to obtain the first-scale encoded data B1; C1 and R1 are respectively downsampled to obtain the second-scale current JPEG image C2 and the second-scale reference JPEG image R2, wherein C2 is conditionally encoded to obtain the second-scale encoded data B2; C2 and R2 are respectively downsampled to obtain the third-scale current JPEG image C3 and the third-scale reference JPEG image R3, wherein C3 is conditionally encoded to obtain the third-scale encoded data B3; The upsampling comprises: The third-scale coded data B3 is conditionally decoded to obtain a third-scale current JPEG image C3, and the third-scale current JPEG image C3 is upsampled to obtain an upsampled image C3' of C3; the second-scale coded data B2 is conditionally decoded to obtain a second-scale current JPEG image C2, and the second-scale current JPEG image C2 is upsampled to obtain an upsampled image C2' of C2; the first-scale coded data B1 is conditionally decoded to obtain a first-scale current JPEG image C1, and the first-scale current JPEG image C1 is used as the final output to obtain the current JPEG image; When encoding the third-scale current JPEG image C3, the third-scale reference JPEG image R3 is used as condition information c to guide the encoding; when encoding the second-scale current JPEG image C2, the data after feature alignment of the upsampled image C3' and the second-scale reference JPEG image R2 is used as condition information c to guide the encoding; when encoding the first-scale current JPEG image C1, the data after feature alignment of the upsampled image C2' and the first-scale reference JPEG image R1 is used as condition information c to guide the encoding; When decoding the third-scale coded data B3, the third-scale reference JPEG image R3 is used as condition information c to guide decoding; when decoding the second-scale coded data B2, the data after feature alignment of the upsampled image C3' and the second-scale reference JPEG image R2 is used as condition information c to guide decoding; when decoding the first-scale coded data B1, the data after feature alignment of the upsampled image C2' and the first-scale reference JPEG image R1 is used as condition information c to guide decoding.

2. The method according to claim 1, characterized in that The conditional coding comprises the following steps: Step 1-1, extracting super-prior features: parse the input current JPEG image to obtain the DCT coefficients y of the three color components, and then pass the encoder function f of the super-prior encoder encoder () Extract super prior features z: z=f encoder (y); Perform entropy coding on the extracted super-prior feature z to obtain the super-prior feature coding data z encoded , the formula is: z encoded =EntropyEncoding(z), Among them, EntropyEncoding is the entropy coding function; Step 1-2, conditional information fusion: The goal of conditional information fusion is to combine the super prior feature z with the conditional information c to generate a fused feature representation z fusion , through the fusion function g fusion (z,c)Complete: z fusion =g fusion (z,c); Step 1-3, conditional probability modeling and entropy coding: The fused feature representation z fusion Input to the entropy parameter network EntropyNet, which is used to predict the mean μ and variance σ of the DCT coefficient y 2 , the formula is: μ,σ 2 =EntropyNet(from fusion ), Then use the mean μ and variance σ 2 Probabilistic modeling of the DCT coefficient y is performed, and the DCT coefficient y is set to conform to the Gaussian distribution. The formula is: p(y|z fusion )~N(μ,σ 2 ), Among them, p(y|z fusion ) indicates that the DCT coefficient y is at z fusion The probability distribution under the condition, N(μ,σ 2 ) indicates that the probability distribution is a distribution with a mean of μ and a variance of σ 2 Gaussian distribution, According to the probability distribution of DCT coefficient y, the DCT coefficient y is encoded using the entropy coding algorithm, and finally the encoded data c of DCT coefficient y is obtained. encoded ; Step 1-4, data packaging and output: encode the super-prior feature data z encoded and the coded data c of the DCT coefficient y encoded Integrate to form the final JPEG compressed data packet B: B=Encode(z encoded ,c encoded ), Among them, Encode is a packaging function.

3. The method according to claim 2, characterized in that The conditional decoding includes: Step 2-1, decapsulation and extraction: Decapsulate the input JPEG compressed data packet B. The formula is: (z encoded ,c encoded )=Decode(B), Among them, Decode is the decapsulation function; Step 2-2, super-prior feature decoding: decode the super-prior feature encoding data z encoded After entropy decoding, the decoded super prior feature z is expressed as: z=EntropyDecoding(z encoded ), Among them, EntropyDecoding represents the entropy decoding function; Step 2-3, conditional information fusion: The decoded super-prior feature z is fused with the input conditional information c to generate the fused feature zf usion : z fusion =g fusion (z,c), Step 2-4, conditional probability modeling and entropy decoding: The fused feature representation z fusion Input to the entropy parameter network EntropyNet in step 1-3 to generate mean μ and variance σ 2 , and use the mean μ and variance σ 2 The DCT coefficient y is probability modeled, and then according to the probability distribution of the DCT coefficient y, y is decoded using the entropy decoding algorithm to obtain the decoded data of the DCT coefficient y, and finally the current JPEG image is output.

4. The method according to claim 3, characterized in that The feature alignment includes: Step 3-1, feature extraction: perform feature extraction on the reference JPEG image and the current JPEG image respectively, and extract the feature representation of each image through the neural network. Suppose the feature representation of the reference JPEG image is F ref , the feature representation of the current JPEG image is F cur , the process of feature extraction is: F ref =FeatureExtrator(I ref ), F cur =FeatureExtrator(I cur ), Among them, I ref and I cur They represent the reference JPEG image and the current JPEG image respectively, and FeatureExtrator represents the neural network model used to extract image features; Step 3-2, feature fusion: generate a synthetic feature representation F fuse : F fuse =α·F ref +(1-α)·F cur , Among them, α is the fusion coefficient, which ranges between 0 and 1. This coefficient controls the importance ratio of the reference image and the current image features; For feature representation F fuse Align and obtain the aligned feature representation F align : F align =T align (F fuse ); Among them, T align It is a transformation function used to align the fused features.

5. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 4.

6. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 4 are executed.

Citation Information

Patent Citations

  • Video coding method based on depth level conditional probability prediction

    CN117692647A