Learned image and video compression by generative neural networks with learned sparse visual representation

WO2025091053A3PCT designated stage Publication Date: 2025-07-10FUTUREWEI TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017751
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing image and video compression methods based on neural networks face challenges in achieving high-quality reconstruction while maintaining efficient compression rates, particularly due to sensitivity to computational mismatches between encoding and decoding processes.

Method used

The implementation of a two-stream framework using generative neural networks with learned sparse visual representation (LSVR) for image and video compression. This approach encodes images and videos into embedding feature tensors, which are then mapped to integer codeword indices for efficient transmission, while a second stream provides fidelity-preserving information for high-quality reconstruction.

Benefits of technology

The proposed method achieves high perceptual quality and fidelity in reconstructed images and videos while maintaining a high compression rate, being robust to computational mismatches and suitable for heterogeneous platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017751_10072025_PF_FP_ABST
    Figure US2025017751_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A Learned Sparse Visual Representation (LSVR)-based Learned Image Compression (LIC) method implemented by a sending device. The LSVR-based LIC method includes encoding an input image into an embedding feature tensor, obtaining, using a set of learned visual Codebooks, a set of integer codeword indices based on embedding feature tensor; encoding the set of integer codeword indices into a compressed codeword string; determining a fidelity-preserving compressed string based on the input image, and transmitting the compressed codeword string and the fidelity-preserving compressed string to a receiving device.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Docket No.4502-82901 (6000680PCT02) Learned Image and Video Compression by Generative Neural Networks with Learned Sparse Visual Representation CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to United States Provisional Patent Application No. 63 / 558,742, filed February 28, 2024, which is incorporated by reference. TECHNICAL FIELD

[0002] This disclosure is related to Learned Image Compression (LIC) and Learned Video Compression (LVC) based on generative neural networks (GNN), and in particular to, systems and methods for LIC and LVC by GNN with Learned Sparse Visual Representation (LSVR). BACKGROUND

[0003] LIC-based on Neural Networks (NN) and LVC-based on NN have been largely studied in recent years and have shown superior performance over traditional coding methods such as Joint Photographic Experts Group (JPEG), Versatile Video Coding (VVC), and High Efficiency Video Coding (HEVC). SUMMARY

[0004] A first aspect relates to an LSVR-based LIC method implemented by a sending device, that includes encoding an input image x into an embedding feature tensor y, obtaining, using a setof learned visual Codebooks ^^, … , ^^ , a set of integer codeword indices ^^, … , ^^ based onembedding feature tensor y ; encoding the set of integer codeword indices ^^, … , ^^ into acompressed codeword string ^^^^^; determining a fidelity-preserving compressed string ^^based on the input image x, and transmitting the compressed codeword string ^^^^^and the fidelity-preserving compressed string ^^toward a receiving device.

[0005] Optionally, in a first implementation according to the first aspect, the input image x isAtty. Docket No.4502-82901 (6000680PCT02)a three-dimensional (3D) tensor having a first shape ^^ × ℎ^ × ^, wherein ^^ ,  ℎ^ ,   c are a firstwidth, a first height, and number of channels of the input image x respectively.

[0006] Optionally, in a second implementation according to the first aspect or anyimplementation thereof, the embedding feature tensor ^ has a second shape ^^ × ℎ^ × ^, wherein^^ ,  ℎ^ ,  d  are a second width, a second height, and number of feature channels respectively.

[0007] Optionally, in a third implementation according to the first aspect or any implementation thereof, the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input image x.

[0008] Optionally, in a fourth implementation according to the first aspect or any implementation thereof, wherein visual Codebook ^^includes ^^codewords, and each codeword being a d-dimension feature vector.

[0009] Optionally, in a fifth implementation according to the first aspect or anyimplementation thereof, wherein each index ^^,^ in ^^ (l = 1, … , ^^ × ℎ^) in the set of integercodeword indices ^^, … , ^^ corresponds to a codeword ^^,^ ∈ ^^ that is nearest to acorresponding latent feature vector ^^of an l-th super-pixel in the embedding feature tensor y,wherein ^^,^ = argmi$^%∈ &' Dist+^,, ^^-, and wherein Dist+^,, ^^- is a distance metric.

[0010] Optionally, in a sixth implementation according to the first aspect or any implementation thereof, the distance metric is L1 or L2 norm.

[0011] Optionally, in a seventh implementation according to the first aspect or any implementation thereof, wherein determining the fidelity-preserving compressed string ^^based on the input image x includes downsampling the input image x to obtain a downsampled image; and determining the fidelity-preserving compressed string ^^based on the downsampled image.

[0012] Optionally, in an eighth implementation according to the first aspect or anyimplementation thereof, wherein encoding the set of integer codeword indices ^^, … , ^^ into thecompressed codeword string ^^^^^includes calculating a frequency of codewords usage in a large training set; reordering the codewords in descending order based on the frequency to obtain areordered codeword indices; reordering the set of integer codeword indices ^^, … , ^^ based onAtty. Docket No.4502-82901 (6000680PCT02) the reordered codeword indices to obtain a reordered set of integer codeword indices; converting the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codeword indices to a Gaussian style bell shape; and encoding, using a Gaussian Mixture- based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^.

[0013] A second aspect relates to an LSVR-based LIC method implemented by a receiving device, that includes receiving a compressed codeword string ^^^^^and a compressed fidelity- preserving string ^^corresponding to an original input image x; determining a set of integercodeword indices ^^, … , ^^ based on the compressed codeword string ^^^^^; determining, usinga set of learned visual Codebooks ^^, … , ^^, a set of decoded feature tensors ^.^, … , ^.^ based onthe set of integer codeword indices ^^, … , ^^, wherein the compressed codeword string ^^^^^ isbased on the set of learned visual Codebooks ^^, … , ^^; determining a low-quality (LQ) substituteinput / 01based on the compressed fidelity-preserving string ^^, wherein the LQ substitute input / 01includes fidelity information of the original input image x; determining an LQ embedding feature tensor ^01based on the LQ substitute input / 01; and generating an output image / .based on the LQ embedding feature tensor ^01 and the set of decoded feature tensors ^.^, … , ^.^ .

[0014] Optionally, in a first implementation according to the second aspect, the method further includes upsampling, prior to determining the LQ embedding feature tensor ^01, the LQ substitute input / 01to an original size of the original input image x.

[0015] Optionally, in a second implementation according to the second aspect or any implementation thereof, for an l-th super-pixel, feature vector ^.^,^in decoded feature tensor ^.^ is a feature vector of codeword ^^,^.

[0016] Optionally, in a third implementation according to the second aspect or any implementation thereof, wherein generating the output image / . based on the LQ embeddingfeature tensor ^01 and the set of decoded feature tensors ^.^, … , ^.^ includes determiningweights ^ , … , ^ bas 01^ ^ ed on the LQ embedding feature tensor ^ ; determining a decodedfeature tensor ^. based on the set of decoded feature tensors ^.^, … , ^.^ and the weightsAtty. Docket No.4502-82901 (6000680PCT02)^ , … , ^ ; and generating t 01^ ^ he output image / . based on the LQ embedding feature tensor ^and the decoded feature tensor ^..

[0017] Optionally, in a fourth implementation according to the second aspect or any implementation thereof, wherein generating the output image / . based on the LQ embedding feature tensor ^01and the decoded feature tensor ^. includes tuning the decoded feature tensor ^. based on the LQ embedding feature tensor ^01through an affine transformation to generates a tuned feature ^.∗; and generating the output image / . based on the LQ embedding feature tensor ^01and the tuned feature ^.∗.

[0018] Optionally, in a fifth implementation according to the second aspect or any implementation thereof, the affine transformation is implemented by a neural network with a set of parameters θ determined through training.

[0019] A third aspect relates to an LSVR-based LVC method implemented by a sending device, that includes encoding an input video frame / 4into an embedding feature tensor ^4, wherein theinput video frame / 4 is a video frame at time stamp t in a set of n video frames X = / ^, … , / 6;obtaining, using a set of learned visual Codebooks ^^, … , ^^ , a first set of codeword indices^4,^, … , ^4,^ based on the embedding feature tensor ^4 ; determining a second set of codewordindices ∆^4,^, … , ∆^4,^ based on the first set of codeword indices ^4,^, … , ^4,^ and a third set ofcodeword indices ^48^,^, … , ^48^,^ corresponding to a previous video frame / 48^ in the set of nvideo frames X = / ^, … , / 6; encoding the second set of codeword indices ∆^4,^, … , ∆^4,^ into acompressed codeword string ^^^^^,4; determining a fidelity-preserving compressed string ^^,4based on the input video frame / 4, and transmitting the compressed codeword stringthe fidelity-preserving compressed string ^^,4toward a receiving device.

[0020] Optionally, in a first implementation according to the third aspect, the input video frame / 4 is a 3D tensor having a first shape ^^9 × ℎ^9 × ^, where ^^9 × ℎ^9 ,  c are a first width, a firstheight, and number of channels of the input video frame / 4respectively.

[0021] Optionally, in a second implementation according to the third aspect or anyimplementation thereof, the embedding feature tensor ^4 has a second shape ^^9 × ℎ^9 × ^ ,Atty. Docket No.4502-82901 (6000680PCT02)wherein ^^9 , ℎ^9 ,  d   are a second width, a second height, and number of feature channelsrespectively.

[0022] Optionally, in a third implementation according to the third aspect or any implementation thereof, the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input video frame / 4.

[0023] Optionally, in a fourth implementation according to the third aspect or any implementation thereof, wherein visual Codebook ^^comprises ^^codewords, and each codeword being a d-dimension feature vector.

[0024] Optionally, in a fifth implementation according to the third aspect or anyimplementation thereof, wherein each index ^^,4,^ in ^^,4 (l = 1, … , ^^9 × ℎ^9) in the first set ofcodeword indices ^4,^, … , ^4,^ corresponds to a codeword ^^,4,^ ∈ ^^ that is nearest to acorresponding latent feature vector ^4,^of an l-th super-pixel in the embedding feature tensor ^4,wherein ^^,4,^ = argmi$^%∈ &' Dist:^, , ^4,^;, and wherein Dist:^,, ^4,^; is a distance metric.

[0025] Optionally, in a sixth implementation according to the third aspect or any implementation thereof, the distance metric is L1 or L2 norm.

[0026] Optionally, in a seventh implementation according to the third aspect or any implementation thereof, wherein determining the fidelity-preserving compressed string ^^,4based on the input video frame / 4includes downsampling the input video frame / 4to obtain a downsampled video frame; and determining the fidelity-preserving compressed string ^^,4 based on the downsampled video frame.

[0027] Optionally, in an eighth implementation according to the third aspect or anyimplementation thereof, wherein encoding the first set of codeword indices ^4,^, … , ^4,^ into thecompressed codeword string ^^^^^,4includes calculating a frequency of codewords usage in a large training set; reordering the codewords in descending order based on the frequency to obtaina reordered codeword indices; reordering the first set of codeword indices ^4,^, … , ^4,^ based onthe reordered codeword indices to obtain a reordered set of integer codeword indices; convertingAtty. Docket No.4502-82901 (6000680PCT02) the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codeword indices to a Gaussian style bell shape; and encoding, using a Gaussian Mixture- based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^,4.

[0028] Optionally, in a ninth implementation according to the third aspect or anyimplementation thereof, each entry ∆^4,^,^ in ∆^4,^ in the second set of codeword indices∆^4,^, … , ∆^4,^ is defined as: ∆^4,^,^ = ^4,^,^ if ^4,^,^ ≠ ^48^,^,^, otherwise ∆^4,^,^ = −1, whereinfor time stamp t=0, ∆^4,^ = ^4,^ .

[0029] A fourth aspect relates to an LSVR-based LVC method implemented by a receiving device, that includes receiving a compressed codeword string ^^^^^,4and a compressed fidelity- preserving string ^^,4corresponding to a video frame / 4at time stamp t in the set of n videoframes X = / ^, … , / 6; determining a first set of integer codeword indices ∆^4,^, … , ∆^4,^ basedon the compressed codeword string ^^^^^,4; determining a second set of integer codeword indices^4,^, … , ^4,^ based on the first set of integer codeword indices ∆^4,^, … , ∆^4,^ and a third set ofinteger codeword indices ^48^,^, … , ^48^,^ corresponding to a previous video frame / 48^ in theset of n video frames X = / ^, … , / 6; determining, using a set of learned visual Codebooks^^, … , ^^ respectively, a set of decoded feature tensors ^.4,^, … , ^.4,^ based on the second set ofinteger codeword indices ^4,^, … , ^4,^ , wherein the compressed codeword string ^^^^^,4 is basedon the set of learned visual Codebooks ^^, … , ^^; determining an LQ substitute input / 014 basedon the compressed fidelity-preserving string ^^,4, wherein the LQ substitute input / 014 comprises fidelity information of the video frame / 4; determining an LQ embedding feature tensor ^014 based on the LQ substitute input / 014 ; and generating an output image / .4based on the LQembedding feature tensor ^014set of decoded feature tensors ^.4,^, … , ^.4,^ .

[0030] Optionally, in a first implementation according to the fourth aspect, for time stamp t=0,^4,^ = ∆^4,^ .

[0031] Optionally, in a second implementation according to the fourth aspect or anyAtty. Docket No.4502-82901 (6000680PCT02) implementation thereof, the method further includes upsampling, prior to determining the LQ substitute input / 014 , the LQ substitute input / 014 to an original size of the video frame / 4.

[0032] Optionally, in a third implementation according to the fourth aspect or anyimplementation thereof, for an l-th super-pixel, feature vector ^.4,^,^ in decoded feature tensor^.4,^  is a feature vector of codeword ^4,^,^..

[0033] Optionally, in a fourth implementation according to the fourth aspect or any implementation thereof, wherein generating the output image / .4based on the LQ embeddingfeature tensor ^014 and the set of decoded feature tensors ^.4,^, … , ^.4,^ includes determiningweights ^ , … , ^ based on the LQ em 01^ ^ bedding feature tensor ^4 ; determining a decodedfeature tensor ^. based on the set of decoded feature, … , ^.4,^ and the^^, … , ^^ ; and generating the output image / .4 based on the LQ embedding feature tensor^014  and the decoded feature tensor ^.4.

[0034] Optionally, in a fifth implementation according to the fourth aspect or any implementation thereof, wherein generating the output image / .4based on the LQ embedding feature tensor ^014 and the decoded feature tensor ^.4includes tuning the decoded feature tensor ^.4based on the LQ embedding feature tensor ^014 through an affine transformation to generates a tuned feature ^.4∗; and generating the output image / .4based on the LQ embedding feature tensor ^014 and the tuned feature ^.4∗.

[0035] Optionally, in a sixth implementation according to the fourth aspect or any implementation thereof, the affine transformation is implemented by a neural network with a set of parameters θ determined through training.

[0036] A fifth aspect relates to an apparatus comprising a memory configured to store instructions; and one or more processors coupled to the memory and configured to execute the instructions to cause the apparatus to perform the method according to any of the preceding aspects or any implementation thereof.Atty. Docket No.4502-82901 (6000680PCT02)

[0037] A sixth aspect relates to a computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium, the computer- executable instructions when executed by one or more processors of an apparatus, cause the apparatus to perform a method according to any of the preceding aspects or any implementation thereof.

[0038] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create a new embodiment within the scope of the present disclosure.

[0039] These and other features, and the advantages thereof, will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF DRAWINGS

[0040] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0041] FIG. 1 is a schematic drawing illustrating a traditional sender-side video coding pipeline.

[0042] FIG. 2 is a schematic drawing illustrating a traditional receiver-side video coding pipeline.

[0043] FIG. 3 is a schematic drawing illustrating a two-stream framework for LSVR-based LIC according to an embodiment of the present disclosure.

[0044] FIG. 4 is a schematic drawing illustrating a two-stream framework for LSVR-based LVC according to an embodiment of the present disclosure.

[0045] FIG. 5 is a schematic drawing illustrating a LSVR-based LIC decoding pipeline according to an embodiment of the present disclosure.

[0046] FIG. 6 is a schematic drawing illustrating a LSVR-based LVC decoding pipelineAtty. Docket No.4502-82901 (6000680PCT02) according to an embodiment of the present disclosure.

[0047] FIG. 7 is a flowchart illustrating a LSVR-based LIC method implemented by a sending device according to an embodiment of the present disclosure.

[0048] FIG. 8 is a flowchart illustrating a LSVR-based LIC method implemented by a receiving device according to an embodiment of the present disclosure.

[0049] FIG.9 is a flowchart illustrating a LSVR-based LVC method implemented by a sending device according to an embodiment of the present disclosure.

[0050] FIG. 10 is a flowchart illustrating a LSVR-based LVC method implemented by a receiving device according to an embodiment of the present disclosure.

[0051] FIG. 11 is a diagram illustrating an apparatus according to an embodiment of the present disclosure. DESCRIPTION OF EMBODIMENTS

[0052] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.

[0053] The present disclosure is related to LIC and LVC based on GNN. In particular, the present disclosure describes a two-stream framework for both end-to-end î using GNN with a highly efficient codebook-based learned sparse visual representation (LSVR) (referenced herein as LSVR-based LIC and LSVR-based LVC). The first stream utilizes a high-quality (HQ) LSVR to generate images and videos with high perceptual quality, and the second stream provides fidelity-preserving information to guide the conditioned image and video generation in the first stream. By combining the power of HQ generative prior modelling through LSVR with highAtty. Docket No.4502-82901 (6000680PCT02) perceptual quality with the fidelity-preserving control signals, the disclosed embodiments achieve HQ reconstruction with both high perceptual quality and high fidelity, while using high compression rate to achieve efficient bits for storage and transmission. Additionally, the disclosed embodiments are robust to transmission or computation error in heterogeneous software and hardware platforms, and can be used for both LIC and LVC.

[0054] In general, for LIC, on the sender side, an input image x is passed through an encoding network to generate an image embedding feature ^ , which is further compressed through quantization and entropy coding into a data string that is efficient for storage and transmission. On the receiver side, a decoded image embedding feature ^. is recovered from the received string through entropy decoding and dequantization. Then a decoding network reconstructs an output image / . based on the decoded image embedding feature ^.. The target is to minimize the restoration loss between the reconstructed output / . and the original input x, and to minimize the bits used to represent the image embedding feature ^ for storage and transmission.

[0055] Similarly, in an LVC-based on NN, given a video segment comprising of n frames A = / ^, … , / 6 on the sender side, a set of embedding features B = ^^, … , ^6, are generated, which arefurther compressed through quantization and entropy coding into a data string that is efficient forstorage and transmission. On the receiver side, a set of decoded embedding features BC =^.^, … , ^.6 are recovered from the received data string through entropy decoding and dequantization,based on which a set of output frames AC = / .^, … , / .6 are reconstructed.

[0056] Most previous LIC / LVC methods are based on a hyperpriors framework, where an entropy model encodes / decodes the quantized latent feature for efficient transmission. One vital issue of the hyperprior framework is the extreme sensitivity to small differences between the encoder and decoder in calculating the hyperpriors. For instance, even floating round-off error can lead to catastrophic error propagation in the decoded latent feature. Existing works mostly assume homogeneous platforms and deterministic processing calculation. Some works use integer NN to prevent non-deterministic processing computation. Some methods design special NN modules that are computational friendly to speed up inference. However, such solutionsAtty. Docket No.4502-82901 (6000680PCT02) cannot be easily generalized to arbitrary network architectures. LSVR is a technique in machine learning that focuses on representing images efficiently using a sparse set of learned features. In particular, LSVR learns one or multiple highly compressed visual codebooks^^, … , ^^ , ^ℎDED F ≥ 1, using a vector-quantized autoencoder that is trained with adversarial andperceptual loss. Non-limiting examples of vector-quantized autoencoders include Vector Quantized Generative Adversarial Network (VQGAN) and Adaptive Code Representation (AdaCode). Each codebook targets at a semantic category of images such as outdoor, portrait, sports, etc. An input image / video can be encoded into the latent space spanned by the visual codebooks as an embedding latent feature ^, which is then mapped into a sequence of codewordindices ^^, … , ^^ , using the learned codebooks ^^, … , ^^ , respectively. The codewords indicesare integers, which can be effectively stored or transferred. Then the receiver side uses the samecodebooks to recover a set of decoded features ^.^, … , ^.^ (e.g., ^., by using codewords incodebook ^,corresponding to the received codeword indices ^,). Then an output image can bereconstructed based on these decoded features ^.^, … , ^.^ .

[0057] Transferring integer indices is very robust to heterogeneous platforms. For instance, by encoding codeword indices instead of latent features, LSVR-based compression does not suffer from the sensitivity of computation mismatch between the sender and receivers. Transferring indices gives the freedom of expanding latent feature dimension (often associated with better representation power for better reconstruction) without increasing bitrate, in comparison to transferring latent features or residues. Generative LSVR increases robustness to input degradations. Realistic and rich textures can be generated using HQ codebooks even for LQ inputs.

[0058] Prior use of LSVR for LIC includes using LSVR for face-centric image compression. For instance, due to the highly structural characteristics of human faces, an HQ codebook can be robustly learned for face reconstruction. LSVR for LIC over general image content has been applied, where codeword features from multiple codebooks that were learned according to diversified semantic categories were weighted combined to improve the semantic-class-agnosticAtty. Docket No.4502-82901 (6000680PCT02) restoration performance. A weight masking and re-prediction technique was used to reduce the transmission overhead of the dense weight map to achieve reasonable bitrate. However, one issue of previous LSVR-based compression methods for LIC is the difficulty in achieving HQ reconstruction. LSVR-based HQ restoration relies on a dense connection of multi-scale features between the embedding network (encoder) and reconstruction network (decoder), e.g., through skip connections. Such intermediate features are too large to transfer, defeating the purpose of the compression task. As a result, previous LSVR-based compression methods for LIC remove such direct feature links. For general images, even if the restored images may look okay perceptually, important fidelity and rich details are usually lost in the reconstruction result.

[0059] FIG. 1 is a schematic drawing illustrating a sender-side video coding pipeline 100 implemented by a sending device in accordance with traditional LVC methods. The sender-side video coding pipeline 100 includes an Intra-coded frame (I-Frame) Selection module 102, an NN- based LIC Encoder module 104, NN-based LIC Decoder module 106, a Reference Frame Selection module 108, an NN-based Motion Estimation module 110, an NN-based Motion Compensation module 112, and an NN-based Decoder 114. As shown in FIG.1, on the sender side, the I-FrameSelection module 102 receives a Group of Pictures (GoP) comprising of n frames A = / ^, … , / 6.The I-Frame Selection module 102 is configured to select m I-frames AH = / H H,I , … , / ,J from then frames, where 1 ≤ L ≤ $. Each I-frame / H,M can be any frame in the GoP, i.e., 1 ≤ NO ≤ $.The m I-frames AH = / H,I , … , / H,J contain a complete image and serve as reference frames for theGoP. The m I-frames AH = / H H,I , … , / ,J are then compressed by the NN-based LIC Encodermodule 104 to generate a set of compressed I-frame features BH = ^H H,I , … , ^,J , which istransmitted to a receiver side. The NN-based LIC Decoder module 106 is configured to computereconstructed I-frames ACH = / .H , … , / .H base H,I ,J d on the set of compressed I-frame features B =^H , … , ^H . PredictiveI,J frames that encode only the differences from aframe) and bidirectional predictive (B) frames (i.e., frames that encode only the differencesAtty. Docket No.4502-82901 (6000680PCT02) between both a preceding frame and a following frame) are then compressed based on thereconstructed I-frames / .H H,I , … , / .,J . In general, to compress a remaining frame / 4 (i.e., the nframes that were not the I-Frame Selection module 102), the Reference FrameSelection module 108 selects a set of reference frames P4comprising the reconstructed m I-frames AH = / H , … , / H an Q Q Q,I ,J d a set of k previously reconstructed frames / .^I , … , / .^' , where / .^Mcan be before or after t. The NN-based Motion Estimation module 110 thena motion vector R4based on the reference frames P4and the remaining frame / 4. Then the NN-based Motion Compensation module 112 computes a residual E4based on the motion vector R4, the reference frames P4, and the remaining frame / 4. The motion vector R4and residual E4are transmitted to the receiver side, usually with additional entropy encoding and decoding process to further reduce transmission overhead. Finally, the NN-based Decoder 114 reconstructs the output frame / .4based on the reference frames P4, the residual E4, and the motion vector R4. The reconstructed output frame / .4is then fed back to the Reference Frame Selection module 108 for future use in selecting reference frames P4.

[0060] FIG. 2 is a schematic drawing illustrating a receiver-side video coding pipeline 200 implemented by a receiver device in accordance with traditional LVC methods. The receiver- side video coding pipeline 200 includes an NN-based LIC Decoder module 202, a Reference Frame Selection module 204, and an NN-Based Decoder 206. As described in FIG. 1, thesending device transmits the set of compressed I-frame features BH = ^H H,I , … , ^,J , the residual E4,and the motion vector R , which are received by the receivingreceivingusing the NN-based LIC Decoder module 202, computes reconstructed I-frames ACH = / .H H,I , … , / .,Jbased on the received I-frame features BH = ^H,I , … , ^H,J Then for each remaining frame / 4, theReference Frame Selection moduleof reference frames P4comprising thereconstructed I-frames / .H , … , / .H and a set of k previ Q Q,I ,J ously reconstructed frames / .^I , … , / .^' .The NN-Based Decoder 206 then reconstructs the output frame / .4based on the reference framesAtty. Docket No.4502-82901 (6000680PCT02) P4, the motion vector R4, and the residual E4. The overall target of LVC is to minimize therestoration loss between the reconstructed output frames AC = / .^, … , / .6 and the original inputframes A = / , … , / , and to minimize the transmitted data including embeddin H^ 6 g features B =^H,I , … , ^H,J for I-frames, the residule E4, and the motion vector R4 for the remaining frames.

[0061] As described above, prior LVC methods implement a traditional video coding pipeline such as the sender-side video coding pipeline 100 in FIG. 1. However, the traditional video coding pipeline is not designed for end-to-end LVC, and thus, the compression performance is innately bounded. For example, errors in each processing module can accumulate, leading to inferior compression performance and computation inefficiency. Also, prior LIC methods have completely different processing pipelines than the LVC methods, thus making prior LIC and LVC methods difficult for industrial implementation and optimization for various applications. To address the above issues, the present disclosure describes a two-stream framework for both end- to-end LSVR-based LIC and LSVR-based LVC.

[0062] FIG. 3 is a schematic drawing illustrating a two-stream LSVR-based LIC framework 300 according to an embodiment of the present disclosure. Beginning on the sender side, the LSVR-based LIC framework 300 includes an Encoding module 302, a set of learned visualCodebooks ^^, … , ^^ 304, a Codeword Encoding module 306, and a Fidelity-PreservingEncoding module 308 on the sender side. In an embodiment, the sending device is provided withor obtains an image x. In an embodiment, the image x is a 3D tensor with shape ^^ × ℎ^ × ^,where ^^ , ℎ^ , ^ are the width, height, and number of channels of the image. For example, insome embodiments, c = 3 for color images, c = 1 for spectral images, or c = 4 for RGB-D(color and depth) images. In a first stream of the LSVR-based LIC framework 300, the Encoding module 302 encodes the input x into an embedding feature tensor ^. In an embodiment, theembedding feature tensor ^ is a 3D tensor with shape ^^ × ℎ^ × ^, where ^^ is the width, ℎ^is the height, and d is the number of feature channels. In an embodiment, the width ^^and height ℎ^depend on the width and height of the input image x as well as the network structure of the Encoding module 302. The embedding feature tensor ^ represents the key features of theAtty. Docket No.4502-82901 (6000680PCT02) image x in a lower-dimensional space and serves as a compact numerical representation of key characteristics of the image x. In an embodiment, the Encoding module 302 is a neural network that has been trained to produce an embedding feature tensor of an input image. Various neural networks can be used as the Encoding module 302. Then, the embedding feature tensor ^ ismapped to a set of integer codeword indices ^^, … , ^^ , using the set of learned visual Codebooks^^, … , ^^ 304, respectively. The set of learned visual Codebooks ^^, … , ^^ 304 are sets ofrepresentative feature vectors (also called “codes” or “codewords”) that are learned from image data (e.g., from still images or video images) that can be used to encode visual patterns in images in a compact and structured way. For example, due to the highly structural characteristics of human faces, an HQ codebook can be robustly learned for facial images. In an embodiment, thevisual Codebook ^^ comprises ^^ codewords, each codeword being a d-dimension featurevector. In an embodiment, each index ^^,^ in ^^ ( U = 1, … , ^^ × ℎ^ ) corresponds to acodeword ^^,^ ∈ ^^ that is nearest to the corresponding latent feature vector ^^ of the l-th “super-pixel” in the embedding feature tensor y: ^^,^ = VEWLN$^%∈ &' XNYZ+^, , ^^-,

[0063] where XNYZ+- is a distance metric such as L1 or L2 norm. Then, the CodewordEncoding module 306 encodes the set of integer codeword indices ^^, … , ^^ into a compressedcodeword string ^^^^^in a lossless way, which can be efficiently transmitted to the receiver side. Additional details of the Codeword Encoding module 306 are further described below.

[0064] Additionally, on the sender side, in a second stream of the LSVR-based LIC framework 300, the input image x is fed into the Fidelity-Preserving Encoding module 308. The Fidelity- Preserving Encoding module 308 is configured to compute a fidelity-preserving compressed string ^^, which is transmitted to the receiver side. Fidelity refers to how accurately a reconstructed image matches the original image. In an embodiment, the fidelity-preserving compressed string ^^carries fine detail image information (e.g., colors, textures, etc.) to achieve HQ reconstruction with both high perceptual quality and high fidelity. In an embodiment, the fidelity-preservingAtty. Docket No.4502-82901 (6000680PCT02) compressed string ^^has a low bitrate that is comparable with the low bitrate of the codeword string ^^^^^. In some embodiments, a conventional image compression encoder designed to optimize the pixel-level fidelity between the decompressed image and the original input image (e.g., a JPEG encoder) is used as the Fidelity-Preserving Encoding module 308. In some embodiments, to achieve a very low bitrate for the fidelity-preserving compressed string ^^, the input image x is first downsampled, and then encoded with either an LIC encoder or a conventional image compression encoder.

[0065] Referring now to the receiver side of the LSVR-based LIC framework 300, the receiver / receiving device includes a Codeword Decoding module 312, the set of learned visualCodebooks ^^, … , ^^ 304, a Decoding module 316, a Fidelity-Preserving Decoding module 318,and an LQ Embedding module 320. As described above, the receiving device receives the codeword string ^^^^^carrying encoded characteristic information of the image x and the fidelity-preserving compressed string ^^carrying fine detail image information for preserving the fidelity of the image x. In an embodiment, the Codeword Decoding module 312 first recoversthe integer codeword indices ^^, … , ^^ from the received codeword string ^^^^^ . Then, usingthe same visual Codebooks ^^, … , ^^ 304 as the sender side, a set of decoded feature tensors^.^, … , ^.^ is determined based on the integer codeword indices ^^, … , ^^. In an embodiment, forthe l-th “super-pixel”, the feature vector ^.^,^in the decoded feature tensor ^.^is the feature vector of the codeword ^^,^.

[0066] Additionally, the Fidelity-Preserving Decoding module 318 computes an LQ substitute input / 01based on the received fidelity-preserving string ^^. The LQ substitute input / 01contains fidelity information of the original input x. In some embodiments, when the image x is downsampled by the sender, the substitute input / 01is also upsampled to be the original size same as the original image x. In an embodiment, the Fidelity-Preserving Decoding module 318 uses the decoder part of the LIC or conventional image compression method corresponding to the Fidelity-Preserving Encoding module 308. The LQ substitute input / 01is further fed into the LQ Embedding module 320. The LQ Embedding module 320 computes an LQ embeddingAtty. Docket No.4502-82901 (6000680PCT02) feature tensor ^01based on the LQ substitute input / 01. The Decoding module 316 reconstructs an output image / . based on the LQ embedding feature tensor ^01and the decodedfeature tensors ^.^, … , ^.^ . In some embodiments, the output image / . is then displayed on adisplay device. Additional details of the Decoding module 316 is further described below.

[0067] FIG.4 is a schematic drawing illustrating a two-stream LSVR-based LVC framework 400 according to an embodiment of the present disclosure. Beginning on the sender side, the LSVR-based LVC framework 400 includes an Encoding module 402, a set of learned visualCodebooks ^^, … , ^^ 404, a Codeword Encoding module 406, and a Fidelity-PreservingEncoding module 408 on the sender side. In an embodiment, the sending device is provided withor obtains a set of n video frames A = / ^, … , / 6. In an embodiment, each video frame / 4 attime stamp t in the set of n video frames is a 3D tensor with shape ^^9 × ℎ^9 × ^ , where^^9 × ℎ^9 , ^ are the width, height, and number of channels of the video frame. For example,3 for color videos, c = 1 for spectral videos, or c = 4 for RGB-D (color and depth) videos.In a first stream of the LSVR-based LVC framework 400, the Encoding module 402 encodes eachinput frame / 4 into an embedding feature tensor ^4 with shape ^^9 × ℎ^9 × ^ representing thewidth, height, and d is the number of feature channels. In an embodiment, the width ^^9and height ℎ^9depend on the width and height of video frame / 4as well as the network structure of the Encoding module 402. The Encoding module 402 may be implemented using one or moreneural networks. The embedding feature tensor ^4 is then mapped to a set of codeword indices^4,^, … , ^4,^ , using a set of learned visual Codebooks ^^, … , ^^ 404, respectively. In anembodiment, the visual Codebook ^^ comprises ^^ codewords, each codeword being a d-dimension feature vector. In an embodiment, each index ^^,4,^ in ^^,4 (U = 1, … , ^^9 × ℎ^9 )corresponds to a codeword ^^,4,^ ∈ ^^ that is nearest to the corresponding latent feature vector^4,^ of the l-th “super-pixel” in the embedding feature tensor ^4:^^,4,^ = VEWLN$^%∈ &' XNYZ+^,, ^4,^-,Atty. Docket No.4502-82901 (6000680PCT02)

[0068] where XNYZ+- is a distance metric such as L1 or L2 norm. In an embodiment, for t >1, the sending device is also provided with a set of codeword indices ^48^,^, … , ^48^,^ for theprevious video frame / 48^ , and a set of integer different codeword indices ∆^4,^, … , ∆^4,^ iscomputed. In an embodiment, each entry ∆^4,^,^ in ∆^4,^ is defined as: ∆^4,^,^ = ^4,^,^ if^4,^,^ ≠ ^48^,^,^ , and ∆^4,^,^ = −1 otherwise. For time stamp t=0, ∆^4,^ = ^4,^ . TheCodeword Encoding module 306 further encodes the set of integer different codeword indices∆^4,^, … , ∆^4,^ into a compressed codeword string ^^^^^,4 in a lossless way, which can beefficiently transmitted to the receiver side.

[0069] In some embodiments, for each codebook ^^, the maximum number of bits needed totransfer codeword indices ^^ in FIG. 3 or ^4,^ in FIG. 4, without any further processing, is[U\\E+log^ ^^- for each “super-pixel” in ^ in FIG. 3 or ^4 in FIG. 4. In some embodiments,the number of bits can be further reduced to save bit consumption of the whole system. In some embodiments, an effective arithmetic coding (AC) method that can losslessly compress the integer indices is used. For example, for natural images, codewords normally show up with different frequencies. For instance, codewords of natural scenes may be used more frequently than those of human faces. In an embodiment, by assigning less bits to more frequently used indices, the total bitrate can be largely reduced.

[0070] In some embodiments, the frequency of codewords’ usage in a large training set is calculated. Then the codewords are reordered in descending order of the occurring frequency. For example, the most frequent occurring codeword with an original index of ix will be reassigned to have index 1. For instance, let ^^̂ or ^4̂,^be the integer sequence using the reordered codeword indices for ^^or ^4,^. Then for each particular ^^̂ or ^4̂,^, the reordered indices are further converted and rescaled as: −,^a^ , N[ N / NY ,^a^^ \^^ ^ , N[ N / NY \^^

[0071] Such operations transform the indices distribution to a Gaussian style bell shape, whichAtty. Docket No.4502-82901 (6000680PCT02) can be efficiently encoded by Gaussian Mixture-based arithmetic coding.

[0072] Additionally, in a second stream of the LSVR-based LVC framework 400, the inputvideo frames A = / ^, … , / 6 is fed into the Fidelity-Preserving Encoding module 408 to computea fidelity-preserving string ^^,^, … , ^^,6, which is transmitted to the receiver side. As describedabove, the fidelity- string ^^,^, … , ^^,6 carries fine detail image information forpreserving the fidelity of the video frame / 4. In an embodiment, each fidelity-preserving string^^,4 corresponds to each frame / 4, which has a low bitrate that is comparable with the low bitrateof the codeword string ^^^^^,4. In some embodiments, an LVC encoder or a conventional video compression encoder designed to optimize the pixel-level fidelity between the decompressed frames and the original input frames (e.g., VVC) is used as the Fidelity-Preserving Encoding module 408. In another embodiment, an LIC encoder or a conventional image compression encoder that encodes each frame as an individual image (e.g., JPEG) is used as the Fidelity- Preserving Encoding module 408. In some embodiments, to achieve a very low bitrate for ^^,4,the input frames A = / ^, … , / 6 are first downsampled, and then encoded with either an LVCencoder or the conventional video compression encoder or an LIC encoder or a conventional image compression encoder.

[0073] Referring now to the receiver side of the LSVR-based LIC framework 400, the receiver / receiving device includes a Codeword Decoding module 412, the set of learned visualCodebooks ^^, … , ^^ 404, a Decoding module 416, a Fidelity-Preserving Decoding module 418,and an LQ Embedding module 420. After receiving the codeword string ^^^^^,4, the CodewordDecoding module 412 recovers the integer different codeword indices ∆^4,^, … , ∆^4,^ from thereceived codeword string ^^^^^,4. In an embodiment, for t > 1, using the previously decodedcodeword indices ^48^,^, … , ^48^,^ , the codeword indices for the current frame ^4,^, … , ^4,^ isrecovered. Again, for time stamp t=0, ^4,^ = ∆^4,^ . Then, using the same visual Codebooks^^, … , ^^ 404 that was used on the sender side, a set of decoded feature tensors ^.4,^, … , ^.4,^ isdetermined based on the codeword indices ^4,^, … , ^4,^ . In an embodiment, for the l-th “super-pixel”, the feature vector ^.4,^,^in the decoded feature tensor ^.4,^is the codeword ^4,^,^feature.Atty. Docket No.4502-82901 (6000680PCT02)

[0074] Additionally, the Fidelity-Preserving Decoding module 418 computes an LQ substitute input / 014 based on the received fidelity-preserving string ^^,4. In some embodiments, when the original input / 4is downsampled in the sender, the substitute input / 014 is also upsampled to be the original size same as the original input / 4. In some embodiments, the Fidelity-Preserving Decoding module 418 uses a decoder part of an LVC or a conventional video compression method, or the previous LIC or conventional image compression method, corresponding to the Fidelity- Preserving Encoding module 408. The LQ substitute input / 014 is further fed into the LQ Embedding module 420 to compute an LQ embedding feature tensor ^014 . The Decoding module 416 reconstructs an output image / .4based on the LQtensor ^014 andthe decoded feature tensors ^.4,^, … , ^.4,^ . In some embodiments, the output image / .4 is thendisplayed on a display device.

[0075] FIG. 5 and FIG. 6 are a schematic diagrams illustrating a detailed workflow of a Decoding module in accordance with an embodiment. In particular, FIG. 5 is a schematic drawing illustrating a LSVR-based LIC decoding pipeline 500 according to an embodiment of the present disclosure. In an embodiment, the Decoding module 316 in FIG. 3 implements the LSVR-based LIC decoding pipeline 500. As shown in FIG.5, for LSVR-based LIC, when K > 1, a final decoded feature tensor ^. is first computed in a Feature Fusion module 502 based on thedecoded feature tensors ^.^, … , ^.^ . In an embodiment, ^. is a weighted combination of^.^, … , ^.^:^. = ∑^^c^ ^^^.^ ,

[0076] where a weight ^^is assigned to the codebook ^^corresponding to the decodedfeature tensor ^.^ . When K=1, ^. = ^.^. When K > 1, the weights ^^, … , ^^ are computed bya Weight Prediction module 504 based on the LQ embedding feature tensor ^01. The LQ embedding feature tensor ^01is then used as a control signal to guide the conditioned generative reconstruction process based on the final decoded feature tensor ^. in a Reconstruction moduleAtty. Docket No.4502-82901 (6000680PCT02) 506. The reconstruction module 506 then reconstructs the output / . based on the LQ embeddingfeature tensor ^01 and the decoded feature tensors ^.^, … , ^.^ .

[0077] FIG. 6 is a schematic drawing illustrating a LSVR-based LVC decoding pipeline 600 according to an embodiment of the present disclosure. In an embodiment, the Decoding module 416 in FIG.4 implements the LSVR-based LVC decoding pipeline 600. As shown in FIG.6, for LSVR-based LVC, in an embodiment, when K > 1, a final decoded feature tensor ^.4is firstcomputed in a Feature Fusion module 602 based on the decoded feature tensors ^.4,^, … , ^.4,^:^.4 = ∑^^c^ ^^^.4,^ .

[0078] For K=1, ^.4 = ^.4,^. When K > 1, the weights ^^, … , ^^ are computed by a WeightPrediction module 604 based on the LQ embedding feature tensor ^014 . The Weight Prediction module 604 can have any network structure including convolutional NN or transformer structures.

[0079] The LQ embedding feature tensor ^014 is then used as a control signal to guide the conditioned generative reconstruction process based on the final decoded feature tensor ^.4in a Reconstruction module 606.

[0080] There are multiple ways to use ^01in FIG. 5 (or ^014 in FIG. 6) to guide the reconstruction process. In some embodiments, the Reconstruction module 506 / 606 has a network structure of multiple CNN layers like the decoding network of a VAE. Then ^01tunes ^. in FIG. 5 (or ^014 tunes ^.4in FIG. 6) through an affine transformation and generates a new tuned feature ^.∗(or ^.4∗): ^.∗ = ^. + e+f^. + g- or ^.∗4 = ^.4 + e+f^.4 + g-,

[0081] where e is a pre-set guidance strength and the affine parameters f, g are given by: f, g = h:^\$+^., ^01-; or f, g = h i^\$:^. 014 , ^4 ;j,

[0082] where ^\$+- is the concatenation operation. In some embodiments, this affineAtty. Docket No.4502-82901 (6000680PCT02) transformation is implemented by an NN with the set of parameters h that are determined through training.

[0083] The network structure of the Reconstruction module 506 in FIG. 5 for LIC and the Reconstruction module 606 in FIG. 6 for LVC can be the same or different. In some embodiments, the Reconstruction module 606 in FIG. 6 can take previous reconstructed frames / . , … , / . or th 01 0148k 48^ e previous LQ embedding feature tensors ^48k ,…, ^48^ ( l ≥ 1- asto help reconstruct the current / .4. For, … , / .48^ and the initialoutput / .4can be stacked together into a 3D tensor and go through a 3Dnetwork to compute an updated / .4. Alternatively, ^0148k,…, ^0148^can be concatenated with the initial ^014 to go through an additional convolutional network to compute an updated ^014 . Such operations may improve the smoothness of the generated video.

[0084] FIG. 7 is a flowchart illustrating a LSVR-based LIC method 700 implemented by a sending device according to an embodiment of the present disclosure. The LSVR-based LIC method 700 begins at step 702 by encoding an input image x into an embedding feature tensor y. In an embodiment, the input image x is a three-dimensional (3D) tensor having a first shape^^ × ℎ^ × ^, wherein ^^ ,  ℎ^ ,  V$^ c are a first width, a first height, and number of channels ofthe input image x respectively. In an embodiment, the embedding feature tensor ^ has a secondshape ^^ × ℎ^ × ^, wherein ^^ ,  ℎ^ ,  V$^ d  are a second width, a second height, and number offeature channels respectively, wherein the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input image. In anembodiment, encoding the set of integer codeword indices ^^, … , ^^ into the compressedcodeword string ^^^^^comprises calculating a frequency of codewords usage in a large training set; reordering the codewords in descending order based on the frequency to obtain a reorderedcodeword indices; reordering the set of integer codeword indices ^^, … , ^^ based on the reorderedcodeword indices to obtain a reordered set of integer codeword indices; converting the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codewordAtty. Docket No.4502-82901 (6000680PCT02) indices to a Gaussian style bell shape; and encoding, using a Gaussian Mixture-based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^.

[0085] At step 704, the sending device obtains, using a set of learned visual Codebooks^^, … , ^^ , a set of integer codeword indices ^^, … , ^^ based on embedding feature tensor y. Inan embodiment, visual Codebook ^^comprises ^^codewords, and each codeword being a d-dimension feature vector. In an embodiment, each index ^^,^ in ^^ (l = 1, … , ^^ × ℎ^) in theset of integer codeword indices ^^, … , ^^ corresponds to a codeword ^^,^ ∈ ^^ that is nearest toa corresponding latent feature vector ^^of an l-th super-pixel in the embedding feature tensor y,wherein ^^,^ = argmi$^%∈ &' Dist+^,, ^^- , and wherein Dist+^,, ^^- is a distance metric. In anembodiment, the distance metric is L1 or L2 norm.

[0086] The sending device, at step 706, encodes the set of integer codeword indices ^^, … , ^^into a compressed codeword string ^^^^^. At step 708, the sending device determines a fidelity- preserving compressed string ^^based on the input image x. In an embodiment, determining the fidelity-preserving compressed string ^^based on the input image x comprises downsampling the input image x to obtain a downsampled image; and determining the fidelity-preserving compressed string ^^based on the downsampled image. The sending device, at step 710,transmits the compressed codeword string ^^^^^ and the fidelity-preserving compressed string^^ toward a receiving device.

[0087] FIG. 8 is a flowchart illustrating a LSVR-based LIC method 800 implemented by a receiving device according to an embodiment of the present disclosure. The LSVR-based LIC method 800 begins at step 802 by receiving a compressed codeword string ^^^^^and a compressed fidelity-preserving string ^^corresponding to an original input image x. At step804, the receiving device determines a set of integer codeword indices ^^, … , ^^ based on thecompressed codeword string ^^^^^. The receiving device, at step 806, determines, using a set oflearned visual Codebooks ^^, … , ^^ , a set of decoded feature tensors ^.^, … , ^.^ based on the setof integer codeword indices ^^, … , ^^ , wherein the compressed codeword string ^^^^^ is basedon the set of learned visual Codebooks ^^, … , ^^ . In some embodiments, for an l-th super-pixel,Atty. Docket No.4502-82901 (6000680PCT02) feature vector ^.^,^in decoded feature tensor ^.^ is a feature vector of codeword ^^,^.

[0088] At step 808, the receiving device determines a low-quality (LQ) substitute input / 01based on the compressed fidelity-preserving string ^^, wherein the LQ substitute input / 01comprises fidelity information of the original input image x. In some embodiments, the receiving device performs upsampling, prior to determining the LQ embedding feature tensor ^01, the LQ substitute input / 01to an original size of the original input image x.

[0089] The receiving device, at step 810, determines an LQ embedding feature tensor ^01based on the LQ substitute input / 01. At step 812, the receiving device generates an output image / . based on the LQ embedding feature tensor ^01and the set of decoded feature tensors^.^, … , ^.^ . In an embodiment, the receiving device determines weights ^^, … , ^^ based on theLQ embedding feature tensor ^01; determines a decoded feature tensor ^. based on the set ofdecoded feature tensors ^.^, … , ^.^ and the weights ^^, … , ^^; and generates the output image / .based on the LQ embedding feature tensor ^01and the decoded feature tensor ^.. In some embodiments, the receiving device tunes the decoded feature tensor ^. based on the LQ embedding feature tensor ^01through an affine transformation to generates a tuned feature ^.∗; and generates the output image / . based on the LQ embedding feature tensor ^01and the tuned feature ^.∗. In some embodiments, the affine transformation is implemented by a neural network with a set of parameters θ determined through training.

[0090] FIG. 9 is a flowchart illustrating a LSVR-based LVC method 900 implemented by a sending device according to an embodiment of the present disclosure. The LSVR-based LVC method 900 begins at step 902 by encoding an input video frame / 4into an embedding feature tensor ^4, wherein the input video frame / 4is a video frame at time stamp t in a set of n videoframes X = / ^, … , / 6. In an embodiment, the input video frame / 4 is a three-dimensional (3D)tensor having a first shape ^^9 × ℎ^9 × ^, where ^^9 × ℎ^9 ,  c are a first width, a first height, andnumber of channels of the input video frame / 4respectively. In an embodiment, the embeddingfeature tensor ^4 has a second shape ^^9 × ℎ^9 × ^, wherein ^^9 , ℎ^9 ,  d  are a second width, aAtty. Docket No.4502-82901 (6000680PCT02) second height, and number of feature channels respectively. In some embodiments, the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input video frame / 4.

[0091] At step 904, the sending device obtains, using a set of learned visual Codebooks^^, … , ^^ , a first set of codeword indices ^4,^, … , ^4,^ based on the embedding feature tensor ^4.In an embodiment, the visual Codebook ^^comprises ^^codewords, and each codeword beinga d-dimension feature vector. In an embodiment, each index ^^,4,^ in ^^,4 (l = 1, … , ^^9 × ℎ^9)in the first set of codeword indices ^4,^, … , ^4,^ corresponds to a codeword ^^,4,^ ∈ ^^ that isnearest to a corresponding latent feature vector ^4,^of an l-th super-pixel in the embedding featuretensor ^4 , wherein ^^,4,^ = argmi$^%∈ &' Dist:^,, ^4,^; , and wherein Dist:^,, ^4,^; is a distancemetric such as L1 or L2 norm. The sending device, at step 906, determines a second set ofcodeword indices ∆^4,^, … , ∆^4,^ based on the first set of codeword indices ^4,^, … , ^4,^ and athird set of codeword indices ^48^,^, … , ^48^,^ corresponding to a previous video frame / 48^ inthe set of n video frames X = / ^, … ,

[0092] At step 908, the sending device encodes the second set of codeword indices∆^4,^, … , ∆^4,^ into a compressed codeword string ^^^^^,4. In some embodiments, the sendingdevice calculates a frequency of codewords usage in a large training set; reorders the codewords in descending order based on the frequency to obtain a reordered codeword indices; reorders thefirst set of codeword indices ^4,^, … , ^4,^ based on the reordered codeword indices to obtain areordered set of integer codeword indices; converts the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codeword indices to a Gaussian style bell shape; and encodes, using a Gaussian Mixture-based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^,4. In some embodiments, eachentry ∆^4,^,^ in ∆^4,^ in the second set of codeword indices ∆^4,^, … , ∆^4,^ is defined as:∆^4,^,^ = ^4,^,^ if ^4,^,^ ≠ ^48^,^,^, otherwise ∆^4,^,^ = −1, wherein for time stamp t=0, ∆^4,^ =^4,^.

[0093] The sending device, at step 910, determines a fidelity-preserving compressed stringAtty. Docket No.4502-82901 (6000680PCT02) ^^,4based on the input video frame / 4. In an embodiment, the sending device downsamples the input video frame / 4to obtain a downsampled video frame; and determines the fidelity- preserving compressed string ^^,4 based on the downsampled video frame. The sending device, at step 912, transmits the compressed codeword string ^^^^^,4and the fidelity-preserving compressed string ^^,4toward a receiving device.

[0094] FIG.10 is a flowchart illustrating a LSVR-based LVC method 1000 implemented by a receiving device according to an embodiment of the present disclosure. The LSVR-based LVC method 1000 begins at step 1002 by receives a compressed codeword string ^^^^^,4and a compressed fidelity-preserving string ^^,4corresponding to a video frame / 4at time stamp t ina set of n video frames X = / ^, … , / 6. At step 1004, the receiving device determines a first setof integer codeword indices ∆^4,^, … , ∆^4,^ based on the compressed codeword string ^^^^^,4.

[0095] The receiving device, at step 1006, determines a second set of integer codeword indices^4,^, … , ^4,^ based on the first set of integer codeword indices ∆^4,^, … , ∆^4,^ and a third set ofinteger codeword indices ^48^,^, … , ^48^,^ corresponding to a previous video frame / 48^ in theset of n video frames X = / ^, … , / 6. In an embodiment, for time stamp t=0, ^4,^ = ∆^4,^ .

[0096] At step 1008, the receiving device determines, using a set of learned visual Codebooks^^, … , ^^ respectively, a set of decoded feature tensors ^.4,^, … , ^.4,^ based on the second set ofinteger codeword indices ^4,^, … , ^4,^ , wherein the compressed codeword string ^^^^^,4 is basedon the set of learned visual Codebooks ^^, … , ^^ , wherein for an l-th super-pixel, feature vector^.4,^,^ in decoded feature tensor ^.4,^  is a feature vector of codeword ^4,^,^.The receiving device, at step 1010, determines a low-quality (LQ) substitute input / 014 based on the compressed fidelity-preserving string ^^,4, wherein the LQ substitute input / 014 comprises fidelity information of the video frame / 4. In some embodiments, the receiving device performs upsampling, prior to determining the LQ substitute input / 014 , the LQ substitute input / 014 to an original size of the video frame / 4. At step 1012, the receiving device determines an LQ embedding feature tensor ^01 014 based on the LQ substitute input / 4.Atty. Docket No.4502-82901 (6000680PCT02)

[0098] The receiving device, at step 1014, generates an output image / .4based on the LQembedding feature tensor ^014 and the set of decoded feature tensors ^.4,^, … , ^.4,^ . In someembodiments, the receiving device determines weights ^^, … , ^^ based on the LQ embeddingfeature tensor ^014 ; determines a decoded feature tensor ^.4based on the set of decoded featuretensors ^.4,^, and the weights ^^, … , ^^; and generates the output image / .4 based on theLQ embedding feature tensor ^014  and the decoded feature tensor ^.4.

[0099] FIG.11 is a diagram illustrating an apparatus 1100 according to an embodiment of the present disclosure. The apparatus 1100 can be used to implement embodiments of the present disclosure. For example, the apparatus 1100 may be configured to perform the functions of a sending device or a receiving device according to any of the embodiments of the present disclosure. The apparatus 1100 includes receiver units (RX) 1120 or receiving means for receiving data via ingress ports 1110. The apparatus 1100 also includes transmitter units (TX) 1140 or transmitting means for transmitting via data egress ports 1150. For example, on the sender side, the sending device may use the RX 1120 or receiving means to obtain an original image or video frames and then use the TX 1140 or transmitting means for transmitting a compressed codeword string and a fidelity-preserving compressed string to a receiving device as described in FIG. 3 and FIG. 4. The receiving device may use the RX 1120 or receiving means to receive the compressed codeword string and a fidelity-preserving compressed string and then use the TX 1140 or transmitting means for transmitting a decoded image of the original image (e.g., the reconstructed output image / .) to a display device or to another computing device.

[0100] The apparatus 1100 includes a memory 1160 or data storing means for storing the instructions and various data. The memory 1160 can be any type of, or combination of, memory components capable of storing data and / or instructions. For example, the memory 1160 can include volatile and / or non-volatile memory such as read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM). The memory 1160 can also include one or more disks, tape drives, and solid-Atty. Docket No.4502-82901 (6000680PCT02) state drives. In some embodiments, the memory 1160 can be used as an over-flow data storage device to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. In some embodiments, the memory 1160 can be memory that is integrated with the processor 1130.

[0101] The apparatus 1100 has one or more processors 1130 or other processing means (e.g., central processing unit (CPU)) to process instructions. The one or more processors 1130 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field- programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The one or more processors 1130 are communicatively coupled via a system bus with the ingress ports 1110, RX 1120, TX 1140, egress ports 1150, and memory 1160. The one or more processors 1130 can be configured to execute instructions stored in the memory 1160. As an example, in one embodiment, the memory 1160 stores an LSVR-based LIC / LVC module 1170. The LSVR-based LIC / LVC module 1170 includes data, executable instructions, and / or one more sub-modules for implementing the disclosed embodiments. Thus, the one or more processors 1130 provide a means for performing any computational, comparison, determination, initiation, configuration, or any other action corresponding to the claims when the appropriate instruction is executed by the processor 1130. Thus, the inclusion of the LSVR- Based LIC / LVC module 1170 substantially improves the functionality of the apparatus 1100.

[0102] The different modules in the disclosed embodiments can be trained altogether or piece by piece. This present disclosure does not place any restriction on the network architectures of various modules or the training methods of the modules. A non-limiting example of a training method may include the following stages:

[0103] Stage 1: The Encoding module, the Codebooks, and the Decoder are firstly trained based on a large set of image dataset. In this stage, the Reconstruction module use unconditioned generative process without using the LQ embedding feature tensor.

[0104] Stage 2: The Encoding module and the Codebooks from Stage 1 are kept fixed, and the LQ embedding module and the Decoder module are trained for the LSVR-based LIC frameworkAtty. Docket No.4502-82901 (6000680PCT02) in FIG.3 based on a large set of image dataset.

[0105] Stage 3: The Encoding module and the Codebooks from Stage 1 are kept fixed, and the LQ embedding module and the Decoder module are for the LSVR-based LVC framework in FIG. 4. If the same network structure is used for the Reconstruction module for both LSVR-based LIC and LSVR-based LVC, the LQ embedding module and the Decoder module can also be finetuned from the Stage 1 version.

[0106] A variety of training loss can be used, including pixel-level distortion metrics like Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) between the reconstruction output / . in FIG. 3 (or / .4in FIG. 4) and the original input / in FIG. 3 (or / 4in FIG. 4), and perceptual quality metrics like Learned Perceptual Image Patch Similarity (LPIPS) between the reconstruction output / . in FIG.3 (or / .4in FIG.4) and the original input / in FIG.3 (or / 4in FIG.4). Other losses like the adversarial Generative Adversarial Network (GAN) loss to improve the naturalism of the generated / . in FIG.3 (or / .4in FIG.4) can also be used.

[0107] The LSVR-based LIC framework 300 in FIG.3 and the LSVR-based LVC framework 400 in FIG.4 include the following novel features:

[0108] High compression rate with high-quality reconstruction at the same time by combining the power of LSVR-based generative modelling with high-perceptual quality and the fidelity- preserving guidance from fidelity-preserving controls. Compared to previous LIC framework based on encoding and transmitting real value latent features, the integer codewords indices are can be highly efficiently encoded and transmitted. Compared to previous LSVR-based compression method, the disclosed embodiments use conditional generation guided by fidelity- preserving controls. The fidelity-preserving controls are drawn from a highly compressed LQ substitute, which provides important guidance to largely improve the reconstruction fidelity with little transmission overhead. Further, by learning powerful visual codebooks from a large number of videos, and by sharing such visual codebooks between sender and receiver, videos can be compressed into integer codeword indices, which are highly robust and efficient to transmit. The receiver side can effectively reconstruct videos based on such indices. Since LSVR is quiteAtty. Docket No.4502-82901 (6000680PCT02) robust to input degradations and small input perturbations, for most video frames, a large portion of the codeword indices can remain the same. Only the different codeword indices need to be transferred. This makes LSVR-based LVC very effective. Different from conventional LVC pipeline, the disclosed LSVR-based LVC is designed for end-to-end LVC, and there is no error accumulation or inefficient computation. In addition, the disclosed LSVR-based LIC and LSVR- based LVC can share a similar processing pipeline, making it convenient for industrial product optimization.

[0109] Moreover, the disclosed embodiments provide high flexibility in accommodating different hardware and software platforms. Transferring integer indices is very robust to heterogeneous software and hardware platforms. By encoding codeword indices instead of latent features, LSVR-based compression does not suffer from the sensitivity of computation mismatch between the sender and receivers. Comparing with previous LSVR-based compression, the disclosed embodiments can accommodate various computing environments. For homogeneous computing platforms, the disclosed processing pipeline can be used to combine the strength of the fidelity cue from existing fidelity-preserving LIC or LVC and the perceptual cue from LSVR-based restoration. For heterogeneous computing platforms where previous LIC or LVC methods may have difficulty to apply, the disclosed embodiments can still provide a decent low-bitrate baseline reconstruction with good perceptual quality using LSVR-based restoration alone, or can pair with traditional image or video compression methods like JPEG or VVC for improved reconstruction.

[0110] Additionally, the disclosed embodiments provide robust HQ reconstruction against input degradation. For example, the disclosed embodiments can generate realistic and rich textures reconstructed images using HQ codebooks even for LQ inputs. Further, the disclosed embodiments provide flexible quality control. In contrast, previous LSVR-based compression does not have an efficient way for bitrate control. The disclosed embodiments provide this capability by tuning the bitrate of the LQ substitute. Also, by tuning the guidance strength in the Reconstruction module, the disclosed embodiments allow tuning the tradeoff between perceptual quality and fidelity. Thus, the disclosed embodiments provide the flexibility to balance the bitrate,Atty. Docket No.4502-82901 (6000680PCT02) perceptual quality and fidelity.

[0111] While several embodiments have been provided in the present disclosure, a person of ordinary skill in the art would understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the disclosure is not to be limited to the details given herein. For example, the various elements or components may be combined or integrated in another system or certain features may be omitted, or not implemented.

[0112] In addition, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled or directly coupled or communicating with each other may be indirectly coupled or communicating through some interface, device, or intermediate component whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

Atty. Docket No.4502-82901 (6000680PCT02) CLAIMS What is claimed is:

1. A Learned Sparse Visual Representation (LSVR)-based Learned Image Compression (LIC) method implemented by a sending device, the LSVR-based LIC method comprising: encoding an input image x into an embedding feature tensor y, obtaining, using a set of learned visual Codebooks ^^, … , ^^, a set of integer codewordindices ^^, … , ^^ based on embedding feature tensor y;the set of integer codeword indices ^^, … , ^^ into a compressed codeword string^^^^^;determining a fidelity-preserving compressed string ^^based on the input image x, and transmitting the compressed codeword string ^^^^^and the fidelity-preserving compressed string ^^toward a receiving device.

2. The LSVR-based LIC method according to claim 1, wherein the input image x is a three-dimensional (3D) tensor having a first shape ^^ × ℎ^ × ^, wherein ^^ ,  ℎ^ ,   c are a first width, afirst height, and number of channels of the input image x respectively.

3. The LSVR-based LIC method according to any of claims 1-2, wherein the embeddingfeature tensor ^ has a second shape ^^ × ℎ^ × ^, wherein ^^ ,  ℎ^ ,  d  are a second width, a secondheight, and number of feature channels respectively.

4. The LSVR-based LIC method according to claim 3, wherein the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input image x.

5. The LSVR-based LIC method according to any of claims 1-4, wherein visual Codebook ^^comprises ^^codewords, and each codeword being a d-dimension feature vector.Atty. Docket No.4502-82901 (6000680PCT02)6. The LSVR-based LIC method according to any of claims 1-5, wherein each index ^^,^ in^^ ( l = 1, … , ^^ × ℎ^ ) in the set of integer codeword indices ^^, … , ^^ corresponds to acodeword ^^,^ ∈ ^^ that is nearest to a corresponding latent feature vector ^^ of an l-th super-pixel in the embedding feature tensor y, wherein ^^,^ = argmi$^%∈ &' Dist+^,, ^^-, and whereinDist+^, , ^^- is a distance metric.

7. The LSVR-based LIC method according to claim 6, wherein the distance metric is L1 or L2 norm.

8. The LSVR-based LIC method according to any of claims 1-7, wherein determining the fidelity-preserving compressed string ^^based on the input image x comprises: downsampling the input image x to obtain a downsampled image; and determining the fidelity-preserving compressed string ^^based on the downsampled image.

9. The LSVR-based LIC method according to any of claims 1-8, wherein encoding the set ofinteger codeword indices ^^, … , ^^ into the compressed codeword string ^^^^^ comprises:calculating a frequency of codewords usage in a large training set; reordering the codewords in descending order based on the frequency to obtain a reordered codeword indices; reordering the set of integer codeword indices ^^, … , ^^ based on the reordered codewordindices to obtain a reordered set of integer codeword indices; converting the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codeword indices to a Gaussian style bell shape; and encoding, using a Gaussian Mixture-based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^.Atty. Docket No.4502-82901 (6000680PCT02) 10. A Learned Sparse Visual Representation (LSVR)-based Learned Image Compression (LIC) method implemented by a receiving device, the LSVR-based LIC method comprising: receiving a compressed codeword string ^^^^^and a compressed fidelity-preserving string ^^corresponding to an original input image x; determining a set of integer codeword indices ^^, … , ^^ based on the compressedcodeword string ^^^^^; determining, using a set of learned visual Codebooks ^^, … , ^^ , a set of decoded featuretensors ^.^, … , ^.^ based on the set of integer codeword indices ^^, … , ^^ , wherein the compressedcodeword string ^^^^^ is based on the set of learned visual Codebooks ^^, … , ^^;determining a low-quality (LQ) substitute input / 01based on the compressed fidelity- preserving string ^^, wherein the LQ substitute input / 01comprises fidelity information of the original input image x; determining an LQ embedding feature tensor ^01based on the LQ substitute input / 01; and generating an output image / . based on the LQ embedding feature tensor ^01and the setof decoded feature tensors ^.^, … , ^.^ .

11. The LSVR-based LIC method according to claim 10, further comprising upsampling, prior to determining the LQ embedding feature tensor ^01, the LQ substitute input / 01to an original size of the original input image x.

12. The LSVR-based LIC method according to any of claims 10-11, wherein for an l-th super- pixel, feature vector ^.^,^in decoded feature tensor ^.^ is a feature vector of codeword ^^,^.

13. The LSVR-based LIC method according to any of claims 10-12, wherein generating the output image / . based on the LQ embedding feature tensor ^01and the set of decoded featuretensors ^.^, … , ^.^ comprises:determining weights ^ 01^, … , ^^ based on the LQ embedding feature tensor ^ ;Atty. Docket No.4502-82901 (6000680PCT02) determining a decoded feature tensor ^. based on the set of decoded feature tensors^.^, … , ^.^ and the weights ^^, … , ^^; andgenerating the output image / . based on the LQ embedding feature tensor ^01and the decoded feature tensor ^..

14. The LSVR-based LIC method according to claim 13, wherein generating the output image / . based on the LQ embedding feature tensor ^01and the decoded feature tensor ^. comprises: tuning the decoded feature tensor ^. based on the LQ embedding feature tensor ^01through an affine transformation to generates a tuned feature ^.∗; and generating the output image / . based on the LQ embedding feature tensor ^01and the tuned feature ^.∗.

15. The LSVR-based LIC method according to claim 14, wherein the affine transformation is implemented by a neural network with a set of parameters θ determined through training.

16. A Learned Sparse Visual Representation (LSVR)-based Learned Video Compression (LVC) method implemented by a sending device, the LSVR-based LVC method comprising: encoding an input video frame / 4into an embedding feature tensor ^4, wherein the inputvideo frame / 4 is a video frame at time stamp t in a set of n video frames X = / ^, … , / 6;obtaining, using a set of learned visual Codebooks ^^, … , ^^ , a first set of codewordindices ^4,^, … , ^4,^ based on the embedding feature tensor ^4;determining a second set of codeword indices ∆^4,^, … , ∆^4,^ based on the first set ofcodeword indices ^4,^, … , ^4,^ and a third set of codeword indices ^48^,^, … , ^48^,^corresponding to a previous video frame / 48^ in the set of n video frames X, / 6;encoding the second set ofindices ∆^4,^, … , ∆^4,^ into a compressed codewordstring ^^^^^,4; determining a fidelity-preserving compressed string ^^,4based on the input video frame / 4, andAtty. Docket No.4502-82901 (6000680PCT02) transmitting the compressed codeword string ^^^^^,4and the fidelity-preserving compressed string ^^,4toward a receiving device.

17. The LSVR-based LVC method according to claim 16, wherein the input video frame / 4is a three-dimensional (3D) tensor having a first shape ^^9 × ℎ^9 × ^, where ^^9 × ℎ^9 ,  c are afirst width, a first height, and number of channels of the input video frame / 4respectively.

18. The LSVR-based LVC method according to any of claims 16-17, wherein the embeddingfeature tensor ^4 has a second shape ^^9 × ℎ^9 × ^, wherein ^^9 , ℎ^9 ,  d  are a second width, asecond height, and number of feature channels respectively.

19. The LSVR-based LVC method according to claim 18, wherein the second width and the second height depend on the first width, the first height, and a network structure used in encoding the input video frame / 4.

20. The LSVR-based LVC method according to any of claims 16-19, wherein visual Codebook ^^comprises ^^codewords, and each codeword being a d-dimension feature vector.

21. The LSVR-based LVC method according to any of claims 16-20, wherein each index ^^,4,^in ^^,4 (l = 1, … , ^^9 × ℎ^9) in the first set of codeword indices ^4,^, … , ^4,^acodeword ^^,4,^that is nearest to a corresponding latent feature vector ^4,^ of an l-th super-pixel in the embedding feature tensor ^4, wherein ^^,4,^ = argmi$^%∈ &' Dist:^, , ^4,^;, and whereinDist:^ , ,^;,^4 is a distance metric.

22. The LSVR-based LVC method according to claim 21, wherein the distance metric is L1 or L2 norm.Atty. Docket No.4502-82901 (6000680PCT02) 23. The LSVR-based LVC method according to any of claims 16-22, wherein determining the fidelity-preserving compressed string ^^,4based on the input video frame / 4comprises: downsampling the input video frame / 4to obtain a downsampled video frame; and determining the fidelity-preserving compressed string ^^,4  based on the downsampled video frame.

24. The LSVR-based LVC method according to any of claims 16-23, wherein encoding thefirst set of codeword indices ^4,^, … , ^4,^ into the compressed codeword string ^^^^^,4 comprises:calculating a frequency of codewords usage in a large training set; reordering the codewords in descending order based on the frequency to obtain a reordered codeword indices; reordering the first set of codeword indices ^4,^, … , ^4,^ based on the reordered codewordindices to obtain a reordered set of integer codeword indices; converting the reordered set of integer codeword indices to transform a distribution of the reordered set of integer codeword indices to a Gaussian style bell shape; and encoding, using a Gaussian Mixture-based arithmetic coding, the reordered set of integer codeword indices into the compressed codeword string ^^^^^,4.

25. The LSVR-based LVC method according to any of claims 16-24, wherein each entry∆^4,^,^ in ∆^4,^ in the second set of codeword indices ∆^4,^, … , ∆^4,^ is defined as: ∆^4,^,^ =^4,^,^ if ^4,^,^ ≠ ^48^,^,^, otherwise ∆^4,^,^ = −1, wherein for time stamp t=0, ∆^4,^ = ^4,^ .

26. A Learned Sparse Visual Representation (LSVR)-based Learned Image Compression (LVC) method implemented by a receiving device, the LSVR-based LVC method comprising: receiving a compressed codeword string ^^^^^,4and a compressed fidelity-preservingstring ^^,4 corresponding to a video frame / 4 at time stamp t in a set of n video frames X = / ^, … , / 6;etermining a first set of integer codeword indices ∆^4,^, … , ∆^4,^ based on theAtty. Docket No.4502-82901 (6000680PCT02) compressed codeword string ^^^^^,4; determining a second set of integer codeword indices ^4,^, … , ^4,^ based on the first set ofinteger codeword indices ∆^4,^, … , ∆^4,^ and a third set of integer codewordindices ^48^,^, … , ^48^,^ corresponding to a previous video frame / 48^ in the set of n video framesX = / ^, …determining, using a set of learned visual Codebooks ^^, … , ^^ respectively, a set ofdecoded feature tensors ^.4,^, … , ^.4,^ based on the second set of integer codeword indices^4,^, … , ^4,^ , wherein the compressed codeword string ^^^^^,4 is based on the set of learned visualCodebooks ^^, … , ^^;determining a low-quality (LQ) substitute input / 014 based on the compressed fidelity- preserving string ^^,4, wherein the LQ substitute input / 014 comprises fidelity information of the video frame / 4; determining an LQ embedding feature tensor ^01 014 based on the LQ substitute input / 4; and generating an output image / .4based on the LQ embedding feature tensor ^014 and theset of decoded feature tensors ^.4,^, … , ^.4,^ .

27. The LSVR-based LVC method according to claim 26, wherein for time stamp t=0, ^4,^=∆^4,^.

28. The LSVR-based LVC method according to any of claims 26-27, further comprising upsampling, prior to determining the LQ substitute input / 01 014 , the LQ substitute input / 4to an original size of the video frame / 4.

29. The LSVR-based LVC method according to any of claims 26-28, wherein for an l-th super- pixel, feature vector ^.4,^,^in decoded feature tensor ^.4,^ is a feature vector of codeword ^4,^,^.Atty. Docket No.4502-82901 (6000680PCT02) 30. The LSVR-based LVC method according to any of claims 26-29, wherein generating the output image / .4based on the LQ embedding feature tensor ^014 and the set of decoded featuretensors ^.4,^, … , ^.4,^ comprises:determining weights ^ 01^, … , ^^ based on the LQ embedding feature tensor ^4 ;determining a decoded feature tensor ^.4 based on the set oftensors^.4,^, … , ^.4,^ and the weights ^^, … , ^^; andgenerating the output image / .4based on the LQ embedding feature tensor ^014  and the decoded feature tensor ^.

4.

31. The LSVR-based LVC method according to any of claims 26-30, wherein generating the output image / .4based on the LQ embedding feature tensor ^014 and the decoded feature tensor ^.4comprises: tuning the decoded feature tensor ^.4based on the LQ embedding feature tensor ^014 through an affine transformation to generates a tuned feature ^.4∗; and generating the output image / .4based on the LQ embedding feature tensor ^014 and the tuned feature ^.4∗.

32. The LSVR-based LVC method according to claim 31, wherein the affine transformation is implemented by a neural network with a set of parameters θ determined through training.

33. An apparatus comprising: a memory or storage means configured to store instructions; and one or more processors or processing means coupled to the memory or the storage means and configured to execute the instructions to cause the apparatus to perform a method according to any of claims 1-32.Atty. Docket No.4502-82901 (6000680PCT02) 34. A computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium, the computer-executable instructions when executed by one or more processors of an apparatus, cause the apparatus to perform a method according to any of claims 1-32.