Image coding and decoding method and device, equipment and storage medium
By using a latent space generative encoding and decoding framework, and combining feature transformation models and generative models with a cross-attention mechanism, the problem of insufficient image encoding and decoding compression performance in existing technologies is solved, and efficient image encoding and decoding effects are achieved.
Patent Information
- Application Number
- CN202410773139.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-12-16
AI Technical Summary
Existing image encoding and decoding technologies struggle to further improve compression performance while ensuring transmission quality. Conventional methods are nearing their performance ceiling, and generation models suffer from limitations in bit rate and computing power.
A latent space-based generative encoding and decoding framework is adopted. Intermediate features are generated through a feature transformation model, and a second latent representation is generated using a generative model to achieve image encoding and decoding. The reconstruction quality is improved by combining a cross-attention mechanism and a prediction model.
Achieving high visual quality and high visual fidelity image reconstruction at extremely low bit rates improves the efficiency and performance of image encoding and decoding, enhances the generative model's ability to capture spatial information, and improves the quality of reconstructed images.
Smart Images

Figure CN121151561A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device and storage medium for image coding. BACKGROUND
[0002] With the continuous development of Internet technology, the increasing demand for image transmission has brought great challenges to network bandwidth. In addition, in recent years, the picture and video content generated based on machine learning technology has grown rapidly, which poses new challenges to the transmission system. In various transmission scenarios, lossy transmission of images plays an indispensable role. Lossy transmission of images is performed by a certain degree of lossy encoding compression at the encoding end, and further decoding at the decoding end to reconstruct the image. Although this can reduce the resource overhead of images in the storage and transmission process, how to further improve the image compression performance while ensuring the quality of image transmission is still a problem to be solved. SUMMARY
[0003] In a first aspect of the present disclosure, a method for image coding is provided. In the method, for conversion between a target image and a bitstream of the target image, at least one intermediate feature is generated based on a first latent representation for the target image by using a feature conversion model. The first latent representation is indicated in the bitstream. Further, a second latent representation for the target image is generated based on the at least one intermediate feature by using a generative model, and the conversion is performed based on the second latent representation.
[0004] In a second aspect of the present disclosure, an apparatus for image coding is provided. The apparatus includes a feature generation module, a latent representation generation module, and a conversion execution module. The feature generation module is configured to: for conversion between a target image and a bitstream of the target image, generate at least one intermediate feature based on a first latent representation for the target image by using a feature conversion model. The first latent representation is indicated in the bitstream. The latent representation generation module is configured to: generate a second latent representation for the target image based on the at least one intermediate feature by using a generative model. The conversion execution module is configured to: perform the conversion based on the second latent representation.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit cause the electronic device to perform the method according to the first aspect of the present disclosure.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon instructions which, when executed by a processor, cause the processor to implement the method according to the first aspect of the present disclosure.
[0007] It should be understood that nothing in the Summary is to be construed as a limitation on the scope of the embodiments of the present disclosure. Other features of the present disclosure will be apparent from the following description, along with the associated drawings. BRIEF DESCRIPTION OF DRAWINGS
[0008] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, wherein:
[0009] Figure 1 A block diagram of an example image coding system that can utilize techniques of the present disclosure is shown;
[0010] Figure 2 An example coding framework according to some embodiments of the present disclosure is shown;
[0011] Figure 3 A schematic diagram of a network structure of a feature transformation model according to some embodiments of the present disclosure is shown;
[0012] Figure 4 A schematic diagram for enhancing local structure of a second latent representation for a target image according to some embodiments of the present disclosure is shown;
[0013] Figure 5 A flowchart of a method for image coding according to some embodiments of the present disclosure is shown;
[0014] Figure 6 A block diagram of an example apparatus for image coding according to some embodiments of the present disclosure is shown; and
[0015] Figure 7 A block diagram of a device in which one or more embodiments of the present disclosure can be implemented is shown. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.
[0017] In the description of embodiments of the disclosure, the term "comprising" and similar terms thereof are to be understood as open-ended, i.e., "including but not limited to". The term "based on" is to be understood as "based at least in part on". The term "one embodiment" or "the embodiment" is to be understood as "at least one embodiment". The term "some embodiments" is to be understood as "at least some embodiments". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the association between various data. For example, the above-mentioned association can be obtained based on various technical solutions known at present and / or to be developed in the future.
[0018] The term "in response to" used herein represents the state in which the corresponding event occurs or the condition is met. It will be understood that the timing of the execution of the subsequent action performed in response to the event or condition is not necessarily strongly associated with the time of occurrence of the event or the condition. For example, in some cases, the subsequent action can be performed immediately after the event occurs or the condition is met; in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is met.
[0019] It can be understood that the data involved in the technical solutions of the present technology (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0020] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0021] For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0022] As an optional but non-limiting embodiment, in response to receiving the active request of the user, the way of sending the prompt information to the user, for example, can be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0023] It can be understood that the above notification and user authorization process is only illustrative and does not limit the embodiments of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the embodiments of the present disclosure.
[0024] As used herein, the term “machine learning model” can learn the association between respective inputs and outputs from training data, so that after training is completed, the corresponding output can be generated for a given input. The generation of a machine learning model can be based on deep learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. Neural network models are one example of models based on deep learning. In the context of the present disclosure, “machine learning model” can also be referred to as “model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably herein.
[0025] As briefly mentioned above, lossy transmission of images is increasingly widely applied. It is desirable to further improve the image compression performance while ensuring the quality of image transmission. Conventional image and video compression algorithms usually utilize the visual and statistical characteristics of image data to design algorithms to improve compression efficiency according to the Shannon rate-distortion optimization theory. However, with the continuous advancement of traditional coding standardization and the full exploitation of performance, the conventional video coding framework has gradually approached the performance ceiling.
[0026] In recent years, the rapid development of generative modeling techniques has promoted the exploration of new coding methods based on generative models. In one of the existing schemes, a generative model is used to enhance the reconstructed image conditioned on the reconstructed image of a conventional end-to-end codec. Since the generative model is only used to enhance the reconstructed image, the code rate in this scheme is still relatively high due to the performance limitation of the conventional end-to-end codec. In another of the existing schemes, a generative model is used to reconstruct an image in the pixel domain. Since the feature representation of the image in the pixel domain still has a large amount of redundant information, the code rate in this scheme is also still relatively high and requires high computing power.
[0027] To this end, embodiments of the present disclosure propose a generative coding framework based on latent space. Specifically, embodiments according to the present disclosure propose a scheme for image coding. In the scheme, for conversion between a target image and a bitstream of the target image, at least one intermediate feature is generated based on a first latent representation for the target image using a feature conversion model. The first latent representation is indicated in the bitstream. Further, a second latent representation for the target image is generated based on the at least one intermediate feature using a generative model, and the conversion is performed based on the second latent representation.
[0028] It will be appreciated by those skilled in the art that, in accordance with embodiments of the present disclosure, a compressed domain latent representation of an image is transmitted in a bitstream, and the compressed domain latent representation is converted into intermediate features for generating a generative domain latent representation for the image with a generative model to perform image coding. In this way, on one hand, a general and compact compressed domain representation for the image can be obtained by means of the latent representation, thereby effectively reducing the code rate. On the other hand, the compressed domain latent representation can be converted into intermediate features that are adapted to the generative model by means of a feature conversion model for reconstructing the image with the generative model. This can enable efficient alignment of the external compressed domain information with the internal prior knowledge of the generative model, thereby enabling image reconstruction with high visual quality and high visual fidelity at an extremely low code rate. Therefore, the efficiency and performance of image coding can be effectively improved.
[0029] Various example implementations of the scheme will be described in detail below with reference to the accompanying drawings. First, refer to Figure 1 which shows a block diagram of an example image coding system 100 that can utilize the techniques of this disclosure. As shown, the image coding system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as an image encoding device, and the destination device 120 can also be referred to as an image decoding device. In operation, the source device 110 can be configured to generate encoded image data, and the destination device 120 can be configured to decode the encoded image data generated by the source device 110. The source device 110 can include an image source 112, an image encoder 114, and an input / output (I / O) interface 116.
[0030] The image source 112 can include a source such as an image capture device. Examples of the image capture device include, but are not limited to, an interface that receives image data from an image content provider, a computer graphics system for generating image data, and / or a combination thereof.
[0031] The image data can include one or more pictures, one or more frames of a video, and so on. The image encoder 114 encodes the image data from the image source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms an encoded representation of the image data. The bitstream can also include other data associated with image coding. The encoded image is an encoded representation of the image. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded image data can be transmitted directly to the destination device 120 via the I / O interface 116 over a network 130A. The encoded image data can also be stored on a storage medium / server 130B for access by the destination device 120.
[0032] Destination device 120 can include I / O interface 126, image decoder 124, and display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded image data from source device 110 or storage medium / server 130B. Image decoder 124 can decode the encoded image data. Display device 122 can display the decoded image data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 configured to interface with an external display device.
[0033] It is to be understood that the structure and function of example image coding system 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.
[0034] Example coding framework
[0035] Figure 2 An example coding framework 200 is shown in accordance with some embodiments of the present disclosure. It is to be understood that Figure 2 The shown example coding framework 200 is merely exemplary, and a coding framework in accordance with embodiments of the present disclosure can also include additional elements not shown, and / or can omit certain element(s) shown. The scope of the present disclosure is not limited in this regard.
[0036] As Figure 2 shown, in the encoding process, encoding model 220 generates a first latent representation for target image 210 based on target image 210 (denoted as y in Figure 2 In the context of the present disclosure, a latent representation for an image is a feature representation of the image in a latent variable domain (i.e., latent space) rather than a feature representation of the image in a pixel domain. A feature representation (which can also be referred to simply as a feature) for an image can be an intermediate representation of the image in the process of a machine learning model. In one example, a feature representation can be implemented in the form of a tensor. In another example, a feature representation can be implemented in the form of a vector. It is to be understood that a feature representation can also be implemented in any other suitable form, and the scope of the present disclosure is not limited in this regard. Furthermore, a feature representation can also be associated with a feature domain (which can also be referred to as a feature space). Exemplarily, multiple feature representations belonging to the same feature domain can have the same characteristics, and serve the same optimization objective. For example, the first latent representation y belongs to a compression domain, and can also be referred to as a compression domain latent representation.
[0037] In some embodiments, the target image 210 can be generated based on information such as text, for example by means of a machine learning model, or hand-drawing, or the like. Alternatively, the target image 210 can be a picture. In some other embodiments, the target image 210 can be a frame in a video, or a portion of a video frame. It should be appreciated that the target image 210 can also be any other suitable visual content data. The scope of the present disclosure is not limited in this respect.
[0038] Further, the first latent representation can be encoded into the bitstream 230, for example by quantization, probability estimation, entropy coding. In other words, the first latent representation is indicated in the bitstream 230. At the decoding side, the decoded first latent representation (denoted as Figure 2 ) can be obtained by inverse processing (e.g., entropy decoding, dequantization, or the like). In the following, no specific distinction is made between the original first latent representation y and the decoded first latent representation , and both can be collectively referred to as the first latent representation. It should be appreciated that the operations described in the following at the decoding side can equally be performed at the encoding side, and the scope of the present disclosure is not limited in this respect.
[0039] At the feature conversion model 240, at least one intermediate feature (denoted as f) is generated based on the first latent representation. In some embodiments, the at least one intermediate feature can include only one intermediate feature. Alternatively, the at least one intermediate feature can include a plurality of intermediate features, for example N intermediate features (denoted as f i , i e [1, N]), where N is an integer greater than 1. The at least one intermediate feature belongs to the same feature domain as the internal feature (denoted as c) in the subsequent processing at the generation model 250. This will be described in detail in the following.
[0040] Figure 3 A schematic diagram 300 of a network structure of the feature conversion model 240 according to some embodiments of the present disclosure is shown. As shown in Figure 3 , the feature conversion model 240 can include a plurality of stacks 310-1, 310-2, …, 310-L of convolution layers and residual blocks, where L is an arbitrary positive integer. For example, L can be equal to N. For another example, L can be greater than N. In addition, the feature conversion model 240 can also include convolution layers, up-sampling layers, down-sampling layers. The spatial domain size of the stacks 310-1 and 310-L is 64x64, i.e., the height and width of the feature tensor is 64. The spatial domain size of the stack 310-2 is 32x32. It should be appreciated that the specific numerical values mentioned in the context of the present disclosure are only exemplary and do not imply a limitation on the scope of the present disclosure, and any other suitable numerical values can also be employed.
[0041] At the feature conversion model 240, at least one intermediate feature (denoted as f) is generated based on the first latent representation. In some embodiments, the at least one intermediate feature can include only one intermediate feature. Alternatively, the at least one intermediate feature can include a plurality of intermediate features, for example N intermediate features (denoted as f i , i e [1, N]), where N is an integer greater than 1. The at least one intermediate feature belongs to the same feature domain as the internal feature (denoted as c) in the subsequent processing at the generation model 250. This will be described in detail in the following. Figure 3In the example, feature transformation model 240 generates N intermediate features (i.e., f i ,i∈[1,N]), for subsequent processing at generative model 250. It should be understood that, Figure 3 The network structure of the feature transformation model 240 shown is merely exemplary, and the feature transformation model 240 may also include additional elements not shown and / or may omit one (or some) of the elements shown; the scope of this disclosure is not limited in this respect. For example, the feature transformation model 240 may also include deconvolution layers, etc.
[0042] Additionally, the feature transformation model 240 can be fine-tuned based on the pre-trained generative model 250 before use, thereby adaptively adapting to the currently used generative model 250. This will be described in further detail below. With the help of the fine-tuned feature transformation model 240, intermediate features aligned with the internal features of the generative model 250's representation prior knowledge can be generated based on the compressed domain information for the image, so that the intermediate features can be subsequently fused with the internal features. In this way, the compressive domain latent representation can be compatible with different generative models using the feature transformation model 240, thereby effectively improving the flexibility and compatibility of the image decoder.
[0043] Furthermore, the generative model 250 generates a second latent representation for the target image 210 based on at least one of the aforementioned intermediate features. This second latent representation may correspond to the generative domain latent representation of the image. Figure 2 In the example shown, the generative model 250 can conditionally generate from the noisy initial latent representation (in) at least one intermediate feature. Figure 2 The Chinese character is written as z. T The initial hidden representation is then generated. For example, the initial hidden representation could be random noise sampled from a prior Gaussian distribution. Alternatively, the initial hidden representation could be white noise, and so on.
[0044] In some embodiments, the generative model 250 can be a generative diffusion model and can be implemented using a network architecture such as U-Net. The generative model 250 can be divided into a set of sub-models, i.e. Figure 2 The set of sub-models includes sub-models 252-1, 252-2, ..., 252-M, where M is any positive integer, for example, M can be equal to N or M can be greater than N. This set of sub-models includes at least one sub-model, each of which corresponds to one of at least one intermediate feature generated by the feature transformation model 240. For example, in... Figure 2 In the model, sub-model 252-2 corresponds to intermediate feature f1, and sub-model 252-M corresponds to intermediate feature f1. N Correspondingly.
[0045] For each of the at least one sub-model, an input feature for the sub-model is determined based on an intermediate feature corresponding to the sub-model and an output feature of a previous sub-model of the sub-model (i.e., an internal feature of the generative model 250), and the sub-model is utilized to generate an output feature of the sub-model based on the input feature. Here, the intermediate feature and the corresponding internal feature belong to the same feature domain. In other words, the intermediate feature and the corresponding internal feature are aligned. Additionally, the intermediate feature and the corresponding internal feature can have the same dimension and size. Alternatively, the intermediate feature and the corresponding internal feature can also have different dimension and / or size, and the scope of the present disclosure is not limited in this respect.
[0046] Exemplarily, the input feature of the sub-model 252-2 can be generated by fusing the output feature (i.e., the internal feature c1) of the previous sub-model (i.e., the sub-model 252-1) of the sub-model 252-2 with the intermediate feature f1; and the input feature of the sub-model 252-M can be generated by fusing the output feature (i.e., the internal feature c N ) of the previous sub-model of the sub-model 252-M with the intermediate feature f N In some embodiments, the feature fusion can be performed by adding the intermediate feature and the output feature of the previous sub-model to obtain the input feature for the current sub-model. Alternatively, the feature fusion can be performed by multiplying the intermediate feature and the output feature of the previous sub-model to obtain the input feature for the current sub-model.
[0047] In some other embodiments, the feature fusion to obtain the input feature for the current sub-model can be performed by means of a cross-attention mechanism. Exemplarily, a first base feature can be generated based on the intermediate feature and the corresponding internal feature, and a first context feature can be generated based on the intermediate feature. For ease of illustration, it is assumed that the intermediate feature and the corresponding internal feature are both 3-dimensional tensors and have a size of CxHxW, where C, H and W are positive integers. The first base feature can be generated by adding the intermediate feature and the corresponding internal feature and performing a dimension transformation (e.g., by a reshape operation or a shuffle operation). The first base feature can be a 2-dimensional tensor with a size of (H*W)xC. That is, the size of the first base feature in the first dimension is equal to the product of H and W, and the size in the second dimension is equal to C. Similarly, the first context feature can be generated by performing a dimension transformation on the intermediate feature. The first context feature can be a 2-dimensional tensor with a size of (H*W)xC.
[0048] Further, the input feature can be generated by applying a cross-attention mechanism to the first base feature and the first context feature. In some embodiments, a Query input representation for the cross-attention mechanism can be determined based on the first base feature, and a Key input representation and a Value input representation for the cross-attention mechanism can be determined based on the first context feature. For example, the Query input representation can be generated by applying a 1x1 convolutional layer to the first base feature, and the Key input representation and the Value input representation can be generated by applying two 1x1 convolutional layers to the first context feature, respectively. The cross-attention weight can be determined with the following equation:
[0049]
[0050] where Q represents the Query input representation for the cross-attention mechanism, K T represents the transpose of the Key input representation for the cross-attention mechanism, C represents the dimension of the Key input representation, Softmax() represents a Softmax function, and Attn(Q, K, V) represents the cross-attention weight.
[0051] In some embodiments, the output of the cross-attention mechanism can be directly taken as the input feature. Alternatively, a residual connection can be introduced, i.e., the input feature is determined based on the sum of the output of the cross-attention mechanism and the Value input representation. For example, the input feature can be determined with the following equation:
[0052]
[0053] where represents the input feature corresponding to the intermediate feature f, V represents the Value input representation for the cross-attention mechanism, FC() represents a fully connected layer, and Attn(Q, K, V) represents the cross-attention weight. In this way, the gradient vanishing problem can be advantageously alleviated, the generalization ability of the model can be improved, and the model performance can be enhanced.
[0054] By fusing the intermediate feature and the corresponding internal feature with the cross-attention mechanism, the feature representation of each position can be dynamically adjusted according to the similarity between the compressed domain information and the internal feature representation of different positions of the generative model 250, so that it not only retains its own information, but also incorporates the context information of other positions. In this way, not only the spatial information capturing ability of the generative model 250 is enhanced, but also the generative model 250 can obtain more accurate spatial information of the feature map at a finer granularity, and effectively utilize the prior knowledge of the pre-trained generative model 250 to improve the subjective compression effect of the image, thereby effectively improving the quality of the subsequent reconstructed image 270.
[0055] In the context of this disclosure, prior knowledge of a model refers to the prior distribution of model parameters or generation processes within a Bayesian framework. It reflects initial assumptions or empirical knowledge about these parameters or generation processes prior to observation of data. In practical applications, prior knowledge of a model can be provided using parameters or structural information already learned in a pre-trained generative model. This prior knowledge helps constrain the model's solution space, enabling the generation or inference process to more effectively approach the target distribution or outcome.
[0056] In the manner described above, at least one intermediate feature generated by the feature transformation model 240 is fused with a corresponding internal feature in the generation model 250. In this way, the generation model 250 conditionally generates the initial latent representation z based on at least one intermediate feature. T Generate implicit representation z T-1 In some embodiments, the output latent representation of the generative model 250 can be set as the input latent representation of the generative model 250 to perform iterative processing using the generative model 250. For example, starting from the initial latent representation z... T Generate implicit representation z T-1 This can be considered the first iteration. For example, in Figure 2 The ground is schematically shown in the dashed box 255, which implicitly represents z. T-1 After performing T-1 iterations using generative model 250, the second hidden representation for the target image 210 is finally obtained (in Figure 2 Chinese record T can be an integer greater than or equal to 1, and when T equals 1, only one iteration is needed to generate the first hidden representation from the initial hidden representation. In some embodiments, the specific value of T may depend on the specific design of the generative model 250 and its training process.
[0057] It should be understood that the above references Figure 2 The described process for generating the second hidden representation is merely exemplary. The second hidden representation can also be generated using the generative model 250 in any other suitable manner. For example, the generative model 250 can also be a generative adversarial network (GAN), or the T iterations of the generative model 250 described above can be implemented using T models. As another example, the generative model 250 can also generate the second hidden representation directly based on at least one intermediate feature generated by the feature transformation model 240. The scope of this disclosure is not limited in this respect.
[0058] In some embodiments, a reconstructed image 270 for the target image 210 can be generated directly based on the generated second latent representation using the decoding model 260. The decoding model 260 (e.g., a convolutional decoder) can reconstruct pixel domain samples from the latent variable domain to obtain the reconstructed image 270.
[0059] In other embodiments, enhancement operations may also be performed on the generated second hidden representation. Figure 4 A schematic diagram 400 of the local structure for enhancing a second hidden representation of a target image 210, according to some embodiments of the present disclosure, is shown. For example... Figure 4 As shown, the prediction model 410 can be used to generate the predicted hidden representation based on the first hidden representation (in... Figure 4 The Chinese character is written as z. y The predicted latent representation and the second latent representation for the target image 210 belong to the same feature domain. Additionally, the predicted latent representation and the first latent representation may have the same dimensions and size. Alternatively, the predicted latent representation and the first latent representation may have different dimensions and / or sizes. The scope of this disclosure is not limited in this respect.
[0060] For example, the prediction model 410 may consist of convolutional layers, stacks of convolutional layers and residual blocks, and upsampling layers. This prediction model 410 may, for example, implement domain mapping from the compressed domain to the generative domain for the latent representation of an image, and the corresponding dimension alignment. It should be understood that the prediction model 410 may also be implemented in other ways, and the scope of this disclosure is not limited in this respect.
[0061] Furthermore, a third hidden representation for the target image 210 can be generated using the fusion model 420 based on the second hidden representation and the predicted hidden representation (in Figure 4 Chinese record In some embodiments, feature fusion can be performed by adding the second hidden representation and the predicted hidden representation to obtain the third hidden representation. Alternatively, feature fusion can be performed by multiplying the second hidden representation and the predicted hidden representation to obtain the third hidden representation.
[0062] In other embodiments, feature fusion can be performed using a cross-attention mechanism to obtain a third latent representation. Exemplarily, a second base feature can be generated based on a second latent representation and a predicted latent representation, and a second context feature can be generated based on the predicted latent representation. For illustrative purposes, it is assumed that both the second and predicted latent representations are 3-dimensional tensors with dimensions C × H × W, where C, H, and W are positive integers. The second base feature can be generated by adding the second and predicted latent representations and performing a dimensionality transformation (e.g., by reshaping or shuffling). This second base feature can be a 2-dimensional tensor with dimensions (H * W) × C. That is, the size of the second base feature in the first dimension is equal to the product of H and W, and the size in the second dimension is equal to C. Similarly, a second context feature can be generated by performing a dimensionality transformation on the second latent representation. This second context feature can be a 2-dimensional tensor with dimensions (H * W) × C.
[0063] Further, a third hidden representation is generated by applying a cross-attention mechanism to the second base feature and the second context feature. In some embodiments, a query input representation Q for the cross-attention mechanism can be determined based on the second base feature, and a key input representation K and a value input representation V for the cross-attention mechanism can be determined based on the second context feature. For example, the query input representation can be generated by applying a 1x1 convolutional layer to the second base feature, and the key input representation and the value input representation can be generated by applying two 1x1 convolutional layers to the second context feature, respectively. Cross-attention weights can be determined with Equation (1) above. In some embodiments, the output of the cross-attention mechanism can be directly taken as the third hidden representation. Alternatively, a residual connection can be introduced, i.e., the third hidden representation is determined based on the sum of the output of the cross-attention mechanism and the value input representation, which is similar to Equation (2) above, and will not be elaborated here. With the help of the residual connection, the gradient vanishing problem can be alleviated advantageously, the generalization ability of the model can be improved, and the model performance can be improved.
[0064] Further, the conversion can be performed based on the third hidden representation. For example, a reconstructed image 270 for the target image 210 can be generated based on the generated third hidden representation with the decoding model 260. By introducing the spatial cross-attention mechanism in the fusion enhancement process of the predicted hidden representation to the second hidden representation, long-distance dependencies across different spatial positions can be flexibly established, so as to fully exploit and utilize the compressed domain information to enhance the accuracy of the generated domain hidden representation. In this way, the compressed domain information contained in the first hidden representation can be fully exploited to further improve the quality of the reconstructed image 270.
[0065] Based on the above description, according to embodiments of the present disclosure, the compressed domain hidden representation of an image is transmitted in a bitstream, and the compressed domain hidden representation is converted into intermediate features for the generated domain hidden representation of the image with a generation model to perform image coding and decoding. In this way, on the one hand, a general and compact compressed domain representation for the image can be obtained with the help of the hidden representation, thereby effectively reducing the code rate. On the other hand, the compressed domain hidden representation can be converted into intermediate features suitable for the generation model with the help of the feature conversion model for the generation model to reconstruct the image. This can achieve efficient alignment of external compressed domain information and internal prior knowledge of the generation model, thereby realizing image reconstruction with high visual quality and high visual fidelity at very low code rate. Therefore, the efficiency and performance of image coding and decoding can be effectively improved.
[0066] In some additional embodiments, the reconstructed image 270 for the target image 210 can also be enhanced. The inventors have noticed that the image compression reconstruction task often pursues the distortion of the reconstructed image 270 relative to the original image (i.e., the target image 210) to be as small as possible, for which the color distribution related parameters of the reconstructed image 270 should be as consistent as possible with the original image. To this end, the reconstructed image 270 can be updated based on the color distribution of the target image 210 and the color distribution of the reconstructed image 270.
[0067] In some embodiments, at least one statistical metric value for at least one color channel of the target image 210 can be obtained, and at least one statistical metric value for the at least one color channel of the reconstructed image 270 can be determined. Exemplarily, the statistical metric value can include mean, variance, standard deviation, etc. Taking an example in which the at least one color channel includes three channels of red, green and blue (RGB), the mean μ and the variance σ of the pixels in the target image 210 can be calculated for the three channels of RGB, respectively, i.e., μ = {μ R , μ G , μ B}, σ = {σ R , σ G , σ B}. Similarly, the mean and the variance of the pixels in the reconstructed image 270 can be calculated for the three channels of RGB, respectively. Additionally, considering the requirement of encoding transmission, the parameters of the target image 210 can be quantized according to the following equation:
[0068]
[0069] wherein denotes the quantized mean, denotes the quantized variance, μ denotes the mean of the pixels in the target image 210, σ denotes the variance of the pixels in the target image 210, Δ denotes the quantization step, and denotes the rounding operation. It should be understood that, in addition to the rounding operation, the quantization can also be performed in any other suitable manner, such as rounding up, rounding down, etc.
[0070] Further, the reconstructed image 270 can be instance normalized based on the at least one statistical metric value of the reconstructed image 270, and the reconstructed image 270 can be updated by adjusting the result of the instance normalization based on the at least one statistical metric value of the target image 210. Exemplarily but not limitatively, the reconstructed image 270 can be determined based on the following equation:
[0071]
[0072] wherein denotes the reconstructed image 270, and denotes the updated reconstructed image 270.
[0073] In this way, the reconstructed image 270 can be enhanced based on the color distribution information, so that the color accuracy of the reconstructed image 270 can be effectively improved, and thus the quality of the reconstructed image 270 can be improved.
[0074] Example model training process
[0075] An example embodiment of training the machine learning model is described below. Since the generative model 250 is often pre-trained, some embodiments of the present disclosure propose a two-stage model optimization strategy of first encoding end and then decoding end. That is, in the first stage, the model of the encoding end (e.g., the encoding model 220) is trained first; in the second stage, the model of the encoding end is fixed, and the model of the decoding end (e.g., the feature conversion model 240, the prediction model 410, and / or the fusion model 420) is trained.
[0076] For example, the encoding model 220 can be trained as part of an end-to-end autoencoder for encoding and reconstructing images. Exemplarily but not limitatively, an end-to-end compression reconstruction task can be used as a pre-optimization task at the encoding end, a full neural network self-decoder is introduced as the decoding model of the pre-task, and an entropy estimation model is used to predict the code rate. The loss function for the first stage training can be determined by the following equation:
[0077]
[0078] wherein denotes the total loss for the first stage training, denotes the code rate, denotes a distortion metric for the reconstructed image, and λ denotes a weight for adjusting the code rate term. For example, the distortion metric can include mean square error (MSE), multi-scale structural similarity metric (MS-SSIM), etc. It should be understood that the loss function for the first stage training can also be implemented in any other suitable manner, and the scope of the present disclosure is not limited in this respect.
[0079] The parameter values of the end-to-end autoencoder are updated by minimizing the loss function. After the first stage training is completed, only the trained encoding model 220 is used to perform the second stage training for the decoding end model.
[0080] In the second stage training, the trained encoding model 220 can be utilized to generate training latent representations for training images (i.e., images used to train the model) based on the training images. At least one training intermediate feature is generated based on the training latent representations by utilizing the feature transformation model 240. Further, at least one denoising operation can be iteratively performed on the noisy training initial latent representation by utilizing the generative model 250 conditioned on the at least one training intermediate feature, and the parameter values of the feature transformation model 240 are updated based on the difference between the output of one of the at least one denoising operation and the training initial latent representation. In this process, the parameters of the trained encoding model 220 and the generative model 250 can be kept fixed.
[0081] By considering the real data distribution for a large number of images, the feature transformation model 240 is optimized by maximizing the conditional likelihood estimation of the generative model 250 during training. Illustratively, the generative diffusion model aims to sample from Gaussian noise and generate images through a process of iterative denoising. The generative diffusion model can include a forward diffusion process and an inverse diffusion generation process, which are implemented based on a physical process model. In the diffusion process, the image is finally completely converted to a noise domain conforming to a Gaussian distribution by iteratively adding Gaussian noise T steps to the image. In the generation process, the noise to be removed at each step is estimated by a noise estimation network, and Gaussian noise is iteratively removed for T steps to generate an image conforming to a target distribution. The latent generative diffusion model adopts a strategy of denoising generation in the latent variable domain rather than in the pixel domain to cope with the challenge of computing power. It is a two-stage diffusion model, which can include, for example, an autoencoder and a denoising network implemented, for example, by means of a U-Net structure θ : a convolutional encoder in the autoencoder converts a training sample from the pixel domain to an initial sample z0in the latent variable domain, and the U-Net network predicts the noise of the input noisy latent variable z t at time step t and performs iterative denoising calculation. Then, a convolutional decoder reconstructs a sample in the pixel domain from the latent variable domain.
[0082] In the context of the latent generative diffusion model, the loss function for training the feature transformation model 240 in the second stage is as follows:
[0083]
[0084] wherein represents the loss for training the feature transformation model 240 in the second stage, E[] represents mathematical expectation, z0represents an initial sample converted by a convolutional encoder from a training sample in the pixel domain to a latent variable domain, t represents a time step randomly selected according to a strategy, f represents at least one intermediate feature generated by the feature transformation model 240, represents the noise map sampled from a Gaussian distribution (i.e., the initial latent representation) following a normal distribution represents the Euclidean distance, ∈ θ (z t , t, f) represents the denoising network ∈ θ the noisy latent variable z at input t time step conditioned on the intermediate feature f t the output. It should be understood that the loss function for training the feature conversion model 240 can also be implemented in any other suitable manner, and the scope of the present disclosure is not limited in this respect.
[0085] In this way, the feature conversion model 240 can effectively learn the mapping pattern from the first latent representation to the internal knowledge of the generation model 250, so that the feature conversion model 240 can be adapted to the specific generation model 250, realizing the fine-tuning of the feature conversion model 240. In this way, the compression domain latent representation can be adapted to different generation models by means of the feature conversion model 240, thereby effectively improving the flexibility and compatibility of the image decoder. In addition, by means of the two-stage model optimization strategy, in the case of replacing the generation model 250, the first stage of training can no longer be performed, and only the feature conversion model 240 needs to be fine-tuned based on the new generation model 250 by means of the second stage training. In this way, the generalization ability, versatility and compatibility of the coding and decoding model to different pre-training generation models 250 can be effectively improved.
[0086] For the training of the prediction model 410 and the fusion model 420 in Figure 4 , the first training latent representation for the training image can be generated by using the encoding model 220 based on the training image. Then, based on the first training latent representation, the second training latent representation and the predicted training latent representation for the training image can be generated by using the feature conversion model 240, the generation model 250 and the prediction model 410. The predicted training latent representation and the second training latent representation belong to the same feature domain. Further, the third training latent representation for the training image can be generated by using the fusion model 420 based on the second training latent representation and the predicted training latent representation, and the reconstructed image for the training image can be generated based on the third training latent representation. This is similar to the process described above with reference to Figures 2 to 4 , and the present disclosure will not be repeated here.
[0087] Further, the parameter values of the prediction model 410 and the parameter values of the fusion model 420 can be updated based on the difference between the training image and the reconstructed image for the training image. In some embodiments, a first loss can be determined based on the difference between the training image and the reconstructed image. For example, the first loss can measure the perceptual distortion between the reconstructed image and the training image, e.g., by means of a Learned Perceptual Image Patch Similarity (LPIPS) score, etc. Additionally, a second loss can be determined based on the difference between the third training latent representation and a target latent representation for the training image. Further, the parameter values of the prediction model 410 and the parameter values of the fusion model 420 can be updated based on the first loss and the second loss. In one example, the sum of the first loss and the second loss can be determined as a total loss. Alternatively, a weighted sum of the first loss and the second loss can be determined as the total loss.
[0088] By way of example and not limitation, the loss function for training the prediction model 410 and the fusion model 420 is as follows:
[0089]
[0090] wherein denotes the total loss; z0denotes a target latent representation for the training image, which can correspond to an initial sample that the convolutional encoder converts the training sample from the pixel domain to the latent variable domain; denotes a predicted training latent representation for the training image, || ||2denotes the Euclidean distance, denotes an LPIPS score for the training image x and the reconstructed image , and λ z denotes a weight for balancing the first loss and the second loss. In this way, the quality of the reconstructed image can be improved while ensuring the prediction accuracy. It should be appreciated that the loss function for training the prediction model 410 and the fusion model 420 can also be implemented in any other suitable manner, and the scope of the present disclosure is not limited in this regard.
[0091] By minimizing the total loss, the parameter values of the prediction model 410 and the parameter values of the fusion model 420 can be updated iteratively until the total loss converges to be less than a predetermined threshold or the training round reaches a predetermined number.
[0092] In some embodiments, a denoising diffusion implicit model (DDIM) accelerated sampling formula can also be employed at training time to obtain the second latent representation for the training image. The DDIM accelerated results are utilized for subsequent training that enhances the second latent representation. In this way, not only can the training speed be accelerated, saving computing resources, but the generated reconstruction quality can also be significantly improved when the number of iteration generation steps is reduced, achieving dual optimization of generation quality and speed.
[0093] In some embodiments, the training for the prediction model 410 and the fusion model 420 can be performed after the two-stage model optimization strategy described above. Alternatively, the training for the prediction model 410 and the fusion model 420 can be performed in the second stage training in the two-stage model optimization strategy described above. The scope of the present disclosure is not limited in this respect.
[0094] The training process for the machine learning model according to embodiments of the present disclosure is described above. It should be appreciated that the training process above is merely exemplary, and the machine learning model according to embodiments of the present disclosure can also be trained in any other suitable manner. The scope of the present disclosure is not limited in this respect.
[0095] Example method
[0096] Figure 5 A flowchart of a method 500 for image coding according to some embodiments of the present disclosure is shown. In some embodiments, the method 500 can be performed at the image encoder 114 and / or the image decoder 124 as shown. It should be appreciated that the method 500 can also include additional blocks not shown and / or can omit certain block(s) shown, the scope of the present disclosure is not limited in this respect. Figure 1
[0097] At block 502, at least one intermediate feature is generated based on a first latent representation for the target image using a feature conversion model for conversion between the target image and a bitstream of the target image. The first latent representation is indicated in the bitstream.
[0098] At block 504, a second latent representation for the target image is generated based on the at least one intermediate feature using a generative model.
[0099] At block 506, the conversion is performed based on the second latent representation.
[0100] In some embodiments, generating the second latent representation includes generating the second latent representation from a noisy initial latent representation using the generative model conditioned on the at least one intermediate feature.
[0101] In some embodiments, generating the second latent representation from the initial latent representation comprises: setting the initial latent representation as the input latent representation; and iteratively performing the following generation operation at least once: generating, with the generative model, an output latent representation from the input latent representation conditioned on the at least one intermediate feature; and setting the output latent representation as the input latent representation.
[0102] In some embodiments, the generative model comprises at least one sub-model, each of the at least one sub-model corresponds to one of the at least one intermediate feature, and generating the output latent representation from the input latent representation comprises: for each of the at least one sub-model, performing the following operations: determining, based on the intermediate feature corresponding to the sub-model and an output feature of a previous sub-model of the sub-model, an input feature for the sub-model, the output feature belonging to a same feature domain as the intermediate feature; and generating, with the sub-model, the output feature of the sub-model based on the input feature.
[0103] In some embodiments, determining the input feature for the sub-model comprises: generating a first base feature based on the intermediate feature and the output feature of the previous sub-model; generating a first context feature based on the intermediate feature; and generating the input feature by applying a cross-attention mechanism on the first base feature and the first context feature.
[0104] In some embodiments, applying the cross-attention mechanism on the first base feature and the first context feature comprises: determining, based on the first base feature, a query input representation for the cross-attention mechanism; and determining, based on the first context feature, a key input representation and a value input representation for the cross-attention mechanism.
[0105] In some embodiments, the input feature is determined based on a sum of an output of the cross-attention mechanism and the value input representation.
[0106] In some embodiments, performing the conversion comprises: generating, with a decoding model, a reconstructed image for the target image based on the second latent representation.
[0107] In some embodiments, the method 500 further comprises: updating the reconstructed image based on a color distribution of the target image and a color distribution of the reconstructed image.
[0108] In some embodiments, updating the reconstructed image comprises: obtaining at least one statistical metric value for at least one color channel of the target image; determining at least one statistical metric value for the reconstructed image for the at least one color channel; performing instance normalization on the reconstructed image based on the at least one statistical metric value of the reconstructed image; and updating the reconstructed image by adjusting a result of the instance normalization based on the at least one statistical metric value of the target image.
[0109] In some embodiments, the generation model is pre-trained, and the training step for the feature conversion model comprises: generating, based on the training image, a training latent representation for the training image using the trained encoding model; generating, based on the training latent representation, at least one training intermediate feature using the feature conversion model; iteratively performing, using the generation model, at least one denoising operation on a noisy training initial latent representation conditioned on the at least one training intermediate feature; and updating the parameter values of the feature conversion model based on a difference between an output of one of the at least one denoising operation and the training initial latent representation.
[0110] In some embodiments, the encoding model is trained as part of an end-to-end autoencoder, and the end-to-end autoencoder is used to encode and reconstruct the image.
[0111] In some embodiments, performing the conversion comprises: generating, based on the first latent representation, a predicted latent representation using a prediction model, the predicted latent representation and the second latent representation belonging to a same feature domain; generating, based on the second latent representation and the predicted latent representation, a third latent representation for the target image using a fusion model; and performing the conversion based on the third latent representation.
[0112] In some embodiments, generating the third latent representation comprises: generating a second base feature based on the second latent representation and the predicted latent representation; generating a second context feature based on the predicted latent representation; and generating the third latent representation by applying a cross-attention mechanism to the second base feature and the second context feature.
[0113] In some embodiments, applying the cross-attention mechanism to the second base feature and the second context feature comprises: determining, based on the second base feature, a query input representation for the cross-attention mechanism; and determining, based on the second context feature, a key input representation and a value input representation for the cross-attention mechanism.
[0114] In some embodiments, the third latent representation is determined based on a sum of an output of the cross-attention mechanism and the value input representation.
[0115] In some embodiments, the training step for the prediction model and the fusion model comprises: generating, based on the training image, a first training latent representation for the training image using the encoding model; generating, based on the first training latent representation, a second training latent representation and a predicted training latent representation for the training image using the feature conversion model, the generation model, and the prediction model, the predicted training latent representation and the second training latent representation belonging to a same feature domain; generating, based on the second training latent representation and the predicted training latent representation, a third training latent representation for the training image using the fusion model; generating, based on the third training latent representation, a reconstructed image for the training image; and updating the parameter values of the prediction model and the parameter values of the fusion model based on a difference between the training image and the reconstructed image for the training image.
[0116] In some embodiments, updating the parameter values of the prediction model and the parameter values of the fusion model comprises: determining a first loss based on a difference between the training image and the reconstructed image for the training image; determining a second loss based on a difference between the third training hidden representation and the target hidden representation for the training image; and updating the parameter values of the prediction model and the parameter values of the fusion model based on the first loss and the second loss.
[0117] In some embodiments, the image comprises a picture, a frame in a video, or a portion of a frame; or the converting comprises reconstructing the target image from a bitstream.
[0118] Example apparatus and devices
[0119] Embodiments of the present disclosure also provide corresponding apparatuses and devices for implementing the above-described methods or processes. Figure 6 A block diagram of an example apparatus 600 for image coding is shown in accordance with some embodiments of the present disclosure. The apparatus 600 may, for example, be used to implement the methods in accordance with some embodiments of the present disclosure. In some embodiments, the apparatus 600 can be implemented at an image encoder 114 and / or an image decoder 124 as shown. Figure 1 In some embodiments, the apparatus 600 can be implemented at an image encoder 114 and / or an image decoder 124 as shown.
[0120] As shown, the apparatus 600 can include a feature generation module 602, a hidden representation generation module 604, and a conversion execution module 606. The feature generation module 602 is configured to generate, for a conversion between a target image and a bitstream for the target image, at least one intermediate feature based on a first hidden representation for the target image using a feature conversion model. The first hidden representation is indicated in the bitstream. The hidden representation generation module 604 is configured to generate, based on the at least one intermediate feature, a second hidden representation for the target image using a generative model. The conversion execution module 606 is configured to perform the conversion based on the second hidden representation. Figure 6 In some embodiments, the hidden representation generation module 604 is further configured to generate, using the generative model, the second hidden representation from a noisy initial hidden representation conditioned on the at least one intermediate feature.
[0121] In some embodiments, generating the second hidden representation from the initial hidden representation comprises: setting the initial hidden representation as an input hidden representation; and iteratively performing the following generation operation at least once: generating, using the generative model, an output hidden representation from the input hidden representation conditioned on the at least one intermediate feature; and setting the output hidden representation as the input hidden representation.
[0122]
[0123] In some embodiments, the generation model comprises at least one sub-model, each of the at least one sub-model corresponds to one of the at least one intermediate feature, and the generating the output latent representation from the input latent representation comprises: for each of the at least one sub-model, performing the following operations: determining, based on the intermediate feature corresponding to the sub-model and an output feature of a previous sub-model of the sub-model, an input feature for the sub-model, the output feature belonging to a same feature domain as the intermediate feature; and generating, based on the input feature, an output feature of the sub-model using the sub-model.
[0124] In some embodiments, the determining the input feature for the sub-model comprises: generating a first base feature based on the intermediate feature and the output feature of the previous sub-model; generating a first context feature based on the intermediate feature; and generating the input feature by applying a cross-attention mechanism on the first base feature and the first context feature.
[0125] In some embodiments, the applying the cross-attention mechanism on the first base feature and the first context feature comprises: determining, based on the first base feature, a query input representation for the cross-attention mechanism; and determining, based on the first context feature, a key input representation and a value input representation for the cross-attention mechanism.
[0126] In some embodiments, the input feature is determined based on a sum of an output of the cross-attention mechanism and the value input representation.
[0127] In some embodiments, the conversion execution module 606 is further configured to generate, based on the second latent representation, a reconstructed image for the target image using a decoding model.
[0128] In some embodiments, the apparatus 600 further comprises an updating module configured to update the reconstructed image based on a color distribution of the target image and a color distribution of the reconstructed image.
[0129] In some embodiments, the updating the reconstructed image comprises: obtaining at least one statistical metric value for at least one color channel of the target image; determining, for the at least one color channel, at least one statistical metric value of the reconstructed image; performing instance normalization on the reconstructed image based on the at least one statistical metric value of the reconstructed image; and updating the reconstructed image by adjusting a result of the instance normalization based on the at least one statistical metric value of the target image.
[0130] In some embodiments, the generation model is pre-trained, and the training step for the feature conversion model comprises: generating, based on the training image, a training latent representation for the training image by the trained encoding model; generating, based on the training latent representation, at least one training intermediate feature by the feature conversion model; iteratively performing, by the generation model, at least one denoising operation on the noisy training initial latent representation conditioned on the at least one training intermediate feature; and updating the parameter values of the feature conversion model based on a difference between an output of one of the at least one denoising operation and the training initial latent representation.
[0131] In some embodiments, the encoding model is trained as part of an end-to-end autoencoder, and the end-to-end autoencoder is used for encoding and reconstructing the image.
[0132] In some embodiments, the conversion performing module 606 is further configured to: generate, based on the first latent representation, a predicted latent representation by the prediction model, the predicted latent representation and the second latent representation belonging to a same feature domain; generate, based on the second latent representation and the predicted latent representation, a third latent representation for the target image by the fusion model; and perform the conversion based on the third latent representation.
[0133] In some embodiments, generating the third latent representation comprises: generating, based on the second latent representation and the predicted latent representation, a second base feature; generating, based on the predicted latent representation, a second context feature; and generating the third latent representation by applying a cross-attention mechanism to the second base feature and the second context feature.
[0134] In some embodiments, applying the cross-attention mechanism to the second base feature and the second context feature comprises: determining, based on the second base feature, a query input representation for the cross-attention mechanism; and determining, based on the second context feature, a key input representation and a value input representation for the cross-attention mechanism.
[0135] In some embodiments, the third latent representation is determined based on a sum of an output of the cross-attention mechanism and the value input representation.
[0136] In some embodiments, the training step for the prediction model and the fusion model comprises: generating, based on the training image, a first training latent representation for the training image by the encoding model; generating, based on the first training latent representation, a second training latent representation and a predicted training latent representation for the training image by the feature conversion model, the generation model and the prediction model, the predicted training latent representation and the second training latent representation belonging to a same feature domain; generating, based on the second training latent representation and the predicted training latent representation, a third training latent representation for the training image by the fusion model; generating, based on the third training latent representation, a reconstructed image for the training image; and updating the parameter values of the prediction model and the parameter values of the fusion model based on a difference between the training image and the reconstructed image for the training image.
[0137] In some embodiments, updating the parameter values of the prediction model and the parameter values of the fusion model comprises determining a first loss based on a difference between the training image and the reconstructed image for the training image, determining a second loss based on a difference between the third training hidden representation and the target hidden representation for the training image, and updating the parameter values of the prediction model and the parameter values of the fusion model based on the first loss and the second loss.
[0138] In some embodiments, the image comprises a picture, a frame in a video, or a portion of a frame; or the converting comprises reconstructing the target image from a bitstream.
[0139] The modules and / or units included in the apparatus 600 can be implemented using various means, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, part or all of the units in the apparatus 600 can be implemented by one or more hardware logic components. Examples of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0140] Figure 6 The modules and / or units illustrated in FIG. 6 can be implemented as, included in, or hosted by hardware modules, software modules, firmware modules, or any combination thereof. In particular, in some embodiments, the procedures, methods, or processes described above can be implemented by hardware of a storage system or a host corresponding to the storage system or other computing device independent of the storage system.
[0141] Figure 7 A block diagram of a device 700 is shown in which one or more embodiments of the disclosure can be implemented. It should be understood that Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to implement the image encoder 114, the image decoder 124, and / or the methods described above shown in FIG. 6. Figure 1 The electronic device 700 shown can be used to implement the image encoder 114, the image decoder 124, and / or the methods described above shown in FIG. 6.
[0142] As Figure 7As shown, the electronic device 700 is in the form of a general electronic device. Components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and capable of executing various processing in accordance with programs stored in the memory 720. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of the electronic device 700.
[0143] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media that is located either internally or externally to the electronic device 700 such as volatile and non-volatile media, removable and non-removable media. The memory 720 can be volatile memory (such as cache, RAM), non-volatile memory (such as ROM, EEPROM, flash memory), or some combination of the two. The storage device 730 can be a removable or non-removable media, and can include machine-readable media such as flash drives, magnetic disks, or any other media that can be used to store information and / or data (e.g., training data for training) and that can be accessed by the electronic device 700.
[0144] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 7 disk drives for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the implementations of the present disclosure.
[0145] The communication unit 740 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0146] Input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 740, as desired, in order to communicate with one or more devices that enable a user to interact with electronic device 700, or any device (e.g., a network card, a modem, etc.) that enables electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0147] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.
[0148] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the size and relative positioning of components can be exaggerated for clarity and / or descriptive purposes. Also, the drawings represent examples of the various aspects of the disclosure and are not limiting of the scope of the present disclosure. It should be further understood that the drawings are not necessarily drawn to scale and that, in particular, the scales of the various components can be different in any drawing.
[0149] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including a manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0150] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0151] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0152] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although being shown and described in terms of certain implementations and overall functions, the implementations are not intended to exclude other implementations or technologies. Modifications and changes can be made in arrangement, operation, and details of the methods and apparatus described. Many modifications and variations of the described implementations are possible and will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. It is therefore intended that the description be considered in all respects as illustrative, rather than limiting, of the disclosed implementations. Changes are made to the descriptions that are intended to be protective of the applications claimed. It is therefore intended that the description be considered in all respects as illustrative, rather than limiting, of the disclosed implementations. Changes are made to the descriptions that are intended to be protective of the applications claimed.
Claims
1. A method for image encoding and decoding, comprising: For the conversion between a target image and its bitstream, at least one intermediate feature is generated using a feature transformation model based on a first hidden representation of the target image, wherein the first hidden representation is indicated in the bitstream; Based on the at least one intermediate feature, a second hidden representation for the target image is generated using a generative model; as well as The transformation is performed based on the second implicit representation.
2. The method of claim 1, wherein generating the second hidden representation comprises: Using the at least one intermediate feature as a condition, the second hidden representation is generated from the noisy initial hidden representation using the generative model.
3. The method of claim 2, wherein generating the second hidden representation from the initial hidden representation comprises: Set the initial implicit representation as the input implicit representation; as well as Perform the following generation operation at least once: Using the at least one intermediate feature as a condition, the generative model generates an output latent representation from the input latent representation; and Set the output implicit representation as the input implicit representation.
4. The method of claim 3, wherein the generating model comprises at least one sub-model, each of the at least one sub-model corresponding to one of the at least one intermediate features, and generating the output latent representation from the input latent representation comprises: For each of the at least one sub-models, perform the following operations: Based on the intermediate features corresponding to the sub-model and the output features of the previous sub-model, the input features for the sub-model are determined, wherein the output features and the intermediate features belong to the same feature domain. as well as Based on the input features, the output features of the sub-model are generated using the sub-model.
5. The method of claim 4, wherein determining the input features for the sub-model comprises: Based on the intermediate features and the output features of the previous sub-model, a first base feature is generated; Based on the intermediate features, a first context feature is generated; as well as The input features are generated by applying a cross-attention mechanism to the first base features and the first context features.
6. The method of claim 5, wherein applying the cross-attention mechanism to the first base feature and the first context feature comprises: Based on the first base feature, a query input representation for the cross-attention mechanism is determined; as well as Based on the first contextual features, the key input representation and value input representation for the cross-attention mechanism are determined.
7. The method of claim 6, wherein the input features are determined based on the sum of the output of the cross-attention mechanism and the value input representation.
8. The method of claim 1, wherein performing the conversion comprises: Based on the second hidden representation, a reconstructed image for the target image is generated using a decoding model.
9. The method according to claim 8, further comprising: The reconstructed image is updated based on the color distribution of the target image and the color distribution of the reconstructed image.
10. The method of claim 9, wherein updating the reconstructed image comprises: Obtain at least one statistical metric value for at least one color channel of the target image; For the at least one color channel, determine at least one statistical metric of the reconstructed image; Instance normalization is performed on the reconstructed image based on at least one statistical metric. as well as The reconstructed image is updated by adjusting the result of instance normalization based on at least one statistical metric of the target image.
11. The method of claim 1, wherein the generative model is pre-trained, and the training step for the feature transformation model comprises: Based on the training images, a training latent representation for the training images is generated using a trained encoding model; Based on the trained latent representation, at least one intermediate training feature is generated using the feature transformation model; Using the at least one intermediate training feature as a condition, the generative model iteratively performs at least one denoising operation on the noisy initial training latent representation. as well as The parameter values of the feature transformation model are updated based on the difference between the output of one of the at least one denoising operations and the initial training latent representation.
12. The method of claim 11, wherein the encoding model is trained as part of an end-to-end autoencoder, and the end-to-end autoencoder is used to encode and reconstruct images.
13. The method of claim 1, wherein performing the conversion comprises: Based on the first hidden representation, a predicted hidden representation is generated using a prediction model, wherein the predicted hidden representation and the second hidden representation belong to the same feature domain; Based on the second hidden representation and the predicted hidden representation, a third hidden representation for the target image is generated using a fusion model; as well as The transformation is performed based on the third implicit representation.
14. The method of claim 13, wherein generating the third hidden representation comprises: Based on the second hidden representation and the predicted hidden representation, a second base feature is generated; Based on the predicted latent representation, a second contextual feature is generated; as well as The third hidden representation is generated by applying a cross-attention mechanism to the second base feature and the second context feature.
15. The method of claim 14, wherein applying the cross-attention mechanism to the second base feature and the second context feature comprises: Based on the second base feature, a query input representation for the cross-attention mechanism is determined; as well as Based on the second contextual feature, the key input representation and value input representation for the cross-attention mechanism are determined.
16. The method of claim 15, wherein the third hidden representation is determined based on the sum of the output of the cross-attention mechanism and the value input representation.
17. The method of claim 13, wherein the training step for the prediction model and the fusion model comprises: Based on the training images, a first training latent representation for the training images is generated using an encoding model; Based on the first training latent representation, a second training latent representation and a predicted training latent representation for the training image are generated using the feature transformation model, the generation model, and the prediction model. The predicted training latent representation and the second training latent representation belong to the same feature domain. Based on the second training latent representation and the predicted training latent representation, a third training latent representation for the training image is generated using the fusion model; Based on the third training hidden representation, a reconstructed image for the training image is generated; as well as The parameter values of the prediction model and the parameter values of the fusion model are updated based on the differences between the training image and the reconstructed image based on the training image.
18. The method of claim 17, wherein updating the parameter values of the prediction model and the parameter values of the fusion model comprises: A first loss is determined based on the difference between the training image and the reconstructed image based on the training image; A second loss is determined based on the difference between the third training latent representation and the target latent representation for the training image; as well as Based on the first loss and the second loss, update the parameter values of the prediction model and the parameter values of the fusion model.
19. The method according to any one of claims 1 to 18, wherein the image comprises a picture, a frame of a video, or a portion of said frame; or The conversion mentioned therein includes reconstructing the target image from the bitstream.
20. An apparatus for image encoding and decoding, comprising: The feature generation module is configured to: for the conversion between a target image and a bitstream of the target image, generate at least one intermediate feature based on a first hidden representation of the target image using a feature conversion model, wherein the first hidden representation is indicated in the bitstream; The latent representation generation module is configured to: generate a second latent representation for the target image based on the at least one intermediate feature using a generative model; as well as The transformation execution module is configured to perform the transformation based on the second implicit representation.
21. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 19.
22. A computer-readable storage medium having instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 19.