Image coding and decoding method, training method and device of image coding and decoding system, equipment, medium and product
By acquiring the compression characteristics and semantic residuals of the image and generating coding information, the problem of image quality degradation at extremely low bit rates in traditional image compression technology is solved, and efficient image transmission and high-quality reconstruction are achieved.
Patent Information
- Application Number
- CN202510639739.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional image compression technology has significantly reduced image quality at extremely low bit rates, resulting in poor image transmission efficiency and making it difficult to achieve efficient image transmission.
By acquiring the compressed features and semantic residuals of the target image, encoding information is generated, and the compressed lost information is characterized by semantic residuals, and image reconstruction is performed in combination with the diffusion model to optimize image transmission efficiency and quality.
While improving image transmission efficiency, it ensures image reconstruction quality, enhances the reconstruction quality of the decoded image, and reduces visual artifacts and information redundancy.
Smart Images

Figure CN120390096A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image encoding and decoding technologies, and in particular, to an image encoding and decoding method, a training method and device, equipment, medium and product of an image encoding and decoding system. Background Art
[0002] With the development of image encoding and decoding technologies, traditional image compression standards such as Joint Photographic Experts Group 2000 (JPEG2000), Better Portable Graphics (BPG), and Versatile Video Coding (VVC) have been widely used in practical applications.
[0003] However, these image compression technologies may exhibit obvious blocking artifacts at extremely low bitrates, resulting in a significant decline in image quality, which makes it difficult to achieve efficient image transmission in the case of extremely limited bandwidth. Summary of the Invention
[0004] The present disclosure provides an image encoding and decoding method, a training method and device, equipment, medium and product of an image encoding and decoding system, so as to at least solve the problems of poor image quality and poor image transmission efficiency in image transmission in related technologies. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, an image encoding method is provided. The image encoding method includes: obtaining a target image to be encoded; compressing original features of the target image to obtain compressed features of the target image; determining a semantic residual based on the target image and a compressed image corresponding to the compressed features, where the semantic residual represents information lost due to the compression; and generating encoded information of the target image based on the compressed features and the semantic residual.
[0006] Optionally, the determining a semantic residual based on the target image and a compressed image corresponding to the compressed features includes: performing semantic extraction on the target image to obtain first semantic information; performing semantic extraction on the compressed image to obtain second semantic information; and determining the semantic residual based on a difference between the first semantic information and the second semantic information.
[0007] Optionally, determining the semantic residual based on the difference between the first semantic information and the second semantic information includes: determining redundant information in the first semantic information and the second semantic information; determining mismatched information in the second semantic information that does not appear in the first semantic information; and determining the semantic residual based on the first semantic information, the redundant information, and the mismatched information.
[0008] According to a second aspect of the embodiments of the present disclosure, an image decoding method is provided. The image decoding method includes: obtaining encoding information of a target image to be decoded; obtaining a compression feature and a semantic residual of the target image based on the encoding information, where the compression feature is obtained by compressing an original feature of the target image, and the semantic residual represents information lost due to the compression; and obtaining a decoded image based on the compression feature and the semantic residual.
[0009] Optionally, obtaining a decoded image based on the compression feature and the semantic residual includes: adding noise to the compression feature to obtain a noisy feature; denoising the noisy feature by predicting noise in the noisy feature under the constraint of the semantic residual to obtain a denoised feature; and obtaining the decoded image based on the denoised feature.
[0010] Optionally, denoising the noisy feature by predicting noise in the noisy feature under the constraint of the semantic residual to obtain a denoised feature includes: generating a control feature based on the compression feature, where the control feature is used to limit parameters in the process of obtaining the decoded image; and denoising the noisy feature by predicting noise in the noisy feature under the constraints of the semantic residual and the control feature to obtain a denoised feature.
[0011] Optionally, denoising the noisy feature by predicting noise in the noisy feature to obtain a denoised feature includes: performing denoising on the noisy feature a preset number of times through the preset number of predictions to obtain a denoised feature, where the preset number is related to the compression degree of the target image.
[0012] Optionally, the image decoding method is executed by a decoding module. The decoding module includes a control module and a generation module. The control module is used to generate the control feature, and the generation module is used to predict noise in the noisy feature and denoise the noisy feature, where the control module is obtained based on at least a part of the network structure in the generation module.
[0013] According to a third aspect of the embodiments of the present disclosure, a training method for an image coding and decoding system is provided. The image coding and decoding system includes an encoding module and a decoding module. The training method includes: obtaining a training sample set, where the training sample set includes training image samples; using the encoding module to compress the original features of the training image samples to obtain compressed features of the training image samples, and determining a semantic residual based on the training image samples and the training compressed images corresponding to the compressed features, where the compressed features are obtained by compressing the features of the training image samples, and the semantic residual represents the information lost due to the compression; using the encoding module to generate encoding information of the training image samples based on the compressed features and the semantic residual; using the decoding module to obtain the compressed features and the semantic residual of the training image samples based on the encoding information, and obtaining a reconstructed image based on the compressed features and the semantic residual; and training the image coding and decoding system based on the training image samples and the reconstructed image.
[0014] Optionally, the obtaining the reconstructed image based on the compressed features and the semantic residual includes: adding noise to the compressed features to obtain noise-added features; denoising the noise-added features by predicting the noise in the noise-added features under the constraint of the semantic residual to obtain denoised features; and obtaining the reconstructed image based on the denoised features.
[0015] Optionally, the training the image coding and decoding system based on the training image samples and the reconstructed image includes: determining a first loss related to the encoding module based on the difference between the reconstructed image and the training image samples; determining a second loss related to the decoding module based on the reconstructed image, the training image samples, and the training predicted noise; and training the image coding and decoding system based on the first loss and the second loss.
[0016] Optionally, the determining the second loss based on the reconstructed image, the training image samples, and the training predicted noise includes: determining a feature domain loss based on the training predicted noise, the image features of the reconstructed image, and the image features of the training image samples; determining a pixel domain loss based on the difference between the pixel features of the reconstructed image and the pixel features of the training image samples; and determining the second loss based on the feature domain loss and the pixel domain loss.
[0017] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image encoding device, including: an obtaining unit configured to obtain a target image to be encoded; a compressing unit configured to compress the original features of the target image to obtain compressed features of the target image; a determining unit configured to determine a semantic residual based on the target image and a compressed image corresponding to the compressed features, where the semantic residual represents information lost due to the compression; and a generating unit configured to generate encoded information of the target image based on the compressed features and the semantic residual.
[0018] According to a fifth aspect of the embodiments of the present disclosure, there is provided an image decoding device, including: an information obtaining unit configured to obtain encoded information of a target image to be decoded; an information determining unit configured to obtain compressed features and a semantic residual of the target image based on the encoded information, where the compressed features are obtained by compressing the original features of the target image, and the semantic residual represents information lost due to the compression; and an image determining unit configured to obtain a decoded image based on the compressed features and the semantic residual.
[0019] According to a sixth aspect of the embodiments of the present disclosure, there is provided a training device for an image encoding and decoding system, where the image encoding and decoding system includes an encoding module and a decoding module, and the training device includes: a training obtaining unit configured to obtain a training sample set, where the training sample set includes training image samples; a training compressing unit configured to use the encoding module to compress the original features of the training image samples to obtain compressed features of the training image samples, and determine a semantic residual based on the training image samples and training compressed images corresponding to the compressed features, where the compressed features are obtained by compressing the features of the training image samples, and the semantic residual represents information lost due to the compression; a training generating unit configured to use the encoding module to generate encoded information of the training image samples based on the compressed features and the semantic residual; a training reconstructing unit configured to use the decoding module to obtain compressed features and a semantic residual of the training image samples based on the encoded information, and obtain a reconstructed image based on the compressed features and the semantic residual; and a training unit configured to train the image encoding and decoding system based on the training image samples and the reconstructed image.
[0020] According to a seventh aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing processor-executable instructions, where when the processor-executable instructions are run by the processor, the processor is caused to execute the image encoding method or the image decoding method or the training method for an image encoding and decoding system according to the embodiments of the present disclosure.
[0021] According to an eighth aspect of embodiments of the present disclosure, there is provided a computer-readable storage medium, which when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the image encoding method, the image decoding method, or the training method of the image encoding and decoding system according to embodiments of the present disclosure.
[0022] According to a ninth aspect of embodiments of the present disclosure, there is provided a computer program product including computer instructions, which when executed by a processor, implement the image encoding method, the image decoding method, or the training method of the image encoding and decoding system according to embodiments of the present disclosure.
[0023] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0024] By compressing the original features of the target image to be encoded to obtain the compressed features of the target image, the semantic residual can be determined based on the target image and the compressed image corresponding to the compressed features. This semantic residual can represent the information lost due to compression. Thus, the encoding information of the target image can be generated based on the compressed features and the semantic residual. In this way, the information loss caused by image compression can be better represented by the semantic residual, ensuring the image quality while improving the image transmission efficiency through compression, which is beneficial to enhancing the quality of image reconstruction on the decoding side.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0027] Figure 1 is a schematic diagram of an exemplary implementation scenario of an image encoding or decoding method shown according to an exemplary embodiment of the present disclosure.
[0028] Figure 2 is a schematic flowchart of an image encoding method shown according to an exemplary embodiment of the present disclosure.
[0029] Figure 3 is a schematic architecture diagram of an image encoding and decoding system shown according to an exemplary embodiment of the present disclosure.
[0030] Figure 4 is a schematic diagram of the working principle of a semantic residual extraction module shown according to an exemplary embodiment of the present disclosure.
[0031] Figure 5 It is a schematic diagram showing the subjective effect of semantic residual extraction on an example data set according to an exemplary embodiment of the present disclosure.
[0032] Figure 6 It is a schematic diagram showing the bit rate saved by semantic residual extraction at different bit rates according to an exemplary embodiment of the present disclosure.
[0033] Figure 7 It is a schematic flowchart showing an image decoding method according to an exemplary embodiment of the present disclosure.
[0034] Figure 8 It is a schematic flowchart showing the steps of obtaining a decoded image in an image decoding method according to an exemplary embodiment of the present disclosure.
[0035] Figure 9 It is a schematic diagram comparing the subjective effects of an image decoding method according to an exemplary embodiment of the present disclosure with a method without constraints on an example data set.
[0036] Figure 10 It is the influence of different preset times on different bit rate points shown according to an exemplary embodiment of the present disclosure.
[0037] Figure 11 It is a schematic flowchart showing a training method of an image codec system according to an exemplary embodiment of the present disclosure.
[0038] Figure 12 It is a schematic flowchart showing training an image codec system based on a first loss and a second loss in a training method of an image codec system according to an exemplary embodiment of the present disclosure.
[0039] Figure 13 It is a block diagram showing an image encoding device according to an exemplary embodiment of the present disclosure.
[0040] Figure 14 It is a block diagram showing an image decoding device according to an exemplary embodiment of the present disclosure.
[0041] Figure 15 It is a block diagram showing a training device of an image codec system according to an exemplary embodiment of the present disclosure.
[0042] Figure 16 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0043] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0044] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0045] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of three parallel cases: "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0046] As described above, traditional image compression technologies have problems of poor image quality and low image transmission efficiency.
[0047] Specifically, traditional image compression standards such as JPEG2000, BPG, and VVC usually adopt a block-based processing method, and obvious block effects may occur at extremely low bit rates, resulting in a significant decrease in image quality. In the case of extremely limited bandwidth, it is difficult to complete efficient image transmission.
[0048] In this regard, with the rapid development of deep learning technology, learning-based image compression methods can show the potential to outperform traditional codecs. According to different optimization objectives, learning-based image compression methods can generally be divided into distortion-oriented methods and perception-oriented methods.
[0049] Distortion-oriented methods are mainly achieved by optimizing the rate-distortion function. Such methods usually emphasize minimizing the reconstruction error of the image during the compression process, thereby improving the compression performance in a mathematical sense. However, this strategy may lead to an unrealistic reconstructed image at extremely low bit rates, usually manifested as being too smooth and blurred, lacking details.
[0050] Different from distortion-oriented methods, perception-oriented methods pay more attention to visual quality. These methods usually adopt techniques such as adversarial training (e.g., Generative Adversarial Network (GAN)) to improve perceptual quality by simulating the preferences of the human visual system. Such methods have made remarkable progress in visual effects and can generate more attractive images. However, in scenarios with extremely low bitrates, this method may produce unpleasant visual artifacts, such as texture distortion or unnatural edges.
[0051] With the development of the Diffusion Model, this technology has demonstrated powerful multi-modal information processing capabilities and can be applied to fields such as image generation and image restoration. In the field of image compression, the method based on the Diffusion Model can further improve the visual reconstruction quality at extremely low bitrates.
[0052] However, in the compression method based on the Diffusion Model, text and other modal information can be combined to control the Diffusion Model to recover the original image from pure noise. However, this method does not fully consider the redundancy of text information and ignores the large amount of latency and uncertainty brought about by generating from pure noise. Therefore, such a design does not conform to the core objective of the image compression task.
[0053] To solve at least some of the above problems, exemplary embodiments of the present disclosure propose an image encoding method, an image decoding method, an image encoding device, an image decoding device, a training method for an image codec system, a training device for an image codec system, an electronic device, a computer-readable storage medium, and a computer program product. The following will be described in detail with reference to Figures 1 to 16 for details.
[0054] The following refers to Figure 1 to give an example implementation scenario of the image encoding method or the image decoding method according to an exemplary embodiment of the present disclosure.
[0055] As Figure 1 shown, when a user requests an image from the server 101 through the network 102 at the user terminal (e.g., mobile phone 103, desktop computer 104, tablet computer 105, etc.), the server 101 can send the encoding information of the corresponding image to the user terminal (e.g., mobile phone 103, desktop computer 104, tablet computer 105, etc.) through the network 102, and the user terminal can display the received image.
[0056] In the above process, the server 101 may encode the transmitted image. During the encoding process, the transmitted image may be encoded according to the image encoding method of the exemplary embodiments of the present disclosure; and / or, the user terminal may decode the encoding information of the received image. During the decoding process, the received image may be decoded according to the image decoding method of the exemplary embodiments of the present disclosure.
[0057] It should be noted that although the above is described by taking the server as an example, it is only an example. The main body of the image encoding method may be any electronic device. Here, the electronic device may include, for example, entity devices such as smart phones, tablet computers, laptop computers, digital assistants, wearable devices, vehicle-mounted terminals, etc. as hardware codecs, or may also include software running on the entity device as a software video encoder.
[0058] It should also be noted that although the application scenario of the user terminal requesting an image is used as an example here for elaboration, it should be understood that the application scenarios of the image encoding method or the image decoding method according to the present disclosure are not limited thereto, and it may also be applied to any other application scenarios involving image encoding or decoding.
[0059] According to the first aspect of the embodiments of the present disclosure, an image encoding method is provided. Figure 2 FIG. is a flowchart showing an image encoding method according to an exemplary embodiment of the present disclosure. This method can better represent the information loss caused by image compression through semantic residuals, so that while improving the image transmission efficiency through compression, the image quality can also be ensured, which is beneficial to enhancing the quality of image reconstruction on the decoding side.
[0060] As shown in Figure 2 FIG., the image encoding method according to the exemplary embodiment of the present disclosure may include the following steps:
[0061] In step S210, a target image to be encoded may be obtained.
[0062] Here, the target image may be an image in any format and with any content. As an example, the target image may be a video frame in a video to be encoded.
[0063] In step S220, the original features of the target image may be compressed to obtain the compressed features of the target image.
[0064] In this step, the target image may be feature-mapped to obtain the original features of the target image; the original features may be feature-compressed to obtain the compressed features.
[0065] Here, a pre-trained Variational Autoencoder (VAE) image codec can be used to map the target image from the pixel domain to the feature space, obtaining the original features. Thus, the complexity and dimension of the data can be reduced by compressing the original features in the feature space, resulting in compressed features. Through image compression, the data volume of the original image during transmission or storage can be significantly reduced.
[0066] In one example, an existing image compression method can be adopted to compress the original features of the target image, obtaining compressed features. In another example, a deep learning-based image compression method can be used, and a pre-trained image codec is utilized to compress the original features of the target image, obtaining compressed features.
[0067] For example, an image compression model architecture such as Efficient Learned Image Compression (ELIC) can be adopted. During the compression process, the original features obtained by mapping can be downsampled (e.g., four times downsampling) through convolution (e.g., two-layer convolution), and then the same hyperprior codec, context model, and entropy model as ELIC are used for compression encoding to obtain the encoded information of the compressed features.
[0068] Specifically, the features of the input target image can be mapped through a feature encoder. For example, through an integrated Stable Diffusion image encoder, it is mapped from the pixel domain to the feature domain. During the mapping process, the pixel information of the target image is converted into a latent feature representation, which can contain high-level semantic information of the target image, such as the shape, structure, and texture of the object, etc. These latent feature representations are a compressed version of the image information and can effectively capture the key semantic content of the image. These latent features can be input into the hyperprior encoder to extract the hyperprior information of the latent features. Here, the hyperprior information not only helps to further compress the features but also provides auxiliary information for the decoding stage. Then, the quantized latent features and the hyperprior information can be entropy-encoded together to generate a binary bitstream, thus obtaining the encoded information of the compressed features.
[0069] Here, after being processed by the feature compression module, the compressed latent feature representation will be used for subsequent semantic residual extraction and image reconstruction. Through compression, redundant information in the feature representation can be reduced, the compression efficiency can be improved, and at the same time, it is ensured that the image maintains a high visual quality at a low bit rate.
[0070] Figure 3 Shows an example architecture of an image codec system according to an embodiment of the present disclosure. As Figure 3As shown, the system may include an encoding module and a decoding module. The encoding module may include a feature compression unit, which may adopt a structure similar to a learning-based image coding framework. For example, it may include a feature encoder, a hyperprior encoder, and an image compression model (such as the ELIC image compression model). In Figure 3 it, the feature encoder can be used to map the target image X to the feature space to obtain the original features; the image compression model can be used to perform feature compression on the original features to obtain compressed features, and the hyperprior encoder can be used to encode the compressed features to obtain the encoded information of the compressed features, and this encoded information can be sent to the decoding side.
[0071] In step S230, the semantic residual can be determined based on the target image and the compressed image corresponding to the compressed features.
[0072] Here, the semantic residual can represent the information lost due to compression, and it can be related to the missing semantic information in the compressed image.
[0073] Specifically, in some cases, the semantic information of the original image can be obtained by directly extracting the text description (caption) of the original image, and the semantic information of the original image and the compressed features can be sent to the decoding side to reconstruct the original image at the decoding side. However, in this process, there may be duplicate information between the compressed features and the semantic information of the original image, and the complete description of the original image may not fully reflect the key information lost during the compression process, but there is information already included in the compressed features. Therefore, it may easily lead to semantic redundancy problems in the transmitted information, which is not conducive to optimizing the image transmission efficiency.
[0074] In the embodiments of the present disclosure, different from directly using the text description of the original image, in step S230, specific prompt words can be designed for the difference between the latent features of the compressed image and the features of the original image. These prompt words are designed to capture the key information lost during the compression process and avoid generating redundant semantic information. In this way, the efficiency and accuracy of semantic information extraction can be improved, and the image transmission efficiency can be optimized.
[0075] In this step, the compressed image can be obtained based on the compressed features. For example, Figure 3 the feature compression unit of can also include a decoder, which can decode the compressed features into the pixel domain to obtain the compressed image. Here, in the decoding stage, the binary bitstream can first go through entropy decoding to recover the compressed features, and the recovered features are sent to the decoder to finally obtain the compressed image feature representation.
[0076] In the case of obtaining a compressed image, in one example, the feature differences between the target image and the compressed image can be compared to determine the feature differences between the two, and the text description corresponding to the feature differences can be extracted as the above semantic residual.
[0077] In another example, this step S230 may include the following steps: performing semantic extraction on the target image to obtain first semantic information; performing semantic extraction on the compressed image to obtain second semantic information; determining the semantic residual based on the difference between the first semantic information and the second semantic information.
[0078] As an example, the image generation ability of multimodal large language models such as GPT-4 and Llama 3 can be used to perform semantic extraction on the image to obtain the text descriptions (captions) of the original image and the compressed representation.
[0079] Specifically, the decoded compressed image and the target image can be fed into a multimodal large language model for processing. The multimodal large language model can generate descriptive prompts for the two images respectively, and the prompts are used to summarize the core semantic information of the images, such as the above first semantic information and second semantic information. To ensure the efficiency and accuracy of the extracted semantic residual, optimized prompts need to be designed to guide the large language model to generate accurate text descriptions related to the image content.
[0080] After obtaining the first semantic information and the second semantic information, the first semantic information and the second semantic information can be input into the large language model again, and the powerful semantic reasoning ability of the model is used to extract the semantic residual information. Specifically, the model can perform differential processing on the first semantic information and the second semantic information to obtain the semantic residual. Through this differential method, the model can automatically identify and extract the key semantic details lost during the compression process, such as but not limited to small objects in the image and local structural features. In this way, the semantic residual information can be accurately extracted and input into the subsequent generation module together with the compression features as supplementary information. The semantic residual can describe the important semantics that exist in the original target image but are missing in the compressed image, so the quality of image reconstruction can be enhanced.
[0081] In the above process, the process of differential processing can be determined according to the actual working principle of the model. An example process of differential processing is given below.
[0082] As an example, the step of determining the semantic residual based on the difference between the first semantic information and the second semantic information may include: determining the redundant information in the first semantic information and the second semantic information; determining the mismatched information in the second semantic information that does not appear in the first semantic information; determining the semantic residual based on the first semantic information, the redundant information, and the mismatched information.
[0083] Redundant information may include, but is not limited to, words, phrases, and sentences contained in both the first and second semantic information. Redundant information may be information that appears repeatedly in both the first and second semantic information. The redundant information is already reflected in the compressed image and compression features, and therefore does not need to be sent to the decoding side as semantic information.
[0084] The mismatching information may be information that appears in the second semantic information but not in the first semantic information. The mismatching information may include, but is not limited to, words, phrases, sentences, etc. Here, the mismatching information may be erroneous information obtained from the compressed image that does not exist in the original image and therefore needs to be removed to avoid affecting the quality of the reconstructed image.
[0085] For example, based on the first semantic information, redundant information and mismatched information may be deleted to obtain a final semantic residual.
[0086] As an example, in Figure 3 In the image encoding and decoding system, the encoding module further includes a semantic residual extraction unit, which may include a multimodal large model (MLLM) that can extract text descriptions of the image. Figure 4 An example process of the semantic residual extraction principle according to an embodiment of the present disclosure is shown.
[0087] like Figure 3 and Figure 4 As shown, the target image X and the compressed image X′ corresponding to the compressed feature can be input into a semantic residual extraction unit (such as a multimodal large model), and the semantic residual extraction unit is used to extract the text descriptions of the target image X and the compressed image X′ respectively to obtain the first semantic information f mllm (X) and the second semantic information f mllm (X′). The first semantic information f mllm (X) and the second semantic information f mllm (X′) is input into the semantic residual extraction unit again, and the semantic residual extraction unit is used to obtain mismatch information (Mismatch) and redundant information (Redundancy).
[0088] by Figure 4 For example, the unmatched information may include "house", which is in the second semantic information f mllm (X′) appears in the first semantic information f mllm (X) does not appear; redundant information may include "Red, surrounded by trees", which is in the first semantic information f mllm (X) and the second semantic information f mllmAll appear in (X′). The first semantic information f can be removed from mllm (X) the mismatched information "house" and the redundant information "Red, surrounded by trees" to obtain the semantic residual "A barn reflected in a pond".
[0089] In the above manner, on the one hand, redundant information can be identified and removed, and on the other hand, the error information introduced during the semantic extraction process of the compressed image can be removed, so as to ensure the efficiency and accuracy of the extracted semantic residual while avoiding redundancy, and without reducing the perceptual quality during image reconstruction.
[0090] Return for reference Figure 2 , in step S240, the encoded information of the target image can be generated based on the compressed feature and the semantic residual.
[0091] In this step, the compressed feature and the semantic residual can be encoded to obtain the encoded information of the target image, and such encoded information can be used to be sent to the decoding side to achieve image transmission.
[0092] As an example, the compressed feature and the semantic residual can be encoded respectively. For example, as Figure 3 shown, the compressed feature can be encoded by using the feature compression unit to obtain the encoded information of the compressed feature; the semantic residual can be encoded by using the arithmetic coding method to obtain the encoded information of the semantic residual, and the encoded information of the compressed feature and the encoded information of the semantic residual together form the encoded information of the target image.
[0093] According to the image coding method of the embodiments of the present disclosure, a semantic extraction module such as a large language model can be used to, when receiving the original target image and the decoded feature after compression, be able to identify the difference between the original image and the decoded image based on its semantic understanding ability, so as to generate a semantic residual for supplementing the image information. Such a semantic residual can be used to supplement the details lost during the compression process, especially some low-level image details and texture information, to ensure the perceptual consistency and realism of the finally reconstructed image.
[0094] Figure 5 shows an example of the subjective effect of semantic residual extraction of the image coding method according to the embodiments of the present disclosure on the CLIC2020 dataset, where bpp represents the code rate of image transmission. As Figure 5 shown, the image reconstructed based on the compressed feature x′ and the semantic residual (Caption res ) and the semantic information of the compressed feature and the original image (Caption ori)The clarity levels of the reconstructed images are almost the same, while the bitrate of the images reconstructed based on the compressed feature x′ and the semantic residual is significantly lower than that of the images reconstructed based on the compressed feature and the semantic information of the original image. It can be seen that the extraction of semantic residuals greatly improves the detailed information of image reconstruction and effectively reduces the visual distortion caused by compression.
[0095] Figure 6 shows the bitrate saved by semantic residual extraction in the image coding method according to an embodiment of the present disclosure at different bitrates. As Figure 6 shown, the latent feature represents the bitrate allocation ratio of the basic compressed feature, or rather, the proportion of the bitrate of the compressed feature. The text residual (Text(res)) represents the proportion of the bitrate occupied by the semantic residual component. The text bitrate saving (Text(saved)) can quantify the bitrate saving effect achieved by the Semantic residual retrieval (Srr) unit, which can represent the proportion of the text bitrate saved by the semantic residual extraction unit. Comparative analysis Figure 6 of the data in the right figure of shows that with the decrease of the target bitrate, the semantic residual extraction process exhibits a significant bitrate optimization efficiency, plays a great role in bitrate saving, and effectively improves the information compression efficiency under low bitrate conditions.
[0096] According to a second aspect of the embodiments of the present disclosure, an image decoding method is provided. Figure 7 is a flowchart showing an image decoding method according to an exemplary embodiment of the present disclosure. This method can decode an image based on semantic residuals that better represent the information loss caused by image compression. In the case of improving the image transmission efficiency through compression, it can also ensure the quality of image reconstruction.
[0097] As Figure 7 shown, the image decoding method according to an embodiment of the present disclosure may include the following steps:
[0098] In step S710, the encoding information of the target image to be decoded can be obtained.
[0099] In step S720, based on the encoding information, the compressed feature and the semantic residual of the target image can be obtained, where the compressed feature is obtained by compressing the original feature of the target image, and the semantic residual represents the information lost due to compression.
[0100] As an example, the encoded information of the target image may include the encoded information of the compressed feature and the encoded information of the semantic residual, and the decoder may perform a restoration process on the encoded information of the compressed feature and the encoded information of the semantic residual. Specifically, the encoded information of the compressed feature may be decoded for image features to obtain the compressed feature; the encoded information of the semantic residual may be decoded for text to obtain the semantic residual.
[0101] In step S730, a decoded image may be obtained based on the compressed feature and the semantic residual.
[0102] In this step, after obtaining the compressed feature and the semantic residual, the information in the semantic residual may be fused with the compressed feature to obtain a reconstructed image, thereby obtaining the decoded image. As an example, an existing feature fusion method or information fusion method may be adopted to obtain the decoded image based on the compressed feature and the semantic residual.
[0103] As another example, in the exemplary embodiments of the present disclosure, image reconstruction may be achieved by predicting noise for a feature with noise and then removing the noise from the feature. For example, as Figure 8 shown, this step S730 may include the following steps:
[0104] In step S810, noise may be added to the compressed feature to obtain a noisy feature.
[0105] In this step, in order to avoid the large latency and uncertainty brought about by reconstructing from pure noise, noise may be added to the compressed feature, and the noisy feature obtained by adding noise may be used as the denoising starting point, thereby reconstructing the image.
[0106] As an example, a diffusion model may be adopted to implement this reconstruction process. Here, the diffusion process of the diffusion model may be represented, for example, by the following formula (1):
[0107]
[0108] where ∈ t represents random Gaussian noise, z0 represents the initial feature of the image, z t represents the image feature after adding noise, represents the parameter in the noise addition process, t represents the time step, and T represents the maximum number of noise addition steps.
[0109] For the diffusion model, during the training process of the model, the image may be added noise to random Gaussian noise, so that the model can learn the ability to remove noise from the image features with noise; while in the testing or application process, the model can reconstruct the image from random noise.
[0110] For noise characteristics such as the above formula (1), since adding random noise will inevitably introduce uncertainty, which is beneficial in pure image generation. However, such noise is not conducive to image compression that has higher requirements for consistency. Specifically, if the starting point of image generation is a random noise, it is equivalent to reconstructing an image from scratch. In this case, for the image generation task, since there is no accurate target value for the image generation task, it is completely reasonable to use such random noise. However, for the image compression task, the goal is to achieve a reconstruction that is exactly the same as the original image. However, if the image is reconstructed from scratch, some error information will inevitably be generated. In addition, based on the above diffusion mechanism, in the training of image generation, the high computational cost of directly training the diffusion model can be significantly reduced, while enhancing the control ability and flexibility. However, in the field of image compression, the consistency between the generated image and the original image greatly affects the reliability of communication. Although this control network has flexible control ability, its control accuracy is still insufficient, and only relying on the control network cannot generate a reconstruction with high Figure 1 consistency with the original. Therefore, in the embodiments of the present disclosure, a new method for adding noise and denoising reconstruction can be proposed.
[0111] Specifically, in this step S810, noise can be added to the compressed feature to obtain a noise-added feature. Such a noise-added feature takes into account the features of the target image, so that in the subsequent image reconstruction process, it is not reconstructed from scratch, and the introduction of error information can be minimized as much as possible.
[0112] As an example, taking the diffusion model as an example, the noise-added feature according to the embodiments of the present disclosure can be obtained based on a preset noise-adding relationship. The noise-adding relationship can be represented, for example, by the following formula (2):
[0113]
[0114] where represents the noise-added feature, z c represents the compressed feature, represents the preset noise (which can be random noise, for example), represents the parameter at the N r th time step during the noise-adding process, such as the noise coefficient.
[0115] As an example, the above noise-adding relationship can be obtained based on the residual between the original feature of the target image and the compressed feature and the forward diffusion mechanism of the diffusion model. Specifically, the forward diffusion mechanism based on semantic residual according to the embodiments of the present disclosure can be represented by the following formula (3):
[0116]
[0117] where ρ resRepresents the residual between the original features and the compressed features of the target image, ρ res = z c - z0, where z0 is the original feature of the target image, and N r is the maximum number of noise addition steps.
[0118] Based on the above equation (3), if we hope to recover z n from z n-1 , the model needs to fit the posterior probability p θ (z n-1 |z n , z c ). Since the forward diffusion mechanism based on the above equation (3) is non-Markovian, z n-1 and z n are independent of each other. Thus, we can obtain the conditional probability distribution q(z n-1 |z n , z0, z c ) = q(z n-1 |z0, z c ), which represents the conditional distribution of z n given z c , z0, and z n-1 . As an example, we can assume that q(z n-1 |z n , z0, z c ) follows the distribution represented by the following equation (4):
[0119]
[0120] where N represents the normal distribution, l n z n + ζ n z0 represents the mean, represents the variance, and I represents the obtained parameters.
[0121] Thus, we can get z n-1 = l n z n + ζ n z0 + σ n ∈ n . Substituting it into the above equation (3), we can obtain the following equations (5) to (7):
[0122]
[0123] where σ n can be set to η = 0 / 1 corresponding to deterministic sampling and stochastic sampling respectively. Since during the testing and application process, z0 cannot be obtained during the decoding process, we can pass Make the following formula (8) hold:
[0124]
[0125] Thus, the noise-adding relationship as shown in the above formula (2) can be obtained. Here, the noise-adding feature is the end point of noise addition and also the starting point of denoising in testing or application. Such a design makes the end point of noise addition no longer a random Gaussian noise, but a noise with a compressed feature in the mean value, which greatly suppresses the randomness brought by the random Gaussian noise. In addition, such a design makes it unnecessary to encode residual information in testing or application, and only the compressed feature needs to be obtained in the decoding process to perform diffusion generation from the noise-adding feature.
[0126] In step S820, the noise-adding feature can be denoised by predicting the noise in the noise-adding feature under the constraint of the semantic residual to obtain a denoised feature.
[0127] In this step, the process of noise prediction can be constrained based on the semantic residual to obtain the predicted noise, so as to obtain a reconstructed image that better conforms to the original image by removing the noise from the noise-adding feature.
[0128] As an example, the semantic residual as text information can be mapped to an embedding vector through a text encoder such as a Contrastive Language-Image Pre-training (CLIP) model, and this embedding vector can be fused with the image feature through an attention mechanism module, such as Figure 3 shown. This embedding vector acts on the attention mechanism layer in the UNet network of the diffusion model to ensure the tight combination of text information and visual information in the generation process, thereby effectively enhancing the perceptual quality of the generated image. Here, by adopting the fusion mechanism of text information and visual information, it not only helps the generated image to more truly reflect the text prompt, but also enhances the ability to reconstruct the details of the image, avoiding the situation of detail loss or distortion.
[0129] As another example, a control feature can also be introduced into the noise prediction and image reconstruction process of the diffusion model. In this example, the control feature is used to limit the parameters in the process of obtaining the decoded image, and it can be obtained based on the compressed feature. For example, step S820 may include: generating a control feature based on the compressed feature, where the control feature is used to limit the parameters in the process of obtaining the decoded image; denoising the noise-adding feature by predicting the noise in the noise-adding feature under the constraints of the semantic residual and the control feature to obtain a denoised feature.
[0130] Specifically, in this example, the compression features and semantic residuals can be used as control conditions and input into the diffusion model to control its denoising operation. Through multiple-step iteration, the diffusion model gradually reconstructs an image with high perceptual quality. During this process, the advantages of compression features and semantic information can be effectively combined to ensure that the finally generated image has excellent perceptual quality.
[0131] As an example, this image decoding method can be executed by a decoding module, such as Figure 3 the decoding module in [reference], which can include a control module and a generation module. The control module can be used to generate the above control features. The control module is, for example, Figure 3 the Visual adapter in [reference]; the generation module can be used to predict the noise in the noisy features and denoise the noisy features. For example, it can be implemented based on Figure 3 the Conditional Diffusion Model (CDM) in [reference].
[0132] Here, the control module can be based on the design idea of the ControlNet architecture, which aims to achieve flexibility and generation control ability at a relatively low computational cost. However, in the embodiments of the present disclosure, control networks with different structures can be used, and a new diffusion process is proposed for image compression.
[0133] Specifically, different from the traditional ControlNet architecture, in the embodiments of the present disclosure, instead of using an additional convolutional neural network to map the image to the feature space, a pre-trained VAE image codec can be used to map the image to the feature space and compress it through a feature compression module. The compressed features and random noise are fused through a convolutional operation and jointly input into the control network, which jointly acts on the subsequent network layers. The output of the control network is connected to each layer of the UNet network of the diffusion model through a Zero Convolution layer for Skip Connection to assist the noise prediction process of the subsequent network layers.
[0134] As an example, the control module can be obtained based on at least a part of the network structure in the generation module. For example, the control module can have the same structure as the network layer of forward diffusion in the generation module. As an example, when the generation module is implemented based on a diffusion model, the core part of the diffusion model, such as the Encoder Block and Middle Block in UNet, can be copied out as the control module to be optimized (or to be learned) to avoid the high resource consumption brought by the training of the diffusion model. During the optimization process of the embodiments of the present disclosure, joint training and a joint loss function can be introduced to comprehensively consider feature reconstruction, compression bit rate, and the accuracy of noise prediction, thereby comprehensively improving the model performance, which will be described in detail in the third aspect below.
[0135] Through the above design, under the joint action of the control feature and the semantic residual, the image finally generated by the diffusion model can retain the details and structure of the original image to the greatest extent under the condition of an extremely low bit rate, jointly guiding image reconstruction, so that the reconstructed or generated image has a low distortion degree and a high fidelity. Through the collaborative work of such multi-modules, the efficient compression and high-fidelity reconstruction of images at a low bit rate are finally realized. As an example, datasets such as MSCOCO and CLIC2020 can be used to compare the method of the embodiments of the present disclosure with traditional methods, as Figure 9 shown, the decoded image obtained by performing noise prediction and denoising under the constraints of the control feature and the semantic residual has greatly improved the consistency between the reconstruction result and the original image compared with the decoded image obtained by the method without using the constraints.
[0136] In the above example, whether under the constraint of the semantic residual or under the constraints of both the semantic residual and the control feature, the processes of noise prediction and denoising can be performed multiple times. Specifically, in step S820, the denoising can be performed on the noise-added feature for the preset number of times through the preset number of predictions to obtain the denoised feature, where the preset number can be related to the compression degree of the target image.
[0137] Specifically, during the noise prediction process, a reverse sampling mechanism based on the semantic residual according to the embodiments of the present disclosure can be adopted. For example, taking the deterministic sampling with η =0 as an example, then Since Based on the above-obtained parameters, the sampling formula can be obtained as shown in the following formula (9):
[0138]
[0139] where represents the feature of the reconstructed image. Here, in each prediction and denoising, the feature For example, it can be expressed by the following formulas (10) and (11):
[0140]
[0141] Thus, in the reverse sampling, the feature of the reconstructed image for each denoising operation can be obtained through the noise predicted each time Here, z can be obtained from z through the above sampling formula n to get z n-1 , and iterated repeatedly until the final reconstruction.
[0142] In the above process, N r can represent the maximum number of noise addition steps in the noise addition process. Correspondingly, it also represents the number of times of prediction and denoising in the denoising process, that is, the above preset number of times. In the embodiments of the present disclosure, different preset numbers N r can be selected for different compression degrees.
[0143] Specifically, in the reverse sampling process of the diffusion model, the information entropy of each step can be measured by the following formula (12):
[0144] R N ~D KL [q(z N-1 |z N , z0)||p θ (z N-1 |z N )] (12)
[0145] Among them, R N represents the information entropy between the features obtained by two denoising operations; D KL represents the KL divergence, which can measure the difference between two probability distributions; q(z N-1 |z N , z0) represents the conditional distribution of the state at the (N - 1)th step under the condition of knowing the noise state z N at the Nth step and the original data z0; p θ (z N-1 |z N )] represents the conditional distribution of the model predicting the previous step state z N based on the current state z N-1 .
[0146] It can be obtained from the above formula (12) that in the early stage of reverse denoising (the larger N is), the noise intensity of the image is large, the information entropy R N is small, and it contains less information, which exactly corresponds to the case of a large compression degree; on the contrary, in the later stage of reverse denoising (the smaller N is), the image noise intensity is greatly reduced, and R NIncreases, corresponding to a small degree of compression, R N is the information entropy. The smaller the information entropy, the smaller the amount of information. And the greater the degree of compression of image compression, the more blurred the image and the smaller the amount of information, which is opposite to the situation in the early stage of reverse denoising, and vice versa. In this case, if a fixed value is set, it may lead to insufficient information recovery or unnecessary generation freedom, which may affect the consistency of the generated image. Therefore, in the embodiments of the present disclosure, different preset numbers N can be selected according to different degrees of compression r , where the degree of compression can be represented by the compression ratio, and the preset number N corresponding to various degrees of compression is determined in advance r . During the test or application process, according to the currently received and decoded compression features, the degree of compression of the compression features can be determined, so as to perform noise prediction and denoising of the denoising features for the preset number of times corresponding to this degree of compression Figure 10 shows an example of an optimal noise starting point corresponding to each compression ratio, as Figure 10 shown. Different preset numbers will affect the bit rate point (bpp). Therefore, it is necessary to determine the preset number corresponding to the compression ratio to achieve the optimal noise reduction starting point, thereby optimizing the bit rate
[0147] In the embodiments of the present disclosure, the compression features are not only input into the generation module as control conditions, but also used as the starting point of the diffusion process together with the preset noise, and the preset noise can be adjusted according to different degrees of compression, and each compression ratio corresponds to an optimal noise starting point. In this way, the generation behavior of the diffusion process can be effectively controlled, and the quality and consistency of image generation can be guaranteed
[0148] Returning to the reference Figure 8 , in step S830, the decoded image can be obtained based on the denoising features
[0149] Here, in the case of obtaining the denoising features, the denoising features can be transformed or decoded from the feature domain to the pixel domain to obtain the decoded image
[0150] By Figure 8 the process shown, the advantages of compression features and semantic information can be combined to ensure that the finally generated image has excellent perceptual quality
[0151] The specific steps, variant methods and beneficial effects in the image decoding method according to the second aspect of the exemplary embodiments of the present disclosure can be the same as or similar to those in the image encoding method according to the first aspect of the exemplary embodiments of the present disclosure described above, so they will not be repeated here Figure 2
[0152] Figure 11 According to the third aspect of the embodiments of the present disclosure, a training method for an image codec system is provided Figure 11It is a flowchart showing a training method of an image codec system according to an exemplary embodiment of the present disclosure. This method can decode an image based on a semantic residual that better represents the information loss caused by image compression. In the case of improving the image transmission efficiency through compression, it can also ensure the quality of image reconstruction.
[0153] The image codec system according to an embodiment of the present disclosure may include an encoding module and a decoding module. As an example, the image codec system may have an architecture as Figure 3 shown. As Figure 11 shown, the training method of this image codec system may include the following steps:
[0154] In step S1110, a training sample set can be obtained.
[0155] Here, the training sample set may include training image samples. The training image samples may, for example, include an image data set in any format and with any content. The image data set may include multiple training images. As an example, the image data set used for training may be, but is not limited to, the high-resolution LSDIR and Flicker2W data sets.
[0156] In addition, as an example, in this step, the training sample set may also include text prompt words for each training image sample, which are used to provide supervision for semantic extraction during training. Here, the text prompt words may be generated using a multimodal large language model such as Llava. In this way, more abundant training supervision signals can be provided for the model during training.
[0157] In step S1120, the encoding module can be used to compress the original features of the training image samples to obtain the compressed features of the training image samples, and determine the semantic residual based on the training image samples and the training compressed images corresponding to the compressed features.
[0158] Here, the compressed features can be obtained by compressing the features of the training image samples, and the semantic residual can represent the information lost due to compression.
[0159] As an example, as Figure 3 shown, the encoding module may include a feature compression unit and a semantic residual extraction unit. The feature compression unit can receive the training image samples and perform feature compression on them, and can output the compressed features of the training image samples. It can also decode the compressed features to obtain the training compressed images corresponding to the compressed features. The semantic residual extraction unit can receive the training image samples and the training compressed images corresponding to the compressed features, and extract the semantic residual between the two.
[0160] The processes and principles of feature compression and semantic residual extraction, as well as the structural examples of the feature compression unit and the semantic residual extraction unit and the related beneficial effects, have been described in detail in the image coding method of the first aspect above, so they will not be elaborated here.
[0161] In step S1130, an encoding module can be used to generate the encoded information of the training image sample based on the compressed feature and the semantic residual.
[0162] Here, the encoding module can also encode the compressed feature and the semantic residual respectively to obtain their encoded information, and can send the encoded information to the decoding module.
[0163] The example processes and related beneficial effects of encoding the compressed feature and the semantic residual have been described in detail in the image coding method of the first aspect above, so they will not be elaborated here.
[0164] In step S1140, a decoding module can be used to obtain the compressed feature and the semantic residual of the training image sample based on the encoded information, and to obtain the reconstructed image based on the compressed feature and the semantic residual.
[0165] Here, the decoding module can decode the received encoded information to obtain the decoded compressed feature and semantic residual. The decoding module can reconstruct the image based on the compressed feature and the semantic residual. Here, the decoding module can be implemented based on, for example, a diffusion model.
[0166] As an example, the steps of obtaining the reconstructed image based on the compressed feature and the semantic residual may include: adding noise to the compressed feature to obtain a noisy feature; predicting the noise in the noisy feature under the constraint of the semantic residual to denoise the noisy feature and obtain a denoised feature; and obtaining the reconstructed image based on the denoised feature.
[0167] Here, during the training process, when adding noise to the compressed feature, different preset numbers N can be selected according to different compression degrees. r Training is carried out, so that the optimal preset number for different compression degrees can be determined.
[0168] In addition, as an example, the decoding module can include a control module and a generation module. The generation module can be used to predict the noise in the noisy feature and denoise the noisy feature. The control module can generate a control feature, so that the generation module predicts the noise and denoises the noisy feature under the constraints of the semantic residual and the control feature to obtain a denoised feature.
[0169] The example processes and related beneficial effects of predicting the noise in the noisy feature and denoising the noisy feature have been described in detail in the image decoding method of the second aspect above, so they will not be elaborated here.
[0170] In step S1150, the image encoding and decoding system can be trained based on the training image samples and the reconstructed images.
[0171] In this step, the encoding module and the decoding module can be jointly trained. For example, the feature compression unit of the encoding module and the control module of the decoding module can be end-to-end trained, while the remaining modules or units in the encoding module and the decoding module are frozen, that is, the parameters are fixed.
[0172] As an example, as Figure 12 shown, the steps of training the image encoding and decoding system based on the training image samples and the reconstructed images can include:
[0173] In step S1210, based on the difference between the reconstructed image and the training image sample, a first loss related to the encoding module can be determined.
[0174] Here, the first loss can be, for example, the loss corresponding to the training feature compression unit. The first loss can be determined based on the bitrate of the compressed features and the difference between the reconstructed image and the training image sample.
[0175] As an example, the first loss can be represented by the following formula (13):
[0176]
[0177] Here, represents the bitrate of the compressed features (including the hyperprior), D(z0,z c ) represents the reconstruction loss between the input features and the decoded features, and λ R represents the weight parameter for weighing the importance of the bitrate and the reconstruction loss.
[0178] In step S1220, based on the reconstructed image, the training image sample, and the training prediction noise, a second loss related to the decoding module can be determined.
[0179] Here, the second loss can be, for example, the loss corresponding to the training control module.
[0180] In an example, the feature domain loss can be determined based on the training prediction noise, the image features of the reconstructed image, and the image features of the training image sample, and the feature domain loss can be used as the second loss.
[0181] Specifically, after knowing q(z n-1 |z n ,z0,z c ), the prior model can minimize the variational lower bound ELBO to obtain the following loss function of formula (14):
[0182]
[0183] Here, combining the parameters defined in Equations (3) and (4) above, Equation (14) can be written as Equation (15) below:
[0184]
[0185] Substituting Equation (15) above into Equation (10) in the above text, Equation (16) below can be obtained:
[0186]
[0187] Equation (16) above can be simplified to Equation (17) below:
[0188]
[0189] Thus far, the second loss can be expressed by Equation (17) above.
[0190] In addition, as an example, during the training process, in order to stabilize the training, the weight parameter in Equation (17) can also be omitted
[0191] In another example, in addition to the feature domain loss, the pixel domain loss can also be considered, and the combined loss of the two is used as the second loss.
[0192] Specifically, the steps for determining the second loss based on the reconstructed image, the training image samples, and the training prediction noise may include: determining the feature domain loss based on the training prediction noise, the image features of the reconstructed image, and the image features of the training image samples; determining the pixel domain loss based on the difference between the pixel features of the reconstructed image and the pixel features of the training image samples; and determining the second loss based on the feature domain loss and the pixel domain loss.
[0193] Here, the process of determining the feature domain loss is as described in the above example. The pixel domain loss can be determined based on the difference between the pixel features of the reconstructed image and the pixel features of the training image samples. As an example, the difference between the pixel features of the reconstructed image and the pixel features of the training image samples can be determined based on pixel statistical differences such as the mean squared error and image perception evaluation metrics.
[0194] Specifically, since during the sampling process, it is necessary to obtain the features of the reconstructed image through the predicted noise ∈ θ (z t , c, z c , t), such as the decoded features described above When there is a systematic bias in the noise prediction, an error accumulation effect will occur during the iterative sampling process, ultimately leading to a deviation of the generated image from the target distribution. Therefore, in order to further improve the predicted decoded features For accuracy, it can be decoded using a decoder, and a supervision constraint is constructed in the pixel domain.
[0195] Specifically, the supervision constraint may include a perceptual similarity constraint and / or a pixel fidelity constraint. Here, the perceptual similarity constraint can be implemented based on, for example, pixel statistical differences, and the pixel fidelity constraint can be implemented based on, for example, an image perception evaluation metric. The pixel statistical difference can be, for example, the Mean Squared Error (MSE), and the MSE can ensure the accurate reconstruction of local textures. The image perception evaluation metric can be, for example, the Learned Perceptual Image Patch Similarity (LPIPS), and the LPIPS can measure the high-dimensional feature matching degree between the generated image and the reference image. Although the MSE and LPIPS are used as examples here to describe the pixel statistical difference and the image perception evaluation metric, the embodiments of the present disclosure are not limited thereto, and other pixel statistical differences and / or image perception evaluation metrics can also be used.
[0196] As an example, in the case of constructing a dual supervision constraint in the pixel domain, the pixel domain loss can be expressed as: where, ∈ θ is the noise predicted by the diffusion model, x0 is the original training image sample, is the feature z denoised by the diffusion model from the t-th step t The directly predicted feature of the reconstructed image Decoded to the pixel domain, λ d and λ p represent preset weight parameters, which can be set according to actual needs.
[0197] In this example, combined with the feature domain loss described above, the second loss can be expressed by the following formula (18):
[0198]
[0199] In step S1230, the image codec system can be trained based on the first loss and the second loss.
[0200] In this step, the sum of the first loss and the second loss can be used as the total training loss. Here, taking the above formulas (13) and (18) as examples, the total training loss can be expressed by the following formula (19):
[0201]
[0202] As an example, substituting each loss term into the above formula (19), the total training loss can be expressed as:
[0203]
[0204] Here, through the supervision in the pixel domain, the accuracy of denoising at each step during the test process can be effectively enhanced, and finally the perceptual quality of the reconstruction result can be enhanced. In this way, the demand for training resources of the diffusion model can be effectively reduced, and the performance of the feature compression module and the control module can be improved in the end-to-end optimization.
[0205] In addition, it should be noted that although the calculation expressions of the respective parameters are described by way of example in the above text using Equations (1) to (19), the embodiments of the present disclosure are not limited thereto. Based on the principle of the embodiments of the present disclosure, the above expressions can be adjusted or modified. For example, the coefficients can be weighted, the operation relationship between the terms can be adjusted, and the like.
[0206] By adjusting the learnable parameters in the image codec system, the training loss is optimized or reduced, thereby realizing the training of the image codec system until a preset training stop condition is reached. For example, when a preset number of training rounds or a preset training loss is reached, the training of the image codec system can be completed.
[0207] The encoding module in the image codec system trained by the training method of the image codec system according to the embodiments of the present disclosure can be deployed on the encoding side and / or the decoding module in the image codec system can be deployed on the decoding side.
[0208] According to the image codec method and the training method of the image codec system of the embodiments of the present disclosure, the semantic residual information between the uncompressed image and its compressed representation can be extracted, thereby reducing semantic redundancy, ensuring that the reconstructed image at an extremely low bit rate reaches the highest fidelity, and the optimized semantic residual and the compressed latent features can be integrated into the diffusion model, thereby realizing high-quality visual reconstruction at an extremely low bit rate. Here, by fusing the residuals into semantic extraction and diffusion generation, the bit rate consumption of text information can be effectively reduced, the quality of the reconstructed image can be improved, and at the same time, the denoising delay of the diffusion model can be greatly reduced.
[0209] In addition, according to the image encoding and decoding method and the training method of the image encoding and decoding system according to the embodiments of the present disclosure, for the encoded text, the semantic residual information of the original image and the compressed feature map can be extracted through a multimodal large language model; for the encoded features, the image can be mapped from the pixel domain to the feature domain and compressed and encoded; during decoding, the compressed features and the text can be decoded first, and then the compressed features after adding a set noise perturbation are denoised under the guidance of these two control conditions to generate a high-quality and highly reliable reconstructed image. Such a method effectively reduces the information redundancy phenomenon of the semantic prompt words in the multimodal image encoding method, while greatly reducing the decoding delay, and exhibits extremely high perceptual quality in each bitrate segment (especially at an extremely low bitrate (<0.005 bpp)).
[0210] According to a fourth aspect of the present disclosure, there is provided an image encoding device, as Figure 13 shown, the image encoding device 1300 may include an acquisition unit 1310, a compression unit 1320, a determination unit 1330, and a generation unit 1340.
[0211] The acquisition unit 1310 is configured to acquire a target image to be encoded.
[0212] The compression unit 1320 is configured to compress the original features of the target image to obtain the compressed features of the target image.
[0213] The determination unit 1330 is configured to determine a semantic residual based on the target image and the compressed image corresponding to the compressed features, where the semantic residual represents the information lost due to compression.
[0214] The generation unit 1340 is configured to generate the encoded information of the target image based on the compressed features and the semantic residual.
[0215] As an example, the determination unit 1330 is configured to: perform semantic extraction on the target image to obtain first semantic information; perform semantic extraction on the compressed image to obtain second semantic information; and determine the semantic residual based on the difference between the first semantic information and the second semantic information.
[0216] As an example, the determination unit 1330 is configured to: determine the redundant information in the first semantic information and the second semantic information; determine the mismatched information in the second semantic information that does not appear in the first semantic information; and determine the semantic residual based on the first semantic information, the redundant information, and the mismatched information.
[0217] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0218] According to a fifth aspect of the present disclosure, there is provided an image decoding apparatus. As shown in Figure 14 FIG. 1400, the image decoding apparatus may include an information acquisition unit 1410, an information determination unit 1420, and an image determination unit 1430.
[0219] The information acquisition unit 1410 is configured to acquire encoding information of a target image to be decoded.
[0220] The information determination unit 1420 is configured to obtain a compression feature and a semantic residual of the target image based on the encoding information, where the compression feature is obtained by compressing an original feature of the target image, and the semantic residual represents information lost due to compression.
[0221] The image determination unit 1430 is configured to obtain a decoded image based on the compression feature and the semantic residual.
[0222] As an example, the image determination unit 1430 is configured to: add noise to the compression feature to obtain a noisy feature; denoise the noisy feature by predicting noise in the noisy feature under the constraint of the semantic residual to obtain a denoised feature; and obtain a decoded image based on the denoised feature.
[0223] As an example, the image determination unit 1430 is configured to: generate a control feature based on the compression feature, where the control feature is used to limit parameters in the process of obtaining a decoded image; and denoise the noisy feature by predicting noise in the noisy feature under the constraints of the semantic residual and the control feature to obtain a denoised feature.
[0224] As an example, the image determination unit 1430 is configured to: perform denoising on the noisy feature a preset number of times through a preset number of predictions to obtain a denoised feature, where the preset number of times is related to the compression degree of the target image.
[0225] As an example, the image decoding apparatus is implemented by a decoding module, and the decoding module includes a control module and a generation module. The control module is used to generate a control feature, and the generation module is used to predict noise in the noisy feature and denoise the noisy feature, where the control module is obtained based on at least a part of the network structure in the generation module.
[0226] Regarding the apparatus in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0227] According to a sixth aspect of the present disclosure, there is provided a training apparatus for an image codec system. The image codec system includes an encoding module and a decoding module. As shown in Figure 15As shown, the training apparatus 1500 of the image encoding and decoding system may include a training acquisition unit 1510, a training compression unit 1520, a training generation unit 1530, a training reconstruction unit 1540, and a training unit 1550.
[0228] The training acquisition unit 1510 is configured to acquire a training sample set, where the training sample set includes training image samples.
[0229] The training compression unit 1520 is configured to use an encoding module to compress the original features of the training image samples to obtain compressed features of the training image samples, and determine a semantic residual based on the training image samples and the training compressed images corresponding to the compressed features, where the compressed features are obtained by compressing the features of the training image samples, and the semantic residual represents the information lost due to compression.
[0230] The training generation unit 1530 is configured to use an encoding module to generate encoded information of the training image samples based on the compressed features and the semantic residual.
[0231] The training reconstruction unit 1540 is configured to use a decoding module to obtain the compressed features and the semantic residual of the training image samples based on the encoded information, and obtain a reconstructed image based on the compressed features and the semantic residual.
[0232] The training unit 1550 is configured to train the image encoding and decoding system based on the training image samples and the reconstructed image.
[0233] As an example, the training reconstruction unit 1540 is configured to: add noise to the compressed features to obtain noisy features; denoise the noisy features by predicting the noise in the noisy features under the constraint of the semantic residual to obtain denoised features; and obtain a reconstructed image based on the denoised features.
[0234] As an example, the training unit 1550 is configured to: determine a first loss related to the encoding module based on the difference between the reconstructed image and the training image samples; determine a second loss related to the decoding module based on the reconstructed image, the training image samples, and the training prediction noise; and train the image encoding and decoding system based on the first loss and the second loss.
[0235] As an example, the training unit 1550 is configured to: determine a feature domain loss based on the training prediction noise, the image features of the reconstructed image, and the image features of the training image samples; determine a pixel domain loss based on the difference between the pixel features of the reconstructed image and the pixel features of the training image samples; and determine the second loss based on the feature domain loss and the pixel domain loss.
[0236] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0237] Figure 16 is a block diagram of an electronic device shown according to an exemplary embodiment. As Figure 16 shown, the electronic device 1600 includes a processor 1610 and a memory 1620 for storing processor-executable instructions. Here, when the processor-executable instructions are run by the processor, the processor executes the image encoding method, the image decoding method, or the training method of the image codec system as described in the above exemplary embodiments.
[0238] As an example, the electronic device 1600 does not have to be a single device, and can also be any aggregate of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 1600 can also be a part of an integrated control system or a system manager, or can be configured as a server that is interconnected locally or remotely (e.g., via wireless transmission).
[0239] In the electronic device 1600, the processor 1610 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 1610 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0240] The processor 1610 can run the instructions or code stored in the memory 1620, where the memory 1620 can also store data. The instructions and data can also be sent and received via a network interface device over a network, where the network interface device can employ any known transmission protocol.
[0241] The memory 1620 can be integrated with the processor 1610, for example, by arranging RAM or flash memory within an integrated circuit microprocessor, etc. In addition, the memory 1620 can include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 1620 and the processor 1610 can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., such that the processor 1610 can read the files stored in the memory 1620.
[0242] In addition, the electronic device 1600 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 1600 can be connected to each other via a bus and / or a network.
[0243] In an exemplary embodiment, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image encoding method, the image decoding method, or the training method of the image encoding and decoding system as described in the above exemplary embodiment. The computer-readable storage medium may be, for example, a memory including instructions. Optionally, the computer-readable storage medium may be: read-only memory (ROM), random access memory (RAM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), card-type memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0244] In an exemplary embodiment, a computer program product may also be provided. The computer program product includes computer instructions that, when executed by a processor, implement the image encoding method, the image decoding method, or the training method of the image encoding and decoding system as described in the above exemplary embodiment.
[0245] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art from the above description and the associated drawings.
[0246] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is merely illustrative, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but may be changed according to the actual situation. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined or deleted according to actual needs.
[0247] The examples are selected and described to explain the principles of the present disclosure and to enable other technicians in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and the various embodiments with various modifications suitable for the intended specific uses. Therefore, it will be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the present disclosure.
[0248] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0249] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image encoding method, characterized in that, The described image encoding method includes: Obtaining a target image to be encoded; Compressing the original features of the target image to obtain the compressed features of the target image; Determining a semantic residual based on the target image and a compressed image corresponding to the compressed features, where the semantic residual represents the information lost due to the compression; Generating the encoded information of the target image based on the compressed features and the semantic residual.
2. The image encoding method according to claim 1, wherein The determining the semantic residual based on the target image and a compressed image corresponding to the compressed features includes: Performing semantic extraction on the target image to obtain first semantic information; Performing semantic extraction on the compressed image to obtain second semantic information; Determining the semantic residual based on the difference between the first semantic information and the second semantic information.
3. The image encoding method according to claim 2, characterized in that, The determining the semantic residual based on the difference between the first semantic information and the second semantic information includes: Determining the redundant information in the first semantic information and the second semantic information; Determining the mismatched information in the second semantic information that does not appear in the first semantic information; Determining the semantic residual based on the first semantic information, the redundant information, and the mismatched information.
4. An image decoding method, characterized in that, The described image decoding method includes: Obtaining the encoded information of a target image to be decoded; Obtaining the compressed features and the semantic residual of the target image based on the encoded information, where the compressed features are obtained by compressing the original features of the target image, and the semantic residual represents the information lost due to the compression; Obtaining a decoded image based on the compressed features and the semantic residual.
5. The image decoding method according to claim 4, wherein The obtaining a decoded image based on the compressed features and the semantic residual includes: Adding noise to the compressed features to obtain noise-added features; Denosing the noise-added features by predicting the noise in the noise-added features under the constraint of the semantic residual to obtain denoised features; Obtaining the decoded image based on the denoised features.
6. The image decoding method according to claim 5, wherein The denosing the noise-added features by predicting the noise in the noise-added features under the constraint of the semantic residual to obtain denoised features includes: Generating control features based on the compressed features, where the control features are used to limit the parameters in the process of obtaining the decoded image; Denosing the noise-added features by predicting the noise in the noise-added features under the constraints of the semantic residual and the control features to obtain denoised features.
7. The image decoding method according to claim 5 or 6, characterized in that, The denosing the noise-added features by predicting the noise in the noise-added features to obtain denoised features includes: Performing the denoising on the noise-added features for a preset number of times through the preset number of predictions to obtain denoised features, where the preset number of times is related to the compression degree of the target image.
8. The image decoding method according to claim 6, characterized in that The image decoding method is executed by a decoding module, and the decoding module includes a control module and a generation module. The control module is used to generate the control features, and the generation module is used to predict the noise in the noise-added features and denoise the noise-added features, where the control module is obtained based on at least a part of the network structure in the generation module.
9. A training method for an image encoding and decoding system, characterized in that, The image encoding and decoding system includes an encoding module and a decoding module, and the training method includes: Obtain a training sample set, where the training sample set includes training image samples; Use the encoding module to compress the original features of the training image samples to obtain the compressed features of the training image samples, and determine a semantic residual based on the training image samples and the training compressed images corresponding to the compressed features, where the compressed features are obtained by compressing the features of the training image samples, and the semantic residual represents the information lost due to the compression; Use the encoding module to generate the encoding information of the training image samples based on the compressed features and the semantic residual; Use the decoding module to obtain the compressed features and semantic residual of the training image samples based on the encoding information, and obtain a reconstructed image based on the compressed features and the semantic residual; Train the image encoding and decoding system based on the training image samples and the reconstructed image.
10. The training method according to claim 9, wherein The obtaining the reconstructed image based on the compressed features and the semantic residual includes: Add noise to the compressed features to obtain noise-added features; Under the constraint of the semantic residual, denoise the noise-added features by predicting the noise in the noise-added features to obtain denoised features; Obtain the reconstructed image based on the denoised features.
11. The training method according to claim 9, characterized in that The training the image encoding and decoding system based on the training image samples and the reconstructed image includes: Determine a first loss related to the encoding module based on the difference between the reconstructed image and the training image samples; Determine a second loss related to the decoding module based on the reconstructed image, the training image samples, and the training predicted noise; Train the image encoding and decoding system based on the first loss and the second loss.
12. The training method according to claim 11, wherein The determining the second loss based on the reconstructed image, the training image samples, and the training predicted noise includes: Determine a feature domain loss based on the training predicted noise, the image features of the reconstructed image, and the image features of the training image samples; Determine a pixel domain loss based on the difference between the pixel features of the reconstructed image and the pixel features of the training image samples; Determine the second loss based on the feature domain loss and the pixel domain loss.
13. An image encoding device, characterized in that, The image encoding device includes: An obtaining unit configured to obtain a target image to be encoded; A compression unit configured to compress the original features of the target image to obtain the compressed features of the target image; A determining unit configured to determine a semantic residual based on the target image and the compressed image corresponding to the compressed features, where the semantic residual represents the information lost due to the compression; A generating unit configured to generate the encoding information of the target image based on the compressed features and the semantic residual.
14. An image decoding apparatus, characterized in that, The image decoding device includes: An information obtaining unit configured to obtain the encoding information of a target image to be decoded; An information determination unit, configured to obtain a compression feature and a semantic residual of the target image based on the encoded information, where the compression feature is obtained by compressing an original feature of the target image, and the semantic residual represents information lost due to the compression; An image determination unit, configured to obtain a decoded image based on the compression feature and the semantic residual.
15. A training device for an image encoding and decoding system, characterized in that, The image encoding and decoding system includes an encoding module and a decoding module, and the training device includes: A training acquisition unit, configured to acquire a training sample set, where the training sample set includes training image samples; A training compression unit, configured to use the encoding module to compress an original feature of the training image samples to obtain a compression feature of the training image samples, and determine a semantic residual based on the training image samples and training compressed images corresponding to the compression features, where the compression feature is obtained by compressing the features of the training image samples, and the semantic residual represents information lost due to the compression; A training generation unit, configured to use the encoding module to generate encoded information of the training image samples based on the compression features and the semantic residuals; A training reconstruction unit, configured to use the decoding module to obtain a compression feature and a semantic residual of the training image samples based on the encoded information, and obtain a reconstructed image based on the compression feature and the semantic residual; A training unit, configured to train the image encoding and decoding system based on the training image samples and the reconstructed images.
16. An electronic device, characterized in that, Comprising: A processor; And A memory for storing processor-executable instructions, wherein when the processor-executable instructions are run by the processor, the processor is caused to execute the image encoding method according to any one of claims 1 to 3, or the image decoding method according to any one of claims 4 to 8, or the training method of the image encoding and decoding system according to any one of claims 9 to 12.
17. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image encoding method according to any one of claims 1 to 3, or the image decoding method according to any one of claims 4 to 8, or the training method of the image encoding and decoding system according to any one of claims 9 to 12.
18. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the image encoding method according to any one of claims 1 to 3, or the image decoding method according to any one of claims 4 to 8, or the training method of the image encoding and decoding system according to any one of claims 9 to 12 is implemented.