Learning-type image compression with masked visual language modeling
By combining sparse visual and textual features with masked visual language modeling (MAVLM), the limitations of existing image compression technologies in terms of reconstruction quality and flexibility are overcome, achieving efficient and flexible image compression suitable for various application scenarios.
Patent Information
- Application Number
- CN202480054065.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-29
- Filing Date
- 2024-08-02
- Publication Date
- 2026-03-20
AI Technical Summary
Existing image compression technologies have limitations in maintaining high reconstruction quality and flexible compression performance, making it difficult to simultaneously meet the needs of different compression objectives and machine analysis tasks.
The Masked Visual Language Modeling (MAVLM) method is adopted to achieve efficient image compression and flexible adjustment of bit rate and reconstruction quality by combining the encoding of sparse visual and textual features with diffuse latent features.
It achieves high compression efficiency and high reconstruction quality image compression, adapts to different compression needs, improves the flexibility and versatility of image compression, and is suitable for a variety of application scenarios.
Smart Images

Figure CN121713482A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 579,465, filed August 29, 2023. The entire disclosure of the above application is incorporated herein by reference. Technical Field
[0002] This invention relates to Learned Image Compression (LIC), and more particularly to LIC using masked visual language modeling for image compression. Background Technology
[0003] Artificial intelligence generated content (AIGC) utilizes a wide range of image generation models, including generative adversarial networks (GANs), diffusion models, and autoregressive (AR) models. The goal is to achieve fast, accessible, high-quality content creation. Various methods have been developed to efficiently manipulate generated content using different types of inputs, such as textual descriptions and / or spatial / spatiotemporal combinations like sketches or segmentation maps.
[0004] Large-scale pre-trained visual-language models (VLMs) have achieved a landmark breakthrough in text-to-image generation at AIGC. By training very large models using very large datasets of described images from the internet, multimodal language-image pre-trained representations can be successfully learned through self-supervised contrastive learning, such as Contrastive Language-Image Pre-training (CLIP) or Bootstrapping Language-Image Pre-training (BLIP). The joint embedding space of text and images is robust to image distribution shifts, enabling language-guided zero-shot image generation. Summary of the Invention
[0005] The first aspect relates to a method for implementing an encoder. The method includes: encoding an original image into sparse visual features; generating sparse text features indicating the content of the original image; calculating control latent features based on a control signal and the original image; calculating adjusted masked sparse visual features, adjusted masked sparse text features, and adjusted masked control latent features based on the sparse visual features, the sparse text features, and the control latent features, respectively; calculating diffusion latent features; and transmitting the adjusted masked sparse visual features, the adjusted masked sparse text features, the adjusted masked control latent features, and the diffusion latent features to a decoder.
[0006] Optionally, in the first implementation according to any one of the first aspects or any implementation thereof, the method further includes: calculating the diffusion latent features based on the original image, the adjusted masked sparse visual features, and the adjusted masked controlled latent features, the diffusion latent features capturing the fidelity and expressive details of the original image.
[0007] Optionally, in a second implementation according to any one of the first aspects or any implementation thereof, the control latent feature indicates the encoded control requirement.
[0008] Optionally, in a third implementation according to any one of the first aspects or any implementation thereof, calculating the control latent feature includes: calculating a text instruction based on the control signal, the text instruction being a text description of the control requirements of the control signal; calculating an input-oriented text instruction and an additional input-oriented prompt instruction based on the text instruction, the original image, the control signal, and the prompt VLM; and calculating the control latent feature based on the input-oriented text instruction and the additional input-oriented prompt instruction.
[0009] Optionally, in a fourth implementation according to any one of the first aspects or any implementation thereof, the original image has a shape The general three-dimensional (3D) tensor, in which w, h, c The image's width, height, and number of channels are used to encode the original image into the sparse visual features, which includes: encoding the original image into a shape... The visual feature tensor, where the width and height Depending on the width and height of the original image, d It is the number of feature channels; sparse visual features are calculated based on the visual feature tensor and the visual codebook, wherein the visual codebook... It includes multiple codewords, each codeword having d dimension.
[0010] Optionally, in a fifth implementation according to any one of the first aspects or any implementation thereof, encoding the original image into the visual feature tensor includes: dividing the original image into multiple image blocks using a visual-language (VL) transformer, and encoding the image blocks into a sequence.
[0011] Optionally, in a sixth implementation according to any one of the first aspects or any implementation thereof, encoding the original image into the visual feature tensor includes: encoding the original image into the entire image in parallel using a convolutional neural network (CNN).
[0012] Optionally, in the seventh implementation according to any one of the first aspect or any implementation thereof, generating the sparse text features includes: calculating text latent features including the text description based on the original image, the text description describing the content of the original image; and generating the sparse text features based on the text latent features.
[0013] Optionally, in the eighth implementation according to any one of the first aspect or any implementation thereof, calculating the text latent features includes: calculating enhanced text latent features; calculating the text latent features based on the enhanced text latent features and / or a text codebook, wherein the text codebook includes a plurality of codewords, each codeword having a natural language element.
[0014] Optionally, in the ninth implementation according to any one of the first aspect or any implementation thereof, calculating the adjusted masked sparse text features based on the sparse text features includes: calculating the masked sparse text features based on the sparse visual features by applying a text mask to the sparse text features to change the corresponding tokens in the sparse text features according to the text mask, wherein the text mask is a binary mask having the same shape as the sparse text features, and the tokens having a value of 0 are set to a specified value indicating removal; and calculating the adjusted masked sparse text features based on the masked sparse text features and using a visual-language (VL) transformer.
[0015] Optionally, in the tenth implementation according to any one of the first aspect or any implementation thereof, calculating the adjusted masked sparse text features based on the sparse text features includes: calculating the masked sparse text features based on the sparse visual features by applying a text mask to the sparse text features to change the corresponding tokens in the sparse text features according to the text mask, wherein the text mask is a binary mask having the same shape as the sparse text features, and the tokens having a value of 0 are set to a specified value indicating removal; and calculating the adjusted masked sparse text features based on the masked sparse text features and using a visual-language (VL) transformer.
[0016] Optionally, in the eleventh implementation according to any one of the first aspects or any implementation thereof, the method further includes: calculating the adjusted masked control potential by applying a visual mask to the control potential based on the control potential features.
[0017] Optionally, in the twelfth implementation according to any one of the first aspects or any implementation thereof, the modified masked control may include modified masked input-oriented text instructions and modified masked appended input-oriented prompt instructions.
[0018] Optionally, in a thirteenth implementation according to any one of the first aspects or any implementation thereof, calculating the diffusion potential features includes: downsampling the original image to a smaller resolution image; and encoding the smaller resolution image using information from the adjusted masked sparse visual features to obtain the diffusion potential features.
[0019] The second aspect relates to a method for implementing a decoder. The method includes: receiving visual language latent features of an original image and diffusion latent features of the original image, the visual language latent features including text and integers; calculating decoded visual language features based on the visual language latent features; reconstructing a baseline image output based on the decoded visual language features; calculating a supplementary output based on the diffusion latent features and the baseline image output; and constructing a final decoded image output based on the supplementary output and the baseline image output.
[0020] The third aspect relates to a method for implementing a decoder. The method includes: receiving adjusted masked sparse visual features, adjusted masked sparse text features, adjusted masked control latent features, and diffusion latent features of an original image; generating restored masked visual features based on the received adjusted masked sparse visual features; calculating encoded masked text features based on the received adjusted masked sparse text features; calculating encoded masked control features based on the adjusted masked control latent features, the restored sparse visual features, and the encoded masked sparse text features; reconstructing a baseline image output based on the restored masked visual features, the encoded masked text features, and the encoded masked control features; calculating a supplementary output based on the diffusion latent features, the masked sparse visual features, and the masked sparse text features; and constructing a final decoded image output based on the supplementary output and the baseline image output.
[0021] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, the recovered masked visual features are based on a visual codebook, and the encoded masked sparse text features are based on a text codebook.
[0022] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, calculating the supplementary output includes: recovering the reconstructed image based on the diffusion latent features; calculating the embedded latent features based on the reconstructed image; and calculating the supplementary output based on the embedded latent features and using the baseline image output as a diffusion condition.
[0023] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, calculating the supplementary output includes: recovering the reconstructed image based on the diffuse latent features; calculating embedded latent features based on the reconstructed image; and using the recovered masked visual features based on the embedded latent features. The supplementary output is calculated using the encoded masked text features as diffusion conditions.
[0024] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, calculating the supplementary output includes: calculating embedded latent features based on the diffusion latent features; and calculating the supplementary output based on the embedded latent features and using the baseline image output as a diffusion condition.
[0025] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, calculating the supplementary output includes: calculating the embedded latent features based on the diffusion latent features; and calculating the supplementary output based on the embedded latent features and using the recovered masked visual features and the encoded masked text features as diffusion conditions.
[0026] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, the method further includes: using a denoising diffusion probabilistic model (DDPM) or a denoising diffusion implicit model (DDIM) to calculate the supplementary output.
[0027] Optionally, in the first implementation according to any of the foregoing aspects or any of its implementations, calculating the supplementary output includes: when using the baseline image output as the diffusion condition, encoding the baseline image output from the pixel domain to the latent domain using an embedding network.
[0028] Optionally, in the first implementation according to any of the foregoing aspects or any implementation thereof, calculating the supplementary output includes: when using the recovered masked visual features When the encoded masked text features are used as the diffusion condition, a transformation network is used to transform the combination result into a dimension corresponding to the embedded latent features.
[0029] The fourth aspect relates to an encoder, including a memory or storage device for storing instructions; one or more processors or processing devices coupled to the memory or storage device and configured to execute the instructions to cause the encoder to perform a method according to any of the foregoing aspects or any implementation thereof.
[0030] The fifth aspect relates to a decoder, including a memory or storage device for storing instructions; and one or more processors or processing devices coupled to the memory or storage device and configured to execute the instructions to cause the decoder to perform a method according to any of the foregoing aspects or any implementation thereof.
[0031] The sixth aspect relates to a computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium, which, when executed by a processor of a device, cause the device to perform a method according to any of the foregoing aspects or any implementation thereof.
[0032] For clarity, any of the above embodiments can be combined with any one or more of the other embodiments described above to create new embodiments within the scope of the present invention.
[0033] These and other features and advantages will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and claims. Attached Figure Description
[0034] To gain a more complete understanding of the invention, reference is made to the following brief description in conjunction with the accompanying drawings and specific embodiments, wherein the same reference numerals denote the same parts.
[0035] Figure 1 A diagram illustrating the general processing flow of AIGC is shown. Figure 2 A diagram illustrating the general framework for LIC is shown; Figure 3 A diagram illustrating a general framework for machine-oriented compression (CfM) is shown. Figure 4 A diagram illustrating the general framework of Learned Sparse Image Representation (LSIR) is shown. Figure 5A An encoding / decoding framework provided by one embodiment of the present invention is shown; Figure 5B An encoding / decoding framework provided by one embodiment of the present invention is shown; Figure 6 The detailed workflow of the main branch provided by one embodiment of the present invention is illustrated; Figures 7A to 7C The detailed workflow of a text code generation module provided in one embodiment of the present invention is illustrated; Figure 8 The detailed operation of a VL converter module provided in one embodiment of the present invention is illustrated. Figure 9 The detailed workflow of a reconstruction module provided in one embodiment of the present invention is illustrated; Figure 10 The processing flow of a control branch provided by one embodiment of the present invention is illustrated; Figure 11 The processing flow of a control branch provided by one embodiment of the present invention is illustrated; Figure 12 The processing flow of a control adjustment module provided in one embodiment of the present invention is illustrated; Figure 13AThe flowchart for processing diffusion branches provided by one embodiment of the present invention is illustrated; Figure 13B The flowchart for processing diffusion branches provided by one embodiment of the present invention is illustrated; Figure 14A and Figure 14B The reverse diffusion modules provided in two embodiments of the present invention are shown; Figure 15 A diagram of an apparatus provided according to an embodiment of the present invention is shown. Detailed Implementation
[0036] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0037] This paper discloses various systems and methods for encoding and decoding images. The present invention proposes a general framework for image compression using masked visual language modeling. Based on various neural network models, including multimodal visual language models, large language models, diffusion models, masked encoders, and query transformers, embodiments of the present invention provide high compression ratios and high reconstruction quality, wherein the bit rate and reconstruction quality can be flexibly adjusted according to different compression requirements.
[0038] In practical applications, the disclosed technology enhances real-world applications by significantly improving encoder functionality. For example, in the telecommunications and digital media sectors, this optimizes data transmission and streaming by efficiently compressing high-resolution images, thereby reducing bandwidth usage while maintaining quality. In healthcare, this facilitates faster and more efficient storage and transmission of medical images, crucial for remote consultations and diagnoses. Furthermore, in consumer electronics, this increases storage capacity in devices such as smartphones and cameras, enabling users to save more high-quality images. Additionally, for cloud storage and autonomous vehicles, this can reduce storage costs and enhance real-time image processing capabilities. Overall, the improved encoder functionality provides high compression rates, maintains reconstruction quality, and offers flexible bitrate adjustments, making it a valuable tool for efficient data management and enhanced user experience across various applications.
[0039] Figure 1 Figure 100 illustrates a general processing flow for AIGC. (Input prompt) y The prompt is transmitted via encoder 102. Prompt input. yThis can be text input, which provides a description of the content to be generated using AI. In some embodiments, input prompts are provided. y It can also include images related to the AI content to be generated. The prompt encoder 102 is used to input prompts. y The components are encoded in a format that the multimodal embedding network 104 can understand and process. In one embodiment, the prompt encoder 102 is used to input prompts. y Generate hints embedding features This represents the prompt input encoded in the input format of the multimodal embedding network 104. y Hints at embedded features Capture input prompts y The semantic meaning and contextual information enable the multimodal embedding network 104 to effectively understand and process the input. 。 The multimodal embedding network 104 is a neural network architecture designed to merge information from multiple modalities, such as text, images, audio, or other types of data. In one embodiment, the multimodal embedding network 104 is used to compute pairs of... Image embedding features modeled based on priors In one embodiment, image embedding features It is the numerical representation of an image in a high-dimensional vector space. Image embedding features It is passed to decoder 106. Decoder 106 is used for image embedding features. and hints embedded features To calculate the output image The decoding neural network aims to achieve the generated image. High visual perception quality (e.g., natural and photorealistic, with low levels of visible artifacts), and Input prompts y Semantic alignment of the described requirements.
[0040] Figure 2Figure 200 illustrates a general framework for LIC. LIC is a modern image compression method that utilizes deep learning techniques to learn efficient representations of images. Traditional image compression techniques rely on manually designed algorithms that transform image data into compressed formats. However, LIC aims to improve compression efficiency by automatically learning the most efficient compression strategies directly from data through training neural networks. Neural network (NN)-based LIC has been extensively studied in recent years and has demonstrated superior performance compared to traditional coding methods such as Joint Photographic Experts Group (JPEG), Versatile Video Coding (VVC), and High Efficiency Video Coding (HEVC). In the depicted embodiment, at the transmitting end, the input image... x Image embedding features are generated by input encoder 202. The image embedding feature is the input image. x The numerical format representation. In one embodiment, the input encoder 202 is used to input the image. x The original pixel values are converted into a neural network that uses compressed and semantically meaningful numerical representations in a high-dimensional vector space. In some embodiments, image embedding features... It is further compressed into a data string through quantization and arithmetic encoding, which is efficient for storage and transmission from sender to receiver.
[0041] In one embodiment, at the receiving end, arithmetic decoding and inverse quantization are used to recover the decoded image embedding features from the received data string sent by the sender. Decoded image embedding features Used as input to decoder 204. Decoder 204 is used to embed features based on the decoded image. To reconstruct the output image The goal is to minimize the reconstructed output. and original input x The recovery loss between them is minimized, and bits are used to represent the image embedding features. For use in storage and transmission.
[0042] Traditionally, compression methods were developed for human consumption. That is, the reconstructed output... The goal is to make it suitable for human viewing in order to maintain high visual quality. Compression artifacts can significantly degrade the performance of some machine analysis tasks, such as detection and recognition, because the information required for these tasks may be altered or lost during compression. To promote machine analysis, standards initiatives such as the Moving Picture Experts Group (MPEG) Video Coding for Machine (VCM) and JPEG-AI have been launched to investigate compression methods suitable for machine analysis tasks.
[0043] Figure 3 Figure 300 illustrates a general framework for CfM. In Figure 3 In the process, a machine-oriented preprocessing module 302 and / or a machine-oriented postprocessing module 304 are used before and / or after a video compression method (e.g., VVC, HEVC, LIC, etc.) to process the input image before the input encoder 202. x Preprocessing (e.g.) Figure 2 The reconstructed output after decoder 204 (as described above) and / or after decoder 204 Post-processing (such as...) Figure 2 The target machine analysis task model 306 is used for connection. In one embodiment, the machine-oriented preprocessing module 302 and / or the machine-oriented postprocessing module 304 are optimized in an end-to-end manner using the machine analysis task model by using task performance loss (while keeping the video compression method and the machine analysis task model unchanged). In one embodiment, a set of machine-oriented preprocessing modules 302 and / or machine-oriented postprocessing modules 304 are used for each specific task model for each machine analysis task.
[0044] Figure 4 Figure 400 illustrates a general framework for Learned Sparse Image Representation (LSIR). In LSIR, a vector quantization autoencoder in the image domain is trained to learn a highly compressed codebook 402 based on adversarial and perceptual losses (e.g., using a Vector Quantized Generative Adversarial Network (VQGAN) method). The learned codebook 402 comprises a set of codewords used in the compression algorithm. The goal is to represent the image using a set of codewords from codebook 402, which is more efficient than directly encoding each vector or group of pixels in the image. End-to-end optimization is performed on the learned codebook 402 to balance codebook efficiency and reconstruction quality. Figure 4 As shown, the input image x The signal is transmitted through the input encoder 202 to generate, as shown below. Figure 2The image embedding features For LSIR, at the transmitting end, the learned codebook 402 is used to embed features into the image. Mapping to code index Within the sequence. Code index. It is an integer that can be efficiently stored or transmitted from the sender to the receiver. At the receiver, the learned codebook 402 is used to recover the decoded image embedding features. (For example, by using the code index in codebook 402 that corresponds to the received code index) (Corresponding codewords). Decoder 204 is then used to embed features based on the decoded image. To reconstruct the output image .
[0045] Image captioning tasks aim to generate text keywords, sentences, and / or paragraphs to describe the content of an image. Pre-trained Large Language Models (LLMs) such as GPT-3 have demonstrated powerful capabilities in natural language understanding and generation. Trained on large-scale text data such as spreadsheets and fiction, LLMs can perform a variety of language tasks. Previous methods combined VLMs and LLMs to improve performance on image captioning tasks by leveraging different knowledge learned from the text domain and the image-text joint domain.
[0046] While text-driven image synthesis has been successfully explored in AIGC, text-driven image editing remains challenging for generative models. The essential requirement of editing methods is to preserve most of the original image content while only changing what the text description indicates. Meeting this requirement in AIGC is not easy due to the lack of control over the content generated based on the text description. To avoid this problem, most text-driven image editing methods require the user to explicitly indicate the area to be edited, for example, using a mask. The more recent Prompt-to-Prompt Editing (P2PE) edits the image by injecting cross-attention maps during the diffusion process to control the generation of pixels corresponding to the text prompts in the diffusion step.
[0047] It has been demonstrated that masked image modeling (MIM) can effectively improve the quality of learned visual latent representations in a self-supervised manner for downstream tasks such as classification and recognition, as well as reconstruction tasks such as image synthesis. Typically, a visual transformer is used to predict masked missing tokens from unmasked missing tokens. Recent masked generative encoders (MAGE) use a unified training framework where a pre-trained VQGAN generates tokens from images, and then the masked encoder and image generator are trained together based on these tokens.
[0048] like Figure 2 The current LIC framework relies on learning a general compact image representation (i.e., image embedding features). Input can be captured x The key points for reconstruction (Latent space). This framework has several serious limitations. First, it is limited in learning general priors in the image domain. In practice, compression performance is inherently limited by model capacity (e.g., the network structure and number of parameters of the input encoder 202 and decoder 204). During training and testing, it is difficult to further improve compression performance beyond a good baseline due to limited model capacity, training data, and computational resources. Secondly, LIC models are taught to balance competing objectives in rate-distortion (RD) loss, where reducing reconstruction distortion and reducing bit rate are contradictory. Because balancing different loss terms is difficult in end-to-end training, it is challenging to simultaneously improve compression performance and perceptual quality.
[0049] Current CfM frameworks, such as Figure 3 As mentioned above, because a set of machine-oriented preprocessing modules 302 and / or machine-oriented postprocessing modules 304 are customized for each task-specific model, the flexibility and versatility are limited. When multiple tasks are required (e.g., multiple levels of recognition), the CfM framework needs to use multiple sets of model parameters to compute and transmit multiple encoded streams.
[0050] Compared to LIC, current LSIR frameworks, such as Figure 4 As described, it performs poorly in image compression. This is because when used for compression, LSIR uses a compact codebook (i.e., Figure 4 The learning codebook 402 in the middle) is used to... The previous complex general images were modeled. The reconstructed images often lacked expressiveness and fidelity in detail. Furthermore, flexibly controlling the compression results to adapt to different compression objectives, such as maintaining fidelity, improving perceptual quality, or meeting other needs of image usage, is quite challenging. Compression methods that are flexible, scalable, and versatile, capable of adapting to various compression requirements, are highly needed in practical applications.
[0051] Large-scale VLMs offer an opportunity to overcome the limitations of representation learning only in the image domain. The joint embedding space of text and images is explored by examining joint priors in the language-image joint domain. This greatly expands the learning capacity. Text-driven image editing, such as P2PE, also offers the opportunity for flexible control over the compression process, easily transferring text prompts using inherently highly compressed text. However, using AIGC is not straightforward for LICs. For example, as... Figure 1 and Figure 2 As mentioned above, AIGC and LIC have different objectives. Figure 2 As shown, LIC requires reconstructing the original input. x , and like Figure 1 As shown, the current AIGC framework is not designed to guarantee this requirement. That is, from the input prompt... y In the generated From joint distribution Extracted from the original input, while the joint distribution is usually not the original input. x The reconstructed version. Furthermore, in P2PE, prompt-based controls, even through a diffusion model, still cannot guarantee the authenticity of unedited content.
[0052] The success of MIM in visual representation learning is encouraging. However, it has not yet been extended to learning visual language representations. It can be expected that masking strategies can help improve the performance of the learned VLM representations. Furthermore, in terms of compression, masking strategies further reduce the number of bits because masked pixels can be skipped and not transmitted.
[0053] This invention proposes a general framework based on the masked vision-language modeling (MAVLM) method, which utilizes the powerful multimodal representation learning in AIGC of LIC to achieve high compression efficiency with high reconstruction quality. The bit rate and reconstruction quality can be flexibly adjusted according to different compression requirements.
[0054] Figure 5A An encoding / decoding framework 500A according to an embodiment of the present invention is shown. The encoding / decoding framework 500A has three processing branches: a main branch 502, a control branch 504, and a diffusion branch 506. For example... Figure 5A As shown, at the sending end, the main branch 502 uses a sparse image encoder 508 to process the raw input using a learned sparse vision-language representation (LSVLR). x Encoding as sparse visual features And sparse text features are generated using an image-based text generator 510. To describe the input x The content. Control branch 504 will control signals. ctl and original input x As input, and through the control generation module 512, a control potential is generated. In one embodiment, sparse visual features are used. Sparse text features and control potential The feed is sent to the MAVLM module 514. The MAVLM module 514 is used to calculate the adjusted masked sparse visual features. Adjusted masked sparse text features and adjusted masked control potential Adjusted masked sparse visual features Adjusted masked sparse text features and adjusted masked control potential Typically composed of integers, text, and a small number of numbers, it can be efficiently transmitted to the decoder.
[0055] Furthermore, the diffusion branch 506 uses the compression module 516 based on the original input. x and adjusted masked sparse visual features and adjusted masked control potential To calculate diffusion potential characteristics Potential characteristics of diffusion Capture the original image x The key points regarding fidelity and expressive detail are added to the main branch 502. Potential diffusion features. It requires a very low bit rate for transmission and can be further compressed through quantization and arithmetic coding. Adjusted masked sparse visual features. Adjusted masked sparse text features Adjusted masked potential and diffusion potential characteristics Transmitted from the encoder at the transmitting end to the decoder at the receiving end.
[0056] exist Figure 5AIn the receiving end, within the main branch, the visual feature recovery module 518 is used to recover visual features based on the received adjusted masked sparse visual features. To generate the restored masked visual features The text encoding module 520 is used to adjust the received masked sparse text features. Calculate the masked text features after encoding Furthermore, in control branch 504, control encoding module 522 is used to control the potential based on the received adjusted masked data. The restored masked visual features and encoded masked text features To calculate the encoded masked control features In one embodiment, the reconstruction module 524 is used to reconstruct the masked visual features. Encoded masked text features and encoded masked control features To reconstruct baseline output .
[0057] like Figure 5A As shown, baseline output This will be combined with supplementary information from the diffusion branch to reconstruct the final output. In one embodiment, the decoder uses diffusion decompression module 526 based on the received diffusion latent features. and baseline output To calculate supplementary output Supplementary output will be provided. and baseline image output The output is passed to the fusion module 528. The fusion module 528 is used to combine and supplement the output. and baseline image output To rebuild the final output The final output represents the original input. x The decoded image. In all disclosed embodiments, the decoder can then output the final image. Transmitted to a display device for displaying the decoded image, or the final output... The image is transmitted to another computing device, such as, but not limited to, the client device that requested the image. In some embodiments, the decoder may pass the decoded image to another application (e.g., an image editing application) for further processing.
[0058] Figure 5B An encoding / decoding framework 500B according to an embodiment of the present invention is shown. The encoding / decoding framework 500B is similar to... Figure 5AIn the encoding / decoding framework 500A, besides the decoder, the post-diffusion decompression module 530 is used for decompression based on the received diffusion latent features. The restored masked visual features and encoded masked text features To calculate supplementary output .
[0059] In the disclosed embodiments, the term "feature" generally refers to a set of values that collectively provide a detailed representation of specific information or features related to the original image or its content, and these features are operated and processed through various stages of the described method. In one example, sparse visual features represent the original image in a reduced or compressed form, preserving necessary visual information while being sparse. In another example, sparse text features indicate the content of the original image in a text or descriptive format, and are also sparse.
[0060] Figure 6 A detailed workflow 600 of the main branch provided by one embodiment of the present invention is illustrated. As shown above... Figures 5A to 5B As discussed in the paper, main branch 502 uses sparse image encoder 508 to process the raw input. x Encoding as sparse visual features .
[0061] In one embodiment, the original image x It has shape The general three-dimensional (3D) tensor, in which w, h, c These are the width, height, and number of channels of the image. For example, for a color image, For spectral images, Or for RGB-D (color and depth) images, or for video clips, ,in c It refers to the number of frames. For example... Figure 6 As shown, the sparse image encoder 508 includes a visual embedding module 602 for embedding the original image... x Encoding as having shape Visual feature tensor , where the width and height Depending on the width and height of the input and the network structure of the visual embedding module 602, where d This refers to the number of feature channels. Various neural networks can be used as the visual embedding module 602. For example, in one embodiment, a visual transformer (ViT) can be used. The ViT is used to transform the original image... xThe image is divided into multiple image patches, and these image patches are encoded into a sequence. In another embodiment, a convolutional neural network (CNN) structure is used, in which the entire original image is processed in parallel. x Encode it.
[0062] Then the visual feature tensor The code is passed to the visual code generation module 604. The visual code generation module 604 is used to generate the code based on the visual feature tensor. and visual codebook 606 to compute latent features based on sparse codebook In one embodiment, the visual codebook 606 includes Each codeword has a unique identifier. d dimension. Each pixel in ( ) corresponds to the corresponding latent features Recent writing : , in It's a distance metric, such as the L1 or L2 norm. The L1 norm is the sum of the absolute values of the entries in a vector. The L2 norm is the square root of the sum of the vector terms. In other words, it represents the entire latent feature based on the sparse codebook. With The index corresponding to each codeword A number of integers. Latent features based on sparse codebooks. It can be transmitted to the decoder efficiently and without loss with very little bit consumption.
[0063] In addition, as mentioned above Figures 5A to 5B As discussed earlier, main branch 502 uses an image-based text generator 510 to generate sparse text features. To describe the input x The content. For example... Figure 6 As shown, at the sending end, the image-based text generator 510 includes a text generation module 608, which takes the original image as input and calculates latent language features. In one embodiment, latent language features It contains text words and / or sentences that can be efficiently transmitted to the decoder. The text generation module 608 generates a description of the original image. x The text of the content. In one embodiment, the text generation module 608 uses a pre-trained multimodal visual language representation, such as CLIP or BLIP, which learns from the original image. x and associated text descriptiony joint priors And based on pre-trained multimodal visual language expression computation conditions The image-based text generator 510 also includes a text code generation module 610 and a large language model (LLM) 612 for generating text based on latent features. To calculate sparse text features .
[0064] Figures 7A to 7C An exemplary embodiment of a text code generation module provided by one embodiment of the present invention is shown. For example... Figure 7A As shown, the text code generation module 702 includes an LLM 704, used to compute enhanced text latent features by generating natural language descriptions. In one embodiment, the text code generation module 702 includes a text code mapping module 706, used to calculate sparse text features based on the text codebook 708. Text codebook 708 includes Codewords, where each codeword can be a natural language element, such as a word, phrase, sentence, or a combination of these elements. Text codebooks can also be learned. 708, each codeword is d T 3D feature vectors. Sparse text features. It consists of codeword indices (e.g., integers), which can be easily transmitted to the decoder.
[0065] like Figure 7B As shown, the text code generation module 710 directly uses the latent features of the text. Sparse text features are calculated using the text code mapping module 712 and the text codebook 714. Instead of using LLM. For example... Figure 7C As shown, in the text code generation module 716, the text code mapping module is not used; instead, the enhanced text latent features generated from LLM 718 are used directly. As a feature of sparse text This also ensures efficient transmission. In all disclosed embodiments, the text codebook can be rule-defined or pre-trained, where LLM is a pre-trained large language model, such as a generative pre-trained transformer (GPT-3).
[0066] Return to reference Figure 6 sparse visual features Sparse text features and control potential The data is fed into the MAVLM module 614. The MAVLM module 614 includes a visual mask module 616 and a text mask module 618. The visual mask module is used for data based on sparse visual features. To calculate masked sparse visual features This text masking module is used for sparse text features To calculate masked sparse text features The visual masking module 616 pairs sparse visual features. Apply a visual mask to change according to the visual mask. The corresponding item (also called a token) in the database. For example, a visual mask could be a... Binary masks with the same shape, A token containing the corresponding item in the visual mask and having a value of 0 is set to a specified value indicating removal. The text mask module 618 handles sparse text features. Apply a text mask to change based on that text mask. The corresponding token in the [text]. For example, the text mask could be [the token that corresponds to the [text] token]. Binary masks with the same shape, The token containing the corresponding item in the text mask and having a value of 0 is set to the specified value indicating that it is to be removed.
[0067] Additionally, it will control potential Feed it into the visual masking module 616 to calculate the adjusted masked latent image. For example, controlling potential Zhongyu The item corresponding to the token removed from the list will also be in Removed from the middle. The MAVLM module 614 also includes a vision-language (VL) transformer 620 for use with masked sparse visual features. Features of sparse text masking To calculate the adjusted masked sparse visual features and adjusted masked sparse text features The VL Transformer 620 is designed to connect visual and linguistic features through cross-modal prediction, while improving the robustness of the learned visual and textual feature representations through masked prediction.
[0068] Figure 8 The detailed operation of a VL converter module provided in one embodiment of the present invention is illustrated. The VL converter 800 is... Figure 6A detailed example of the VL transformer described herein is provided. In one embodiment, the VL transformer 800 uses a query transformer structure, such as Q-Former. The VL transformer 800 includes a visual embedding module 802 and a text embedding module 804, the visual embedding module being used to embed text based on masked sparse visual features. To calculate masked visual embedding features This text embedding module is used to embed text based on masked sparse text features. To calculate masked text embedding features In one embodiment, the visual embedding module 802 can simply be based on The codeword index in the codebook retrieves codewords from the visual codebook. For example, unmasked... Each "pixel" ) is the index for The typing. For The masked (removed) tokens in the code can use pre-defined (either predefined or learned during training) tokens. d 3D feature vector. In one embodiment, the structure of the text embedding module 804 depends on... The format. In one example, when When text is included, the text embedding module 804 typically uses an encoder network structure to... Each token in the calculation d T 3D feature vector. In one example, when When the text embedding module 804 includes the codeword index of the learned text codebook, it can simply... The codeword index in the text codebook retrieves codewords. In one embodiment, for... The masked tokens in the code can use pre-defined (predefined or learned during training) tokens. d T 3D feature vectors.
[0069] The VL transformer 800 also includes a visual self-attention block module 806 for passing masked visual embedding features. In one embodiment, the visual self-attention block module 806 has a self-attention transformer layer, wherein the query, key, and value are all derived from masked visual embedding features. The generated VL transform 800 also includes a text self-attention block module 808 for passing masked text embedding features. In one embodiment, the text self-attention block module 808 has a self-attention transformer layer, wherein the query, key, and value are all derived from masked text embedding features. Generated.
[0070] In one embodiment, the outputs of the visual self-attention block module 806 and the text self-attention block module 808 are passed through the visual cross-attention block module 810 and the text cross-attention block module 812. In one embodiment, the visual cross-attention block module 810 has a cross-attention transformer layer, where the keys and values are derived from masked visual embedding features. The generated query is derived from masked text embedding features. The generated text self-attention block module 812 has a cross-attention transform layer, where the keys and values are derived from masked text embedding features. Generated, while the query is generated from visually embedded features via a mask. Generated.
[0071] The visual cross-attention block module 810 is followed by the visual predictor module 814, which is used to compute the adjusted masked sparse visual features. The visual predictor module 814 is designed to predict using a pre-trained system. A prediction network using codeword indices of masked visual tokens is used to generate unmasked visual embedding features. The text cross-attention block module 812 is followed by the text predictor module 816, which is used to compute the adjusted masked sparse text features. The text predictor module 816 is designed to predict texts using a trained predictor. A prediction network based on the codeword index of the masked text token is used to generate unmasked text embedding features. .
[0072] The VL converter 800 also includes a visual mask adjustment module 818 and a text mask adjustment module 820. The visual mask adjustment module 818 is used to compute adjusted masked sparse visual features using a network trained to compute masks. This mask preserves the most important information to reconstruct the input from the adjusted masked output. The text masking module 820 is used to compute adjusted masked sparse text features using a network trained to compute the mask. This mask preserves the most important information to reconstruct the input from the adjusted masked output. Typically, the VL converter 800 adjusts the original output by considering the cross-modal and unimodal relationships of visual and linguistic tokens. and .
[0073] In one embodiment, during the testing phase, the visual masking module 616 and the text masking module 618 can be considered as an additional bit compression method, and the masked tokens can be further indicated by the compressed bit consumption. This masking ratio during the testing phase can be flexibly adjusted and can be skipped based on actual compression requirements (i.e., no mask is applied to remove any elements). During training, a relatively high non-zero masking ratio is used to train the VL transformer 800.
[0074] Return to reference Figure 6 Adjusted masked sparse visual features from VT converter 620 and adjusted masked sparse text features And the potential for adjusted masking control The encoder at the sending end sends the data to the decoder at the receiving end.
[0075] At the receiving end, the decoder includes a visual feature recovery module 622. In this embodiment, the visual feature recovery module 622 is used to recover visual features based on the received adjusted masked sparse visual features. Based on the same visual codebook as the sender 624 to generate a shape The restored masked visual features In one embodiment, ( Each "pixel" is indexed. The code words.
[0076] In one embodiment, the decoder further includes a text encoding module 626, configured to base its encoding on the received adjusted masked sparse text features. Calculate the encoded masked text features z _x^(LM). Text encoding module 626 depends on The format. In one example, when When text is included, the text encoding module 626 typically employs an encoder network structure to encode text. Each token in the calculation d T 3D feature vector. In one example, when When the codeword index of the learned text codebook is included, the text encoding module 626 can simply... The codeword index retrieves codewords from the text codebook. The masked token in the middle can be set to Predetermined (pre-trained or learned during training) d 3D feature vectors. The masked token in the middle can be set to Predetermined (pre-trained or learned during training) d T 3D feature vector In one embodiment, the decoder further includes a control encoding module 628 for controlling the latent based on the received adjusted masked data. and the restored masked visual features To calculate the encoded masked control features Details of the control encoding module 628 will be provided below. Figure 10 As described in the text. Figure 6 As shown, the restored masked visual features and the restored masked text features With encoded masked control features Together they are fed into the reconstruction module 630 to reconstruct the baseline output. In reconstruction module 630, there are multiple methods to combine the recovered masked visual features. Features of the restored masked text and encoded masked control features .
[0077] Figure 9 The detailed workflow of a reconstruction module 900 provided in one embodiment of the present invention is illustrated. For example... Figure 6 As shown, the restored masked visual features and the restored masked text features With encoded masked control features Together they are fed into the reconstruction module 630 to reconstruct the baseline output. .
[0078] In one embodiment, similar to the VL converter in the transmitter encoder, the recovered masked visual features are... The data is fed into the visual self-attention block module 902, followed by the visual cross-attention block 904 and the visual predictor 906 to compute the recovered visual features. In one embodiment, the recovered masked text features are... The text features are fed into the text self-attention block 908, followed by the text cross-attention block 910 and the text predictor 912 to compute the recovered text features. In one embodiment, generator 914 is based on the restored visual features. Restored text features and encoded masked control features To calculate baseline output .
[0079] Generator 914 typically has a network structure with multiple CNN layers, such as the decoding network of a variational autoencoder (VAE). In some embodiments, it is tuned via affine transformation. To and and Perform weighted combination and generate a weighted combination. and affine parameters , New combination features . ,in It is a cascading operation. It is an aggregation from and The information is manipulated, for example, through convolution. In other embodiments, the reconstruction module 900 may be a decoder diffusion model, such as a text-modulated image generation model or a guided language to image diffusion for generation and editing (GLIDE) model, or other cue-modulated image generation models. The recovered text features... and encoded masked control features It provides guidance for the image diffusion process.
[0080] Figure 10 The processing flow of a control branch 1000 provided in one embodiment of the present invention is illustrated. In one embodiment, given an input control signal... ctl The instruction generation module 1002 first uses the prompt LLM 1004 based on control signals. ctl To generate text instructions Note that LLM 1004 can be the same as or different from the LLM used in the main branch in Figure 7. (Text instructions) It is a description ctl The prompt generation module 1006 provides a textual description of the control requirements. In one embodiment, the prompt generation module 1006 is based on textual instructions. Input image x Control signals ctl And prompt VLM 1008 to generate input-oriented text instructions. and additional input-oriented prompts .
[0081] In one embodiment, control signal ctl It can take many different forms, such as one or a combination of the following control mechanisms: text description, sketch, bounding box, color panel, etc. The instruction generation module 1002 is based on control signals. ctl The text description uses a pre-trained transformer (GPT-3) and other prompts in the VLM 1008 to compute text instructions. Text instructions Typically for control signals ctl The instructions described in Chinese are further enhanced with more detailed text descriptions. Control signals are processed in the prompt generation module 1006. ctl Other forms of raw input and enhanced text instructions To use the prompt VLM1008 to calculate input-oriented text instructions. and additional input-oriented prompts The VLM 1008 for prompts typically models multimodal embedding representations between text descriptions, various forms of prompts, and images.
[0082] In one embodiment, the multimodal embedding of the image and text description learned under the guidance of the image segmentation mask can be used as a cue VLM 1008, which generates additional input-oriented cue instructions. As a positioning control signal ctl The bounding box of the region of interest. A concrete example is a control signal. ctl Aimed at ensuring the reconstruction quality of specific objects in the scene (e.g., ctl It is a text description of "HQ Bear". Enhanced text instructions. These requirements are described in detail (e.g., It is "to ensure high quality and high resolution for animal bears"). Calculate input-oriented text instructions. To reflect input x The actual content (e.g., It is "to ensure high-quality, high-resolution images of brown bears fishing in the river". Additional input-oriented prompts are also included. It could be the junction box of the bear in the image that needs attention.
[0083] In one embodiment, control signal ctl This can vary depending on different compression requirements. For example, the control signal could be "Fidel Bear" instead of "HQ Bear" to emphasize the fidelity of the reconstructed content rather than perceptual quality. Such requirements are very useful for continuous detection and recognition tasks in machine analysis. Besides textual descriptions, other types of cues can also be used as control signals. ctl For example, selecting a texture style or a color style. In another example, control signals... ctl It can include the text description "leaf background" and a warm leaf color panel. Enhanced text commands. This could be "adjust the image to have the color of leaves". The two examples above can be combined into a complex control signal. ctl Enhanced text instructions for "Fidel Bear, leafy background". This could be something like, "Adjust the image to have the colors of leaves while preserving the original colors of the bear animal." Accordingly, this is a text-based instruction directed at the input. and additional input-oriented prompts These control instructions will be changed to reflect the changes in the image after it has been reconstructed.
[0084] In one embodiment, a cross-attention mechanism trained to capture attentional responses across multiple modalities, including images, text, and various cues, can be used to implement the cue VLM module 1008, such as the cross-attention used in P2PE. An exemplary structure for the cue VLM module 1008 is a multimodal encoder with cross-attention, followed by a multimodal generator. In one embodiment, the cue VLM module 1008 can be implemented using a network structure that adds conditional cue controls, where desired types of cue controls (e.g., sketches, masks, bounding boxes, etc.) can be added to the base VLM for text-image embedding.
[0085] In one embodiment, text instructions oriented towards input and additional input-oriented prompts Formation of control potential The control potential is fed into the visual masking module 1010 to generate an adjusted masked control potential based on masking operations. For example, text instructions that are input-oriented can be used. Adjusted to the masked input text command The instruction is to "adjust the image to have the colors of leaves while preserving the original colors of the bear animal, focusing on the bear animal with less masking." Additionally, input-oriented prompts will be provided. Adjusted to the modified input-oriented prompts with added masking. In one embodiment, control potential No changes are needed. and Same. Then, the adjusted masked control potential... Send to the receiving end.
[0086] At the receiving end, the received The information is fed to the prompt embedding module 1012, which calculates the encoded masked control feature using the prompt VLM module 1014. In one embodiment, because the cue VLM 1014 models the multimodal embedding representation between text descriptions, various forms of cues, and images, the received, adjusted, masked input text instructions can be... And the adjusted masked input-oriented prompts The data is fed into this multimodal representation space to obtain masked multimodal encoded features via a multimodal encoder. In this case, the cue embedding module 1012 can be an encoder with a cross-attention mechanism from the cue VLM 1014.
[0087] Figure 11 The processing flow of a control branch 1100 provided in one embodiment of the present invention is illustrated. In one embodiment, an input control signal is... ctl ,and Figure 10 The process is the same. The control generation module 1102 and the visual masking module 1104 generate the adjusted masked latent image. This includes adjusted masked input text instructions. And the adjusted masked input-oriented prompts In one embodiment, the same visual feature recovery module 1106 as the receiver of the main branch is used to calculate the recovered masked visual features. The same text encoding module 1108 as the receiving end is used to calculate the features of the recovered masked text. All , And the adjusted masked control potential The signal is fed into the same control encoding module 1112 as the receiving end of the control branch, which calculates the encoded masked control features. Then, the same reconstruction module 1110 as the receiver of the main branch uses the encoded masked control features. The restored masked visual features and the restored masked text features To calculate the baseline output image In one embodiment, the control adjustment module 1114 outputs an image based on the reconstructed baseline. and the original input image x To adjust the potential masked control after the adjustment Finally, after control adjustments, the adjusted masked control will be applied to the potential... The control encoding module 1116 is transmitted to the receiving end.
[0088] Figure 12 The processing flow 1200 of a control adjustment module provided in one embodiment of the present invention is illustrated. For example... Figure 12 As shown, the loss and execution update module 1202 uses the original input. x and baseline output image The distortion between the two values is used to calculate the distortion loss (e.g., MSE), which, along with the gradient of the loss, is further updated by the execution update module via automatic backpropagation of the update loss to adjust the masked latent image. The text description and prompts. In some other embodiments, by observing the original input x and baseline output image It can be manually adjusted to change the underlying masking control. The text description and prompts. In some other embodiments, direct operations can be performed on the encoded image embedding, text embedding, and / or prompt embedding to modify the adjusted masked latent image, for example, by injecting random noise. .
[0089] Furthermore, the control adjustment module 1200 can employ various strategies to update the adjusted masked control potential. An exemplary configuration of the control adjustment module 1200 is a multimodal encoder 1204, followed by a multimodal generator 1206. The multimodal encoder 1204 takes images, text descriptions, and other cues as input and computes encoded image embeddings, text embeddings, and cue embeddings in a multimodal visual language space. In one embodiment, the multimodal generator 1206 uses the encoded image embeddings, text embeddings, and cue embeddings to generate images, text descriptions, and cues.
[0090] Figure 13A The processing workflow of a diffusion branch provided by an embodiment of the present invention is illustrated. Typically, the diffusion branch 1300A uses a conditional diffusion model (CDM) to provide the original input image with minimal transmission bit cost. x The fidelity and expressive details are enhanced to supplement the reconstructed output from the main branch, resulting in a final output. For the original input x It is true. Furthermore, in the depicted embodiment, the degradation module 1302 includes a downsampling module 1304 and an encoding module 1306. Original input x First, downsampling is performed by downsampling module 1304 to produce a smaller resolution, and then encoding module 1306 uses the adjusted masked sparse visual features. The information is encoded as diffusion potential features In some embodiments, the downsampling module 1304 may use a learned downsampling network or a bicubic filter, or other preset methods. In some embodiments, the encoding module 1306 may use a compression method (e.g., conventional encoding tools such as VVC / HEVC / JPEG, or LIC methods). In some embodiments, the adjusted masked sparse visual features... This guides the downsampling module 1304 and the encoding module 1306. For example, masked regions can be downsampled less and allocated more bits, while unmasked regions can be downsampled and allocated fewer bits, thus allowing for better bit allocation to preserve the input. x The content. The downsampling module 1304 and the encoding module 1306 can also be selected to downsample or average the encoding of all regions. Overall, the encoding module 1306 frequently uses high compression ratios to diffuse latent features. It has a relatively low transmission bit rate (further compressed through quantization and arithmetic coding).
[0091] At the receiving end, the received decoding diffusion potential features (Typically after arithmetic decoding and dequantization) used to calculate supplementary output via diffusion decompression module 1308. In one embodiment, will The image is fed into the pixel recovery module 1310. The pixel recovery module 1310 is used to recover the reconstructed image. The pixel restoration module 1310 is the image decoding process corresponding to the image encoding process used by the transmitting end encoding module 1306. The latent embedding module 1312 is used for image decoding based on the reconstructed image. To compute the latent features of the embedding Then, output the baseline image. (corresponding to) Figure 5A (workflow) or restored masked visual features and the restored masked text features (corresponding to) Figure 5B The reverse diffusion module 1314 uses the embedded latent features as a diffusion condition (the workflow). To generate supplementary output The potential embedded module 1312 is an encoder network, such as the encoder portion of a VAE.
[0092] Figure 13B The processing workflow of a diffusion branch provided by an embodiment of the present invention is illustrated. At the transmitting end, the diffusion branch 1300B and... Figure 13A The diffusion branch in the 1300A is the same. At the receiving end, such as... Figure 13B As shown, The data is fed into the latent transformation module 1316. The latent transformation module 1316 is used to calculate the embedded latent features. Typically, the latent transform module 1316 improves resolution and feature channels to enhance the transformation. Execute enhancements and obtain In some embodiments, the potential transformation module 1316 may be eliminated. and Similar. Then, with Figure 13A Similarly, the backdiffusion module 1314 uses embedded latent features and baseline output image (corresponding to) Figure 5A (workflow) or restored masked visual features and the restored masked text features (corresponding to) Figure 5B Use any of the workflows to generate supplementary output. .
[0093] Figure 13A and Figure 13B The back diffusion module can use any diffusion process, including the denoising diffusion probabilistic model (DDPM) or the updated denoising diffusion implicit model (DDIM). The back diffusion module can operate as DDPM in the pixel domain or as a latent diffusion model (LDM) in the latent domain. In the case of DDPM, it uses... Figure 10 In an embodiment of A, the potential embedding module is skipped. In this case, the reconstructed image... The feed is directly sent to the backdiffusion module to calculate the supplementary output. .
[0094] In one embodiment, when the encoding module 1306 uses a conventional compression method such as VVC / HEVC / JPEG, it will use... Figure 13A The framework in which the pixel restoration module 1310 is the corresponding decoding process of the compression method to calculate the reconstructed image. When the encoding module 1306 uses the LIC method, the pixel recovery module 1310 can be... Figure 13A The corresponding decoding process of the LIC method in the framework, or you can use Figure 13B The framework, in which intermediate decoding features from LIC can be directly transformed into This invention does not impose any restrictions on the compression method or intermediate decoding features used.
[0095] Figure 14A and Figure 14B The reverse diffusion modules provided in two embodiments of the present invention are shown; Figure 14A and Figure 14B The reverse diffusion module is in Figure 13A and Figure 13BAn example of a reverse diffusion module implemented in the code, which uses CDM to generate supplemental image details. Figure 14A The 1400A reverse diffusion module and Figure 14B The reverse diffusion module 1400B includes a condition module 1402 and a reverse prediction module 1104.
[0096] Given a baseline, output image (corresponding to) Figure 5A (workflow) or restored masked visual features and the restored masked text features (corresponding to) Figure 5B (Workflow), Condition module 1402 is used to calculate diffusion conditions In some embodiments, when using At that time, conditional module 1402 is an embedded network used to output a baseline image from the pixel domain. Encoded as latent (e.g., having latent features with the embedding) (Same dimension). In one embodiment, when using and When used as a diffusion condition, condition module 1402 is a transformation network used for combination. and And transform the combined result to the desired dimension (e.g., with...). (Same). Then, the back-prediction module 1404 is used to calculate the back-diffusion step of the conditional DDPM. Or the reverse diffusion step of LDM In one embodiment, a total of T iteration t =1, …, T . T It can be preset or configured for each input. x Confirmed. Or, T It can be determined at the receiving end, or it can be determined at the sending end and synchronized with the receiving end. Send them together to the receiving end. For DDPM, such as Figure 14A As shown, in T After the next iteration, it can be used directly. As output For LDM, such as Figure 14B As shown, in T After the last iteration, finally The final output is further processed by the decoding network 1406 (e.g., the upsampling part of UNet). The output from the main branch can be... Considered The initial estimate, which is compared with the output from the diffusion branch. Combine to generate the final output. In some embodiments, the fusion module 1408 may simply be an addition operation, i.e. Other interpolation networks can be used in fusion model 1408 to further enhance the merging results.
[0097] Without sacrificing generalization ability, CDM can be viewed as using the main branch to compute deterministic initial estimates. and generate the initial estimate (or generate) Potential characteristics and This serves as a condition to guide the diffusion process in generating residual details to supplement the initial estimate. The CDM method provides robustness in controlling the generation process to recover the content of the original image by shifting the objective from generating the entire natural image to generating a residual of the image. A similar CDM method is used for text-to-speech generation to reduce generation complexity.
[0098] In some embodiments, the back prediction module can employ either a primitive fractional diffusion model using ordinary differential equations (ODEs) or a consistent diffusion model based on probability-flow ordinary differential equations (PF-ODEs). The number of iterations T can vary between single-step and multi-step iterations. .
[0099] Figure 15 A diagram illustrating an apparatus 1500 provided according to an embodiment of the present invention is shown. The apparatus 1500 can be used to implement embodiments of the present invention, such as, but not limited to, an encoder or decoder. For example, the apparatus 1500 can be used to perform… Figures 5A to 14B The encoder or decoder functionality provided in any of the embodiments shown is illustrated. Apparatus 1500 includes a receiving unit (RX) 1520 or a receiving component for receiving data through an inlet port 1510. Apparatus 1500 also includes a transmitting unit (TX) 1540 or a transmitting component for transmitting data through an outlet port 1550. For example, at the transmitting end, the encoder can use RX 1520 or the receiving component to obtain the raw image and / or control commands, and then use TX 1540 or the transmitting component to transmit encoded image information (e.g., visual language latent features, diffusion latent features, and control latent requirements) to the receiving end. At the receiving end, the decoder can use RX 1520 or the receiving component to obtain the encoded image information, and then use TX or the transmitting component to decode the raw image (e.g., the final output image). It is sent to a display device or another computing device.
[0100] The apparatus 1500 includes a memory 1560 or data storage component for storing instructions and various types of data. The memory 1560 can be any type or combination of memory components capable of storing data and / or instructions. For example, the memory 1560 may include volatile and / or non-volatile memory, such as read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM). The memory 1560 may also include one or more disk drives, tape drives, and solid-state drives. In some embodiments, the memory 1560 may be used as an overflow data storage device to store such programs when program execution is selected, and to store instructions and data read during program execution.
[0101] The device 1500 has one or more processors 1530 or other processing units (e.g., a central processing unit, CPU) to process instructions. The one or more processors 1530 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The one or more processors 1530 are communicatively coupled to an ingress port 1510, an RX port 1520, a TX port 1540, an egress port 1550, and a memory 1560 via a system bus. The one or more processors 1530 may be used to execute instructions stored in the memory 1560. Therefore, the one or more processors 1530 provide a component for performing any calculation, comparison, determination, initiation, configuration, or any other action corresponding to the claims when the processor executes appropriate instructions. In some embodiments, the memory 1560 may be a memory integrated with the processor 1530.
[0102] In one embodiment, memory 1560 stores a MAVLM-based LIC module 1570. The MAVLM-based LIC module 1570 includes data, executable instructions, and / or another submodule for implementing the disclosed embodiments. Therefore, including the MAVLM-based LIC module 1570 substantially improves the functionality of device 1500.
[0103] The embodiments of the present invention provide at least the following technical advantages.
[0104] AIGC utilizes a powerful multimodal VLM representation to generate high-quality images with high compression ratios. The main branch uses a learned sparse visual language representation, consisting of integers and text. This type of representation is highly efficient for transmission. The VLM models the joint distribution of the image representation and the corresponding text description based on a sparse codebook. Trained on large-scale image-text data pairs, the VLM extracts richer features from both the image and text domains to better describe the input image compared to using the image domain alone. Compression performance is significantly improved compared to previous LIC methods that learned models only in the image domain.
[0105] MAVLM for learning general and robust sparse visual-language representations. This invention extends masked image modeling to masked image and text modeling, where the VL transformer connects sparse visual and linguistic features through cross-modal prediction, while improving the robustness of the learned visual and text feature representations through mask prediction.
[0106] Flexible quality control: The diffusion branch provides supplementary fidelity and representational details extracted from the current input image to improve the reconstruction fidelity of the original input. These details can be added selectively, and their quality can be flexibly adjusted based on specific circumstances such as computational power, time constraints, and quality requirements. For example, with limited computational power and strict time constraints, these details can be skipped, and the decompressed output can be passed through the main branch in a single inference pass, potentially reducing the fidelity of the original input. When the goal is to provide high-quality, high-fidelity output, and computational power or time is not an issue, multiple diffusion iterations can be used to add rich details to the output.
[0107] Flexible, task-oriented control through prompts: The control branch uses a prompt VLM, which models multimodal embedding representations between text, images, and various forms of prompts to enable guided compression using prompt commands. Control signals can take a default form (e.g., ensuring fidelity or perceptual quality) or can be configured to adapt to specific compression objectives (e.g., emphasizing a specific object so that it can be reconstructed in a certain way). The transmission control potential can be adjusted automatically or manually to reduce online loss.
[0108] Furthermore, without departing from the scope of the invention, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods. Other items shown or described as coupled to each other, or directly coupled, or communicating with each other, may be indirectly coupled or communicated electrically, mechanically, or otherwise through some interface, device, or intermediate component. Those skilled in the art can identify other examples of changes, substitutions, and modifications, and make changes, substitutions, and modifications without departing from the spirit and scope of the invention.
Claims
1. A method for implementing an encoder, characterized in that, include: Encode the original image into sparse visual features; Generate sparse text features that indicate the content of the original image; Calculate control latent features based on the control signals and the original image; The adjusted masked sparse visual features, the adjusted masked sparse text features, and the adjusted masked control latent features are calculated based on the sparse visual features, the sparse text features, and the control latent features, respectively. Calculate the potential diffusion characteristics; The adjusted masked sparse visual features, the adjusted masked sparse text features, the adjusted masked control latent features, and the diffusion latent features are transmitted to the decoder.
2. The method according to claim 1, characterized in that, It also includes: calculating the diffusion latent features based on the original image, the adjusted masked sparse visual features, and the adjusted masked controlled latent features, the diffusion latent features capturing the fidelity and expressive details of the original image.
3. The method according to any one of claims 1 and 2, characterized in that, The control potential features indicate the encoded control requirements.
4. The method according to any one of claims 1 to 3, characterized in that, Calculating the control potential characteristics includes: The text instruction is calculated based on the control signal, and the text instruction is a text description of the control requirements of the control signal; The input-oriented text instructions and additional input-oriented prompts are calculated based on the text instructions, the original image, the control signals, and the visual-language model (VLM). The control potential features are calculated based on the input-oriented text instructions and the additional input-oriented prompt instructions.
5. The method according to any one of claims 1 to 4, characterized in that, The original image has shape The general three-dimensional (3D) tensor, in which w, h, c The image's width, height, and number of channels are used to encode the original image into the sparse visual features, which include: The original image is encoded into a shape. The visual feature tensor, where the width and height Depending on the width and height of the original image, d It is the number of feature channels; Sparse visual features are calculated based on the aforementioned visual feature tensor and the visual codebook, wherein the visual codebook includes multiple codewords, each codeword having... d dimension.
6. The method according to any one of claims 1 to 5, characterized in that, Encoding the original image into the visual feature tensor includes: using a visual transformer to divide the original image into multiple image blocks, and encoding the image blocks into a sequence.
7. The method according to any one of claims 1 to 5, characterized in that, Encoding the original image into the visual feature tensor includes: encoding the original image into the entire image in parallel using a convolutional neural network (CNN).
8. The method according to any one of claims 1 to 7, characterized in that, Generating the sparse text features includes: Based on the original image, latent text features including the text description are calculated, wherein the text description describes the content of the original image; The sparse text features are generated based on the latent text features.
9. The method according to any one of claims 1 to 8, characterized in that, Calculating the latent features of the text includes: Calculate the latent features of the enhanced text; The text latent features are calculated based on the enhanced text latent features and / or the text codebook, wherein the text codebook includes multiple codewords, each codeword having natural language elements.
10. The method according to any one of claims 1 to 9, characterized in that, Calculating the adjusted masked sparse visual features based on the sparse visual features includes: Based on the sparse visual features, a masked sparse visual feature is calculated by applying a visual mask to the sparse visual features to change the corresponding token in the sparse visual features according to the visual mask. The visual mask is a binary mask with the same shape as the sparse visual features, and the token with a value of 0 is set to a specified value indicating that it is removed. The adjusted masked sparse visual features are calculated based on the masked sparse visual features and using a visual-language (VL) transformer.
11. The method according to any one of claims 1 to 10, characterized in that, Calculating the adjusted masked sparse text features based on the sparse text features includes: Based on the sparse visual features, a masked sparse text feature is calculated by applying a text mask to the sparse text feature to change the corresponding token in the sparse text feature according to the text mask. The text mask is a binary mask with the same shape as the sparse text feature, and the token with a value of 0 is set to a specified value indicating that it is removed. The adjusted masked sparse text features are calculated based on the masked sparse text features and using a visual-language (VL) transformer.
12. The method according to any one of claims 1 to 11, characterized in that, Also includes: Based on the control potential features, the adjusted masked control potential is calculated by applying a visual mask to the control potential.
13. The method according to any one of claims 1 to 12, characterized in that, The modified masked control may include modified masked input-oriented text instructions and modified masked input-oriented prompt instructions.
14. The method according to any one of claims 1 to 13, characterized in that, Calculating the diffusion potential characteristics includes: The original image is downsampled to a smaller resolution image; The lower-resolution image is encoded using information from the adjusted masked sparse visual features to obtain the diffusion potential features.
15. A method for implementing a decoder, characterized in that, include: The received original image is processed with adjusted masked sparse visual features, adjusted masked sparse text features, and adjusted masked controlled latent and diffusion latent features. The recovered masked visual features are generated based on the received adjusted masked sparse visual features. The encoded masked text features are calculated based on the received adjusted masked sparse text features; The encoded masked control features are calculated based on the adjusted masked control latent, the recovered sparse visual features, and the encoded masked sparse text features. The baseline image is reconstructed based on the recovered masked visual features, the encoded masked text features, and the encoded masked control features; The supplementary output is calculated based on the diffusion potential features and the baseline image output; The final decoded image output is constructed based on the supplementary output and the baseline image output.
16. A method for implementing a decoder, characterized in that, include: The received original image is processed with adjusted masked sparse visual features, adjusted masked sparse text features, and adjusted masked controlled latent and diffusion latent features. The recovered masked visual features are generated based on the received adjusted masked sparse visual features. The encoded masked text features are calculated based on the received adjusted masked sparse text features; The encoded masked control features are calculated based on the adjusted masked control latent, the recovered sparse visual features, and the encoded masked sparse text features. The baseline image is reconstructed based on the recovered masked visual features, the encoded masked text features, and the encoded masked control features; The supplementary output is calculated based on the diffusion potential features, the masked sparse visual features, and the masked sparse text features; The final decoded image output is constructed based on the supplementary output and the baseline image output.
17. The method according to any one of claims 15 and 16, characterized in that, The restored masked visual features are based on a visual codebook, and the encoded masked sparse text features are based on a text codebook.
18. The method according to any one of claims 15 to 17, characterized in that, The calculation of the supplementary output includes: The reconstructed image is recovered based on the aforementioned diffusion potential features; The embedded latent features are calculated based on the reconstructed image; The supplementary output is calculated based on the embedded latent features and using the baseline image output as a diffusion condition.
19. The method according to any one of claims 15 to 17, characterized in that, The calculation of the supplementary output includes: The reconstructed image is recovered based on the aforementioned diffusion potential features; The embedded latent features are calculated based on the reconstructed image; Based on the embedded latent features and using the recovered masked visual features The supplementary output is calculated using the encoded masked text features as diffusion conditions.
20. The method according to any one of claims 15 to 17, characterized in that, The calculation of the supplementary output includes: The embedded latent features are calculated based on the aforementioned diffusion latent features; The supplementary output is calculated based on the embedded latent features and using the baseline image output as a diffusion condition.
21. The method according to any one of claims 15 to 17, characterized in that, The calculation of the supplementary output includes: The embedded latent features are calculated based on the aforementioned diffusion latent features; The supplementary output is calculated based on the embedded latent features and using the recovered masked visual features and the encoded masked text features as diffusion conditions.
22. The method according to any one of claims 15 to 17, characterized in that, Also includes: The supplementary output is computed using either the Denoising Diffusion Probabilistic Model (DDPM) or the Denoising Diffusion Implicit Model (DDIM).
23. The method according to any one of claims 15 to 22, characterized in that, Calculating the supplementary output includes: when using the baseline image output as the diffusion condition, using an embedding network to encode the baseline image output from the pixel domain to the latent domain.
24. The method according to any one of claims 15 to 23, characterized in that, The calculation of the supplementary output includes: when using the restored masked visual features When the encoded masked text features are used as the diffusion condition, a transformation network is used to transform the combination result into a dimension corresponding to the embedded latent features.
25. An encoder, characterized in that, include: A memory or storage device used to store instructions; One or more processors or processing devices are coupled to the memory or storage device and are used to execute the instructions to cause the encoder to perform the method according to any one of claims 1 to 14.
26. A decoder, characterized in that, include: A memory or storage device used to store instructions; One or more processors or processing devices are coupled to the memory or storage device and are used to execute the instructions to cause the decoder to perform the method according to any one of claims 15 to 24.
27. A computer program product, characterized in that, Includes computer-executable instructions stored on a non-transitory computer-readable storage medium, which, when executed by one or more processors of the device, cause the device to perform the method according to any one of claims 1 to 27.