Image coding method and device, storage medium and program product

By using a perceptual quantization network and large language model for cross-modal mapping and context compression, the method enhances image compression efficiency and reduces file sizes while preserving image quality.

CN120321403APending Publication Date: 2025-07-15PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510448085.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing end-to-end learning image compression method sacrifices compression rate when pursuing subjective image quality, resulting in a large compressed file size, which is not conducive to storage and transmission.

Method used

Image tokens are determined through perceptual quantization networks, text information is generated using cross-modal mapping technology, and context compression is carried out through large language models to generate code streams.

Benefits of technology

On the basis of ensuring that image information is not lost, the compression rate of image compression is improved, the storage and transmission costs are reduced, and the compressed image maintains a good human-eye perception effect after decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321403A_ABST
    Figure CN120321403A_ABST
Patent Text Reader

Abstract

The invention discloses an image coding method and device, a storage medium and a program product, and relates to the technical field of image compression, and the image coding method comprises the steps: processing an original image through a perceptual quantization network, and determining an image token corresponding to each sub-block in the original image; generating text information corresponding to the original image through a cross-modal mapping technology and the image tokens; and performing context compression on the text information through a preset large language model to generate a code stream corresponding to the original image. According to the invention, through the context association capability of the perceptual quantization network, the cross-modal mapping and the large language model, the code rate of the code stream is reduced, the compression rate of image compression is improved, and the image is convenient to store and transmit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image compression, and particularly to an image encoding method, device, storage medium and program product. Background Art

[0002] With the development of deep learning, many end-to-end learning image compression methods have emerged, such as ELIC (Efficient Learned Image Compression), LIC-TCM (Learned Image Compression with Transform Coding Module), etc. These methods have shown good performance in terms of distortion metrics. However, a simple distortion metric cannot fully reflect the subjective perception of human beings on image quality, and thus the Perceptual-lossy Image Compression (PIC) task has emerged.

[0003] Currently, methods for the PIC task usually adopt small models to improve the human eye perception metric. However, this small-scale end-to-end training strategy sacrifices part of the compression ratio when dealing with complex and diverse image data due to excessive pursuit of subjective image quality, resulting in a still large file size after compression, which is not conducive to image storage and transmission. Summary of the Invention

[0004] The main purpose of the present application is to provide an image encoding method, device, storage medium and program product, aiming to solve the technical problem of how to improve the compression ratio of image compression.

[0005] To achieve the above object, the present application proposes an image encoding method, and the method includes:

[0006] Processing the original image through a perceptual quantization network to determine image tokens corresponding to each sub-block in the original image;

[0007] Generating text information corresponding to the original image through a cross-modal mapping technique and each of the image tokens;

[0008] Compressing the context of the text information through a preset large language model to generate a bitstream corresponding to the original image.

[0009] In one embodiment, the step of generating text information corresponding to the original image through a cross-modal mapping technique and each of the image tokens includes:

[0010] Generating the text information according to each of the image tokens and a preset mapping relationship between the image tokens and the text information.

[0011] In one embodiment, the perceptual quantization network includes an extractor. The steps of processing the original image through the perceptual quantization network and determining the image tokens corresponding to each sub-block in the original image include:

[0012] Extracting the image features of each sub-block through the extractor;

[0013] Determining each image token according to the image codebook and each image feature.

[0014] In one embodiment, the steps of compressing the context of the text information through a preset large language model to generate the bitstream corresponding to the original image block include:

[0015] Dividing the text information into multiple text tokens, and adding a start token and an end token;

[0016] Determining the start token as the current token, and determining the probability distribution of the next text token of the current token through the large language model;

[0017] Determining the next text token as the current token, and performing the step of determining the probability distribution of the next text token of the current token through the large language model until the next text token is the end token;

[0018] Generating the bitstream through an arithmetic encoder according to the probability distribution of each text token.

[0019] In one embodiment, before the step of compressing the context of the text information through a preset large language model, it further includes:

[0020] Pre-training the large language model based on text training samples;

[0021] Fine-tuning the large language model based on image training samples.

[0022] In one embodiment, after the step of generating the bitstream corresponding to the original image, it further includes:

[0023] Decompressing the bitstream through the large language model to obtain the text information;

[0024] Generating a decompressed image according to the text information.

[0025] In one embodiment, the step of generating a decompressed image according to the text information includes:

[0026] Determining each image token according to the mapping relationship between the preset image tokens and the text information and the text information;

[0027] Through the generator of the perceptual quantization network, the decompressed image is generated according to each of the image tokens.

[0028] In addition, to achieve the above object, the present application also proposes an image encoding device, which includes:

[0029] An image token determination module, configured to process the original image through a perceptual quantization network to determine the image tokens corresponding to each sub-block in the original image;

[0030] A text information generation module, configured to generate the text information corresponding to the original image through a cross-modal mapping technique and each of the image tokens;

[0031] A bitstream generation module, configured to perform context compression on the text information through a preset large language model to generate the bitstream corresponding to the original image.

[0032] In addition, to achieve the above object, the present application also proposes an image encoding device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the image encoding method as described above.

[0033] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the image encoding method as described above are implemented.

[0034] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the image encoding method as described above are implemented.

[0035] One or more technical solutions proposed in this application have at least the following technical effects: First, by processing the original image through the perception quantization network, the image tokens corresponding to each sub-block in the original image are determined, initially realizing the quantization and refinement of image information, reducing the storage requirements of the image, and achieving data compression; furthermore, through the cross-modal mapping technology and each image token, the text information corresponding to the original image is generated, effectively converting the visual information in the image into semantic information, providing a more efficient representation method for subsequent processing, and helping to reduce data redundancy; furthermore, through the preset large language model, the text information is contextually compressed to generate a bitstream corresponding to the original image block. By virtue of the context understanding ability and data compression ability of the large language model itself, the context information in the text information is accurately captured and converted into an efficient binary bitstream, improving the compression ratio of image compression on the basis of ensuring that the image information is not lost, which is conducive to the efficient storage and transmission of image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0037] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the image encoding method of this application;

[0039] Figure 2 It is a schematic diagram of the image-text mapping relationship provided for Embodiment 2 of this application;

[0040] Figure 3 It is a schematic flowchart of the application process of the perception quantization network provided for Embodiment 2 of this application;

[0041] Figure 4 It is a schematic diagram of context compression based on a large language model provided for Embodiment 2 of this application;

[0042] Figure 5 It is a schematic flowchart of the overall process of image encoding and decoding provided for Embodiment 3 of this application;

[0043] Figure 6 It is a schematic diagram of the module structure of the image encoding device of the embodiments of this application;

[0044] Figure 7It is a schematic diagram of the device structure of the hardware operating environment involved in the image encoding method in the embodiments of this application. Detailed implementation manners

[0045] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0046] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and the image compression implementation manners.

[0047] It should be noted that the image encoding method described in this application is also related to other image processing processes such as image decoding. Therefore, the image encoding method of this application can also be called an image processing method or an image encoding and decoding method.

[0048] In this embodiment, for the convenience of description, the following will be described with the terminal as the execution subject.

[0049] Since small models are usually adopted for the task of lossy image compression to improve the human eye perception index. However, this small-scale end-to-end training strategy sacrifices part of the compression rate when dealing with complex and diverse image data due to excessive pursuit of subjective image quality; at the same time, due to the limitations of small models in feature extraction and expression ability, it is difficult to capture complex structures and detailed information in images, resulting in its inability to effectively remove redundant information during the compression process. Therefore, although the images obtained after lossy image compression may have better quality in human eye perception currently, their file sizes are still large, the compression rate is low, and the storage and transmission costs of images are increased.

[0050] The present application provides a solution. By means of a perception quantization network, the image tokens corresponding to each sub-block in the original image block can be determined, which can effectively extract image features, reduce the storage requirements of the image, and achieve data compression. Furthermore, according to each image token and the mapping relationship between the image token and the text information preset, text information is generated, converting the visual information in the image into semantic information, enabling the subsequent large language model to more accurately understand and process the image information. Furthermore, through the preset large language model, the probability distribution of each text token obtained by dividing the text information is determined. With the powerful language understanding and generation ability of the large language model, the context information in the text information can be accurately captured, improving the accuracy of the probability distribution of each text token. Furthermore, through an arithmetic encoder, a bitstream is generated according to the probability distribution of each text token. Since the accuracy rate of the above-obtained probability distribution is relatively high, when encoding through the arithmetic encoder, digital intervals can be more accurately assigned to each text token, thereby reducing the bit rate of the bitstream obtained by compressing the original image block. Therefore, the present application can improve the compression ratio of image compression on the basis of ensuring that the image information is not lost, which is beneficial to the efficient storage and transmission of image data. At the same time, it ensures that the compressed image can maintain a good human eye perception effect after decoding, avoiding common image distortion and detail loss problems in traditional compression methods.

[0051] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, taking the terminal as the execution subject as an example, this embodiment and the following embodiments will be described.

[0052] Based on this, the embodiment of the present application provides an image encoding method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the image encoding method of the present application.

[0053] In this embodiment, the image encoding method includes steps S10 to S30:

[0054] Step S10, processing the original image through a perception quantization network to determine the image tokens corresponding to each sub-block in the original image;

[0055] Optionally, the perception quantization network is a deep learning model for processing image data. By learning the feature representation of the image, the image is divided into multiple sub-blocks, and an image token is assigned to each sub-block.

[0056] Optionally, the original image refers to a region in the image data, which can contain the complete image data or a part of the image, and can be adjusted according to the specific image task. The image token refers to the quantized sub-block representation generated by the perceptual quantization network, which is itself a compact encoding of the image block features and can improve the compression ratio and reduce the storage pressure.

[0057] Exemplarily, for an original image containing "a panda sleeping on a branch", the original image can be segmented into multiple sub-blocks through the perceptual quantization network. Among them, each sub-block contains local image information (such as parts of the panda, parts of the branch, etc.); furthermore, through the quantization operation of the perceptual quantization network, these sub-blocks are converted into compact token representations. This process not only retains the core content of the image but also reduces the data volume of the image blocks obtained after decompression, thereby improving the compression ratio of image compression.

[0058] Exemplarily, a deep learning model (such as a neural network) can also be used to predict the content and complexity of the image blocks, and dynamically adjust the quantization parameters of the perceptual quantization network according to the prediction results, such as the quantization step size, quantization level, etc., so as to achieve more efficient compression while ensuring the image quality.

[0059] It can be understood that by generating image tokens, redundant and noisy information can be removed while retaining the key information of the image, thereby effectively reducing the storage requirements of the image and achieving data compression.

[0060] Step S20, generating the text information corresponding to the original image block through the cross-modal mapping technology and each image token;

[0061] Optionally, the cross-modal mapping technology refers to a technology that can convert data of different modalities (such as images, texts, etc.) into each other. By learning the internal relationship between data of different modalities, a mapping relationship is established so that data of one modality can be mapped to data of another modality.

[0062] Optionally, the text information refers to the semantic summary of the image block. In this embodiment, it is manifested as the natural language description output by the image description generation technology, or it can also be the description of the image features obtained through the mapping relationship between the image token and the text information. The text information includes text sentences, paragraphs, words, etc. Experiments have shown that the large language model preset in this embodiment has the best compression effect on text sentences.

[0063] Exemplarily, the visual features of an image can be converted into a natural language description through a deep learning model, which generally includes two stages: image feature extraction and text generation. For example, based on an encoder-decoder architecture, a convolutional neural network can be used as the encoder to extract the visual features of the image and convert the input image patches into a fixed-length feature vector; a recurrent neural network can be used as the decoder, which is mainly responsible for gradually converting the feature vector generated by the encoder into a natural language description (text information), and the output of each step depends on the output of the previous step and the features of the encoder.

[0064] It can be understood that by converting the visual information of the image patches into text information, the efficient compression and semantic processing of the image information are realized, providing more efficient data input for subsequent compression and encoding steps, which is beneficial to improving the overall compression ratio of image compression.

[0065] Step S30: Through a preset large language model, perform context compression on the text information to generate a bitstream corresponding to the original image.

[0066] Optionally, a large language model (LLMs) is a pre-trained deep learning model. By learning from large-scale text data, this model has powerful language generation and understanding capabilities, and can capture long-range dependencies and complex semantic relationships in the text. Common large language models include Qwen2.5, GPT2, Llama3.2, etc. A bitstream refers to a continuous data stream composed of binary digits, which is used to represent the compressed image data for easy storage and transmission.

[0067] Exemplarily, use the large language model to perform intelligent analysis on the text information to generate the probability distribution of the occurrence of each word or character; then these probability distributions are input into an arithmetic encoder, and by continuously narrowing an interval to represent the probability of the input sequence, a compact binary bitstream is finally generated, and this bitstream is the compressed representation of the image patch.

[0068] It can be understood that by performing context compression on the text information through the large language model, the redundant information in the text data is reduced, while the key context information is retained, which can improve the compression ratio and reduce the storage and transmission costs of redundant information.

[0069] This embodiment provides an image encoding method. By using a large language model for image processing, the potential of the large language model itself as a data compressor for data compression is fully utilized, and the form of the bitstream is also more conducive to the storage and transmission of images.

[0070] In a feasible implementation manner, before step S30, it further includes:

[0071] Step S201, pre-train a large language model based on text training samples;

[0072] Step S202, fine-tune the large language model based on image training samples.

[0073] Optionally, the text training samples refer to a collection of text data for pre-training large models, such as papers, books, etc.; the image training samples refer to a data set composed of images, which is used to enhance the large language model's recognition ability of image content and prompt the large language model to transform from a text generator to an image compressor.

[0074] Exemplarily, in the pre-training stage, the large language model interacts with the text training samples, and gradually masters the rules and usage of language by understanding and learning the language phenomena and context information in the text; in the fine-tuning stage, the model interacts with the image training samples, and improves its ability in processing text compression tasks related to images by establishing the association between images and language.

[0075] Exemplarily, the methods for fine-tuning the large language model include full-parameter fine-tuning and low-rank adaptation fine-tuning, etc., and this embodiment does not make specific limitations on this. Among them, full-parameter fine-tuning means comprehensively updating all the parameters of the model, which requires a large amount of memory and computing resources; low-rank adaptation fine-tuning freezes the parameters of the large language model and introduces an additional low-rank adapter to update the model, reducing the training parameters and computational complexity while making the model adapt to new tasks. Although full-parameter fine-tuning requires a large amount of memory resources, from the experimental performance, the compression performance of the large language model after low-rank adaptation fine-tuning is not as good as that of full-parameter fine-tuning. Therefore, users can choose the fine-tuning method of the large language model according to the terminal configuration.

[0076] In this embodiment, the two stages of pre-training and fine-tuning the large language model complement each other and jointly constitute the training process of the large language model from a text generator to an image compressor, so that the model can adapt to relevant image compression tasks.

[0077] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, step S20 includes:

[0078] Step S21, generate text information according to each image token and the mapping relationship between the preset image token and text information.

[0079] Optionally, the mapping relationship between image tokens and text information refers to a predefined rule or model that associates image tokens with text descriptions. This mapping relationship includes an image space and a text space, and the text space corresponds to the large language model used. The preset mapping relationship between image tokens and text information can be rule-based or deep learning-based. Among them, the rule-based mapping relationship embodies a forced cross-modal mapping, while the deep learning-based mapping relationship is a stochastic mapping obtained after alignment training with image-text paired data.

[0080] Experiments have shown that in the case of a small amount of data (not exceeding 1024), the two methods for determining the mapping relationship between the above-mentioned image tokens and text information have basically the same compression effect on images. Compared with the method based on image-text paired training, the rule-based mapping method does not require an additional training module or a training set supplemented with image-text paired data, has lower costs, and is easy to expand.

[0081] Exemplarily, please refer to Figure 2 , Figure 2 A schematic diagram of an image-text mapping relationship is provided. This image-text mapping relationship is the above-mentioned mapping relationship between image tokens and text information, including an image space and a text space. For example, in the process of generating text information according to each image token and the preset image-text mapping relationship, first, according to the image token, find the digital encoding in the image space, and according to the mapping relationship, determine the text character corresponding to this digital encoding in the text space. For example, when the image token is 0, the corresponding text character after conversion is "a". Then, through cross-modal mapping, convert the image tokens of each sub-block into complete text information.

[0082] Exemplarily, the text information determined based on the quantized image tokens can be used not only for the image compression task described in this application but also for other image tasks, such as image generation and image inpainting. Specifically, in the image inpainting task, the generated image information can be used as a condition, and through a generation model (such as a generative adversarial network or a diffusion model), gradually generate the missing part of the image content, and finally generate a complete picture.

[0083] In this embodiment, by converting image blocks into image tokens, the image information is effectively compressed, and the compression rate of image compression is improved. Furthermore, by converting the image tokens into text information, the large language model can more accurately understand and process the image information, so that the large language model can be used to compress the image information.

[0084] In a feasible implementation manner, the perceptual quantization network includes an extractor, and step S10 includes:

[0085] Step S11: Extract the image features of each sub-block through an extractor.

[0086] Optionally, the extractor refers to an algorithm or model component used to extract useful image features or information from input data.

[0087] Exemplarily, the extractor can be composed of a convolutional neural network and a Transformer, and is used to extract all features within the sub-block. Specifically, through the convolutional neural network, convolutional and pooling operations are performed on the input original image block to extract the local features of the original image block, and these features are organized into a feature map; furthermore, the feature map is divided into a series of non-overlapping sub-blocks, and positional encoding is added to each sub-block so that the Transformer can identify the positions of each sub-block; furthermore, these sub-blocks with positional encoding are processed by the Transformer. Based on the self-attention mechanism of the Transformer, the global context information between these sub-blocks is captured, and a new high-level feature representation is generated; finally, the Transformer outputs the image features, which contain all the features within the sub-block, including both local features and high-level features obtained from the global context understanding.

[0088] Step S12: Determine the image tokens corresponding to each sub-block according to the image codebook and each image feature.

[0089] Optionally, the image codebook is a mapping relationship that maps image features to image tokens. The image features can be color, texture, shape, etc., and the image tokens are the unique identifiers of these features.

[0090] Exemplarily, for the image features of a certain sub-block, search in the image codebook for the token feature with the highest similarity to the image features, and determine the image token to which the token feature belongs as the image token corresponding to the sub-block.

[0091] Exemplarily, please refer to Figure 3 , Figure 3 A schematic diagram of the application process of a perceptual quantization network is provided. The perceptual quantization network can include an extractor and a generator. Specifically, in the image compression stage, the extractor extracts the image features of the input image block, and through the image codebook, maps the image features to image tokens; in the image decompression stage, the generator reconstructs the discrete image tokens back into an image, obtains the decompressed image, and outputs the image.

[0092] In this embodiment, by mapping the image features of the image block to image tokens, directly storing the entire image block or storing complex image features is avoided, effectively improving the compression ratio of image compression.

[0093] In a feasible embodiment, step S30 includes:

[0094] Step S31: Divide the text information into multiple text tokens, and add a start token and an end token.

[0095] Exemplarily, divide the text information into multiple text tokens according to predefined rules (such as spaces, punctuation marks, etc.) to obtain a token list; then, add a start token <bos_id> at the beginning of the token list and an end token <eos_id> at the end. For example, for the text information "A panda is lying on a branch sleeping", the token list obtained after division is "<bos_id>", "A", "panda", "is", "lying", "on", "a", "branch", "sleeping", "<eos_id>".

[0096] It can be understood that by dividing the text information into text tokens, the natural language description of the image patch is successfully structured into a format that can be processed by a large language model, facilitating subsequent compression and encoding.

[0097] Step S32: Determine the start token as the current token, and through the large language model, determine the probability distribution of the next text token of the current token.

[0098] Step S33: Determine the next text token as the current token, and execute the step of determining the probability distribution of the next text token of the current token through the large language model until the next text token is the end token.

[0099] Optionally, the probability distribution refers to the set of probabilities of the possible values of the next text token for the current token, that is, all the possibilities and their relative magnitudes that the large language model predicts for the next text token that may appear based on the context and semantic information of the text information.

[0100] Exemplarily, the probability distribution of the next text token can be calculated one by one through LLM autoregression. Specifically, after obtaining the token list, create an empty list to store the probability distributions of each text token; determine the start token as the current token and pass it to the large language model; furthermore, through the large language model, predict the next possible text token and the corresponding probability to obtain the probability distribution of the next text token; furthermore, read the next text token from the token list, determine it as the current token, and repeat the above steps until the next text token read is the end token.

[0101] Exemplarily, the probability distribution can also be predicted through NAR (Non-Autoregressive Transformer). It can generate the probability distributions of all text tokens in parallel without marking the start token and the end token.

[0102] It can be understood that by using the large language model to predict the probability distribution of the next text token, the context and semantic information of the text sequence can be more accurately understood, and the uncertainty of the obtained probability distribution is lower, thereby improving the accuracy of the probability distribution of each generated text token.

[0103] Step S34, through an arithmetic encoder, generate a bitstream according to the probability distribution of each text token.

[0104] It should be noted that when using the probability distribution generated by the LLM for arithmetic coding, it is necessary to ensure that the probability distribution corresponds to the currently encoded token. Therefore, it is necessary to shift the input text information one position to the left so that the position of each token is aligned with the position of the probability distribution generated by the model.

[0105] Optionally, the arithmetic encoder can map the input data stream to a continuous real number interval by calculating the probability of the input data symbol, and achieve the technical effect of reducing the bit rate through encoding selective substitution; the bitstream generated by the arithmetic encoder contains the encoded representation of the original data and can be decoded to restore the original data.

[0106] Exemplarily, starting from the text token with a probability distribution through an arithmetic encoder, according to the actual token list and the probability distribution of the currently encoded text token, divide in the continuous real number interval to obtain the encoding interval corresponding to each text token; further, convert the midpoint of the encoding interval corresponding to the end token into a binary representation to obtain a compact bitstream representation (bitstream), thereby achieving efficient image compression.

[0107] It can be understood that due to the high accuracy of the probability distribution determined by the large language model, the smaller the interval division of the arithmetic encoder, the smaller the corresponding encoding bit rate, thereby improving the compression rate of image compression.

[0108] Exemplarily, please refer to Figure 4 , Figure 4A schematic diagram for context compression based on a large language model is provided. In the case where the text information is a text sentence, first, the text sentence is divided to obtain multiple text tokens, and a start token <bos_id> and an end token <eos_id> are added. Then, through the large language model after full-parameter fine-tuning or low-rank adaptation fine-tuning, the probability distribution of the next text token for each text token is predicted. Among them, in the process of full-parameter fine-tuning, all parameters of the large language model will be trained, but a large amount of GPU (Graphics Processing Unit) graphics card resources are required. While in the process of low-rank adaptation fine-tuning, the parameters of the large language model are frozen, and an additional low-rank adapter is used to update the model to adapt to a specific text compression task. Which fine-tuning method to specifically adopt can be determined according to actual needs. Further, the text sentence is shifted to the left to align the text tokens and their corresponding probability distributions. At the same time, the probability distributions of the text tokens are input into an arithmetic encoder. Finally, a bitstream is generated through the arithmetic encoder. For text information such as words and paragraphs, the same compression process is also based on the above, so it will not be elaborated here.

[0109] In this embodiment, a large language model is used to output a probability distribution, and in combination with an arithmetic encoder, the probability distribution is converted into a bitstream. Further, by utilizing the characteristic of the arithmetic encoder to save the bit rate, the compression ratio is improved.

[0110] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the content that is the same as or similar to the above-mentioned first embodiment and second embodiment can be referred to the above introduction, and will not be elaborated hereinafter. On this basis, after step S30, it further includes:

[0111] Step S40, decompress the bitstream through a large language model to obtain text information;

[0112] Exemplarily, through an arithmetic decoder, the binary representation of the bitstream is restored to a specific value. Then, according to the probability distribution predicted by the large language model and its own prediction ability, each text token is determined in turn until the complete text information is obtained.

[0113] It can be understood that through the prediction ability of the large language model, the text information can be accurately restored from the bitstream, providing an accurate basis for the subsequent generation of compressed images.

[0114] Step S50, generate a decompressed image according to the text information.

[0115] Exemplarily, for the text information generated based on an encoder-decoder architecture, corresponding compressed images can be generated through a generative adversarial network, a variational autoencoder, etc. However, this generation from text to image is full of uncertainties and is usually quite different from the original image blocks.

[0116] In a feasible implementation, step S50 includes:

[0117] Step S51, determine each image token according to the mapping relationship between the preset image tokens and the text information and the text information;

[0118] Optionally, the text information and image tokens obtained by the above decompression are exactly the same as the text information and image tokens generated during the process of compressing the original image to generate a bitstream, and this part belongs to lossless compression.

[0119] Step S52, through the generator of the perceptual quantization network, generate the decompressed image according to each image token.

[0120] Optionally, the generator is usually a neural network responsible for converting the image tokens or feature vectors generated by the extractor back into an image. Extracting features through the extractor and restoring the image through the generator belongs to lossy compression, and there are differences between the decompressed image blocks obtained and the original image blocks.

[0121] Exemplarily, through the generator, each discrete image token is decoded into a pixel representation of a sub-block, and through a series of deconvolution layers, upsampling layers, activation functions, etc., the pixel representation of each sub-block is gradually reconstructed back into an output image (i.e., the decompressed image block) with the same size as the image block.

[0122] In this implementation, the restoration based on discrete image tokens often retains the key features of the image block to the greatest extent, and can also bring a better human eye perception experience to the user.

[0123] In this embodiment, through an effective decoding and reconstruction process, the quality and details of the decompressed image are improved, so as to output a high-quality decompressed image.

[0124] Exemplarily, in order to help understand the implementation process of the image coding method obtained by combining the above-mentioned embodiment 1 and embodiment 2, please refer to Figure 5 , Figure 5 A schematic diagram of the overall process of image coding and decoding is provided. In the case where the original image block is a complete original image and the mapped text information is a text sentence, specifically:

[0125] Compression step 1: Quantization, that is, through the perceptual quantization network, convert the original image into image tokens;

[0126] Compression step 2: Cross-modal map the image tokens into a text sentence;

[0127] Compression step 3: Context compression, that is, through the fine-tuned large language model, compress the text sentence into a bitstream.

[0128] The resulting bitstream has a lower bit rate and a higher compression ratio compared to the original image, facilitating the storage and compression of the original image. Subsequently, it can enter the corresponding decompression process:

[0129] The first step of decompression: context decompression, that is, through a fine-tuned large language model, the bitstream is decompressed back into the corresponding text sentences;

[0130] The second step of decompression: cross-modal mapping of the text sentences into image tokens;

[0131] The third step of decompression: dequantization, that is, through a perceptual quantization network, the image tokens are reconstructed into the decompressed image.

[0132] Through such a decompression process, ensure that each step retains and restores the information of the original image as much as possible, so as to ensure that the decoded image obtained can bring a better human eye perception experience to the user.

[0133] Therefore, the image processing (including image encoding and decoding) method as described in this application can make the decompressed image have a better human eye perception effect when generating a smaller bitstream.

[0134] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image encoding method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0135] The embodiment of this application also provides an image encoding device. Please refer to Figure 6 , the image encoding device includes:

[0136] The image token determination module 10 is used to process the original image through a perceptual quantization network to determine the image tokens corresponding to each sub-block in the original image;

[0137] The text information generation module 20 is used to generate the text information corresponding to the original image through cross-modal mapping technology and each image token;

[0138] The bitstream generation module 30 is used to perform context compression on the text information through a preset large language model to generate the bitstream corresponding to the original image.

[0139] The image encoding device provided by the embodiment of this application adopts the image encoding method in the above embodiment and can solve the technical problem of how to improve the compression ratio of image compression. Compared with the prior art, the beneficial effects of the image encoding device provided by this application are the same as those of the image encoding method provided in the above embodiment, and other technical features in the image encoding device are the same as those disclosed in the above embodiment method, which will not be elaborated here.

[0140] An embodiment of the present application provides an image encoding device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the image encoding method in the first embodiment above.

[0141] Reference is made below to Figure 7 , which shows a schematic structural diagram of an image encoding device suitable for implementing the embodiments of the present application. The image encoding device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The image encoding device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0142] As Figure 7 shown, the image encoding device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the image encoding device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image encoding device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image encoding device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0143] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0144] The image encoding device provided by the embodiments of the present application adopts the image encoding method in the above embodiments, and can solve the technical problem of how to improve the compression ratio of image compression. Compared with the prior art, the beneficial effects of the image encoding device provided by the present application are the same as those of the image encoding method provided by the above embodiments, and other technical features in the image encoding device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0145] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0146] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0147] The embodiments of the present application provide a computer-readable storage medium, which has computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image encoding method in the above embodiments.

[0148] The computer-readable storage medium provided by the embodiments of the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0149] The above computer-readable storage medium can be included in an image encoding device; or can exist separately without being assembled into the image encoding device.

[0150] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by an image encoding device, the image encoding device is caused to: process the original image through a perception quantization network to determine image tokens corresponding to each sub-block in the original image; generate text information corresponding to the original image through a cross-modal mapping technique and each image token; and perform context compression on the text information through a preset large language model to generate a bitstream corresponding to the original image.

[0151] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0153] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0154] The readable storage medium provided by the embodiments of this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned image coding method, and can solve the technical problem of how to improve the compression ratio of image compression. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image coding method provided by the above embodiments, and will not be elaborated here.

[0155] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of the image encoding method as described above when executed by a processor.

[0156] The computer program product provided by the embodiment of the present application can solve the technical problem of how to improve the compression ratio of image compression. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the image encoding method provided by the above embodiment, and will not be elaborated here.

[0157] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. An image encoding method, characterized in that, The described image encoding method includes: Processing the original image through a perceptual quantization network to determine image tokens corresponding to each sub-block in the original image; Generating text information corresponding to the original image through a cross-modal mapping technique and each of the image tokens; Performing context compression on the text information through a preset large language model to generate a bitstream corresponding to the original image.

2. The image encoding method according to claim 1, wherein The step of generating text information corresponding to the original image through a cross-modal mapping technique and each of the image tokens includes: Generating the text information according to each of the image tokens and a preset mapping relationship between the image tokens and the text information.

3. The image encoding method according to claim 1, wherein The perceptual quantization network includes an extractor, and the step of processing the original image through the perceptual quantization network to determine image tokens corresponding to each sub-block in the original image includes: Extracting image features of each of the sub-blocks through the extractor; Determining each of the image tokens according to an image code table and each of the image features.

4. The image encoding method according to claim 1, wherein The step of performing context compression on the text information through a preset large language model to generate a bitstream corresponding to the original image block includes: Dividing the text information into multiple text tokens and adding a start token and an end token; Determining the start token as the current token and, through the large language model, determining the probability distribution of the next text token of the current token; Determining the next text token as the current token and executing the step of determining the probability distribution of the next text token of the current token through the large language model until the next text token is the end token; Generating the bitstream through an arithmetic encoder according to the probability distribution of each of the text tokens.

5. The image encoding method according to claim 1, wherein Before the step of performing context compression on the text information through a preset large language model, it further includes: Pre-training the large language model based on text training samples; Fine-tuning the large language model based on image training samples.

6. The image encoding method according to claim 1, wherein, After the step of generating the bitstream corresponding to the original image, it further includes: Decompressing the bitstream through the large language model to obtain the text information; Generating a decompressed image according to the text information.

7. The image encoding method according to claim 6, characterized in that, The step of generating a decompressed image according to the text information includes: Determining each of the image tokens according to a preset mapping relationship between the image tokens and the text information and the text information; Generating the decompressed image according to each of the image tokens through a generator of the perceptual quantization network.

8. An image encoding device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image encoding method according to any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image encoding method according to any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image encoding method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image transmission and reception system, transmitter, receiver, computer program, and image transmission and reception method

    JP7838169B1