Image Processing Method, Apparatus, Image Coding and Decoding System, Device, and Storage Medium

Through multi-stage decoding and fusion information processing, the problem of image quality degradation in image encoding is solved, and clear and natural images are generated under high compression ratio, reducing blur and block effects, and improving image processing effect.

CN119562070BActive Publication Date: 2025-07-04PEKING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510109328.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-07-04
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

In the process of image encoding, the image video quality decreases with the increase of compression ratio, and the reconstructed image is distorted on visual characteristics such as edges, colors, textures, etc., and it is difficult to reconstruct a clear image that conforms to the subjective effects of the human eye under high compression ratio, and artifacts such as blur and block effects are prone to occur.

Method used

The first decoder is used to decode the text code stream to obtain semantic information, the second decoder decodes the feature code stream, processes the feature decoding information through the texture information acquisition module, uses the hidden layer to represent the diffusion model to fusion semantics and texture information, and finally generates the decoded image through the image generation module, including the image generator and the enhancement network module to remove the block effect.

Benefits of technology

The image clarity is significantly improved under high compression ratio, reducing image blur and block effects, and generating clear images that conform to the subjective effects of the human eye, improving image processing effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119562070B_ABST
    Figure CN119562070B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, image coding and decoding system, device, and storage medium. The image processing method includes: decoding a text bitstream from an encoding end through a first decoder to obtain semantic information of an original image; decoding a feature bitstream from the encoding end through a second decoder to obtain feature decoding information of the original image; processing the feature decoding information through a texture information acquisition module to obtain texture information of the original image; processing the semantic information and the texture information based on a hidden layer representation diffusion model to obtain fusion information; and processing the fusion information through an image generation module to obtain a decoded image. The image processing method provided by the embodiments of the present application processes the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information, and processes the fusion information through the image generation module to obtain a decoded image, which can obtain a relatively clear image, greatly reduce the situation of image blurring, and has a good processing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and particularly relates to an image processing method, apparatus, image coding and decoding system, device, and storage medium. Background Art

[0002] With the explosive growth of the quantity of image content in the big data era, there is an urgent need to design an efficient coding scheme. Currently, with the adoption of more and more efficient adaptive coding tools in traditional coding standards, the coding efficiency has been continuously improved. However, with the emergence of boundary effects, it is increasingly difficult to obtain coding performance gains simply by increasing the coding mode and expanding the parameter search range. The traditional coding methods based on signal processing encounter a bottleneck in performance improvement. In the image processing methods in related technologies, it is difficult to obtain a clear image after decoding and reconstructing the encoded image, and the image is prone to blurring, resulting in poor processing effects.

[0003] The above statements are only used to provide background technical information related to the present application, and do not necessarily constitute prior art. Summary of the Invention

[0004] The purpose of the present application is to provide an image processing method, apparatus, image coding and decoding system, device, and storage medium. To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.

[0005] According to one aspect of the embodiments of the present application, an image processing method is provided, which is applied to a decoding end. The decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the method includes:

[0006] Decoding a text code stream from an encoding end through the first decoder to obtain semantic information of the original image;

[0007] Decoding a feature code stream from the encoding end through the second decoder to obtain feature decoding information of the original image;

[0008] Processing the feature decoding information through the texture information acquisition module to obtain texture information of the original image;

[0009] Processing the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information;

[0010] Processing the fusion information through the image generation module to obtain the decoded image.

[0011] In some embodiments of the present application, the texture information acquisition module includes a latent space representation module, an image reconstruction module, and a feature extractor that are connected in sequence; processing the feature decoding information through the texture information acquisition module to obtain the texture information of the original image includes:

[0012] Inputting the feature decoding information into the latent space representation module for processing to obtain latent space representation information;

[0013] Processing the latent space representation information through the image reconstruction module to obtain image reconstruction information;

[0014] Processing the image reconstruction information through the feature extractor to obtain the texture information of the original image.

[0015] In some embodiments of the present application, processing the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information includes:

[0016] Through the hidden layer representation diffusion model, based on the visual knowledge of the hidden layer representation diffusion model, fusing the semantic information and the texture information to obtain fusion information.

[0017] In some embodiments of the present application, the image generation module includes an image generator and an enhancement network module that are connected in sequence; processing the fusion information through the image generation module to obtain the decoded image includes:

[0018] Processing the fusion information through the image generator to obtain an initial generated image;

[0019] Removing block artifacts from the initial generated image through the enhancement network module to obtain the decoded image.

[0020] In some embodiments of the present application, the encoding end includes a text-to-image large model, a latent space representation extraction model, a first encoder connected to the text-to-image large model, and a second encoder connected to the latent space representation extraction model;

[0021] The text bitstream is obtained by encoding first semantic information through the first encoder; the first semantic information is the semantic information of the original image extracted through the text-to-image large model.

[0022] According to another aspect of the embodiments of the present application, there is provided an image processing method applied to an encoding end, the encoding end including a text-to-image model, a latent space representation extraction model, a first encoder connected to the text-to-image model, and a second encoder connected to the latent space representation extraction model, the method including:

[0023] Process the original image through the described image-to-text model to obtain semantic information;

[0024] Process the original image through the described latent space representation extraction model to obtain texture information;

[0025] Encode the semantic information using the first encoder to obtain a text bitstream;

[0026] Encode the texture information using the second encoder to obtain a feature bitstream;

[0027] Send both the text bitstream and the feature bitstream to the decoding end, so that the decoding end executes the image processing method applied to the decoding end described in any embodiment of the present application.

[0028] According to another aspect of the embodiments of the present application, there is provided an image processing apparatus, which is characterized in that it is applied to the decoding end. The decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the apparatus includes:

[0029] A first decoding module, configured to decode the text bitstream from the encoding end through the first decoder to obtain the semantic information of the original image;

[0030] A second decoding module, configured to decode the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image;

[0031] A texture information acquisition module, configured to process the feature decoding information through the texture information acquisition module to obtain the texture information of the original image;

[0032] A fusion information acquisition module, configured to process the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information;

[0033] An image generation module, configured to process the fusion information through the image generation module to obtain the decoded image.

[0034] According to another aspect of the embodiments of the present application, there is provided an image processing apparatus, which is applied to the encoding end. The encoding end includes an image-to-text model, a latent space representation extraction model, a first encoder connected to the image-to-text model, and a second encoder connected to the latent space representation extraction model. The apparatus includes:

[0035] A semantic information acquisition module, configured to process the original image through the image-to-text model to obtain semantic information;

[0036] A texture information acquisition module, configured to process the original image through the latent space representation extraction model to obtain texture information;

[0037] A text bitstream acquisition module, configured to encode the semantic information by using the first encoder to obtain a text bitstream;

[0038] A feature bitstream acquisition module, configured to encode the texture information by using the second encoder to obtain a feature bitstream;

[0039] A sending module, configured to send both the text bitstream and the feature bitstream to a decoding end, so that the decoding end executes the image processing method applied to the decoding end in any embodiment of the present application.

[0040] According to another aspect of the embodiments of the present application, there is provided an image encoding and decoding system, including an encoder and a decoder. The decoder is configured to execute the image processing method applied to the decoding end in any embodiment of the present application, and the encoder is configured to execute the image processing method applied to the encoding end in any embodiment of the present application.

[0041] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the method in any embodiment of the present application.

[0042] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the method in any embodiment of the present application.

[0043] One aspect of the technical solution provided by the embodiments of the present application may include the following beneficial effects:

[0044] The image processing method provided by the embodiments of the present application decodes the text bitstream from the encoding end through a first decoder to obtain the semantic information of the original image, and decodes the feature bitstream from the encoding end through a second decoder to obtain the feature decoding information of the original image. The texture information acquisition module processes the feature decoding information to obtain the texture information of the original image. Based on the latent layer representation diffusion model, the semantic information and the texture information are processed to obtain fusion information, and the image generation module processes the fusion information to obtain the decoded image, which can obtain a relatively clear image, greatly reduce the situation of image blurring, and has a good processing effect.

[0045] The above description is only an overview of the technical solution of the embodiment of the present application. In order to understand the technical means of the embodiment of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiment of the present application more obvious and understandable, the following specifically illustrates the specific implementation manners of the present application. Brief Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0047] Figure 1 Shows the flowchart of the image processing method according to an embodiment of the present application.

[0048] Figure 2 Shows the flowchart of processing the feature decoding information by the texture information acquisition module according to an embodiment of the present application.

[0049] Figure 3 Shows the flowchart of the image processing method according to another embodiment of the present application.

[0050] Figure 4 Shows the schematic diagram of the generative encoding and decoding framework structure according to an embodiment of the present application.

[0051] Figure 5 Shows the schematic diagram of the encoder framework structure according to an embodiment of the present application.

[0052] Figure 6 Shows the schematic diagram of the decoder framework structure according to an embodiment of the present application.

[0053] Figure 7 Shows the visualization diagram of the comparison result of comparing the processing result of the image processing method according to an embodiment of the present application with the processing results of other methods.

[0054] Figure 8 Shows the block diagram of the image processing device according to an embodiment of the present application.

[0055] Figure 9 Shows the block diagram of the image processing device according to another embodiment of the present application.

[0056] Figure 10 Shows the block diagram of the electronic device according to an embodiment of the present application.

[0057] Figure 11Schematic diagram of a computer-readable storage medium according to an embodiment of the present application is shown. Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0059] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the technical field to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0060] With the development of deep learning, end-to-end coding schemes based on neural networks have become increasingly mature. Through the optimization of non-linear transformations and entropy models, significant performance improvements have been achieved. However, with the increase in the compression ratio, the loss of data signals will lead to a significant decline in the quality of image and video. Especially under the condition of a compression ratio of thousands, the reconstructed image and video will have sharp distortions in visual characteristics such as edges, colors, and textures, and may also lead to the loss of key visual information such as concepts and semantics. Since the compression-domain representation based on signal transformation fails to explicitly model and retain more critical visual information, the bitstream still needs to transmit more feature information to ensure the reconstruction quality at the decoding end, which limits the compression efficiency of coding. To further improve the compression efficiency, it is necessary to go beyond the traditional optimization objectives and develop new image and video coding methods with higher compression efficiency to meet the growing data processing requirements and diverse application scenarios. The mainstream compression methods in the related art reconstruct images through the statistical characteristics of data. In the case of a high compression ratio, it is difficult to reconstruct images that meet the subjective effects of the human eye, and artifacts such as blurring and blocking effects are likely to occur.

[0061] In view of the problems existing in the related art, an embodiment of the present application provides an image processing method. By using a first decoder to decode the text bitstream from the encoding end, the semantic information of the original image is obtained. By using a second decoder to decode the feature bitstream from the encoding end, the feature decoding information of the original image is obtained. The texture information acquisition module processes the feature decoding information to obtain the texture information of the original image. Based on the latent representation diffusion model, the semantic information and the texture information are processed to obtain the fusion information. The image generation module processes the fusion information to obtain the decoded image. In this way, a clearer image can be obtained, the situation of image blurring is greatly reduced, and the processing effect is better.

[0062] The following describes an image processing method, apparatus, image encoding and decoding system, device, and storage medium according to an embodiment of the present application with reference to the accompanying drawings.

[0063] Reference Figure 1 As shown, an embodiment of the present application provides an image processing method applied to a decoding end. The decoding end includes a first decoder, a second decoder, and a generative diffusion model. The generative diffusion model includes a texture information acquisition module, a latent representation diffusion model, and an image generation module connected in sequence. The image processing method may include steps S10 - S50:

[0064] S10. Use the first decoder to decode the text bitstream from the encoding end to obtain the semantic information of the original image.

[0065] Exemplarily, the encoder converts the input image or image-related information (such as visual features of the image, semantic labels, etc.) into a compact text bitstream or latent representation. This text bitstream may be data after a certain form of encoding (such as natural language description, symbolic encoding, or vector representation), representing the high-level semantic information of the image. The decoder is a component that converts this compact text bitstream back into a structure or semantic information in a human-readable form. The decoder may include multiple decoding layers and relies on a certain form of neural network (such as Transformer, LSTM, etc.) to gradually convert the high-dimensional information back into the semantic representation of the image. The first decoder restores the semantic information of the image through step-by-step decoding of the bitstream. This semantic information may include the category, location, relationship, scene description, etc. of the objects in the image. For example, whether there is a cat in the image, whether the position of the cat is on the left, and whether the cat's state is sitting or standing, etc.

[0066] The decoder can use the Attention Mechanism, which helps to pay more attention to the important parts in the bitstream during decoding. During the decoding process, the decoder can learn the complex semantic structure of the image based on the context information. For example, the semantic information may not be limited to single object recognition, but also involve the relationships between objects or the understanding of the background scene.

[0067] Each decoding step during the decoding process may generate an intermediate representation, representing a certain part of the semantic information of the image. For example, through the first decoder, the rough structural information of the image (such as the positions and categories of the main objects) may be decoded first, and then further refined to decode more specific details (such as the spatial relationships between objects or more context information).

[0068] Finally, the semantic information of the original image decoded by the first decoder usually appears in a structured representation form, which may be a vector, matrix or text description containing labels of the image content, scene descriptions, object positions or attributes. This information provides a basis for subsequent tasks (such as image reconstruction, image generation or classification).

[0069] The first decoder decodes the text bitstream from the encoding end to obtain the semantic information of the original image, specifically including receiving and decoding the encoded text stream, extracting the high-level semantic information of the image, and applying context association and the attention mechanism, so as to generate an accurate semantic representation and lay a foundation for subsequent processing or tasks.

[0070] S20. Decode the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image.

[0071] Exemplarily, the feature bitstream from the encoding end contains high-level semantic feature information of the image, such as edges, shapes, textures, etc. If quantization or encoding processing is performed at the encoding end, the decoder will first perform inverse quantization or decoding operations. Inverse quantization is the process of restoring the quantized discrete values to values close to the original features, usually restoring them to floating-point numbers with higher precision. If the feature bitstream undergoes other transformations (such as wavelet transform, Fourier transform, etc.), the decoder will perform the corresponding inverse transformation operations to restore a spatial representation closer to the original features. The decoder reorganizes the decoded features into the spatial information of the original image or its corresponding feature maps, and these feature maps represent the local structures, textures, edges, etc. of the image in the spatial domain.

[0072] The decoder gradually reconstructs the detailed features of the original image (such as color, brightness, shape, etc.). By utilizing the previously extracted features and the restored high-level semantic information, it gradually fills in the missing information to restore the integrity of the image. The decoder also predicts the specific pixel values of the image based on the restored feature information, gradually refining the image restoration process. The predicted pixel values are processed and optimized multiple times between the decoding layers and finally approach the original image. After the above steps, the final output of the decoder will be the feature decoding information of the image, usually including the key detail features, structure, and texture of the image.

[0073] By using the first decoder and the second decoder to decode the text bitstream and the feature bitstream respectively, the semantic information and feature information of the original image can be decoded independently. This helps to retain more image details and semantic structures, making the decoded image more expressive and effectively avoiding information loss.

[0074] S30. Process the feature decoding information through the texture information acquisition module to obtain the texture information of the original image.

[0075] Processing the feature decoding information through the texture information acquisition module to obtain the texture information of the image helps to enhance the detail expressiveness of the image, making the decoded image visually more realistic and delicate. Especially in the decoding of low-resolution images, the enhancement of texture information can supplement and improve the quality of the image.

[0076] In some embodiments, the texture information acquisition module includes a latent space representation module, an image reconstruction module, and a feature extractor connected in sequence; refer to Figure 2 As shown, processing the feature decoding information through the texture information acquisition module to obtain the texture information of the original image may include steps S301 - S303:

[0077] S301. Input the feature decoding information into the latent space representation module for processing to obtain latent space representation information.

[0078] Processing the feature decoding information through the latent space representation module can efficiently represent the latent features of the image in a low-dimensional latent space. This helps to reduce data redundancy while retaining the main information of the image, thereby improving computational efficiency.

[0079] S302. Process the latent space representation information through the image reconstruction module to obtain image reconstruction information.

[0080] Converting the latent space representation information into image reconstruction information through the image reconstruction module can effectively restore the structure and details of the image, thereby enhancing the image information through the learned image features and improving the clarity and accuracy of the image during the reconstruction process.

[0081] S303. Process the image reconstruction information through the feature extractor to obtain the texture information of the original image.

[0082] Texture information is the detailed information in an image and has strong visual distinctiveness. Through step S303, the texture features in the image can be effectively extracted and restored, thereby providing a more accurate basis for subsequent image processing, recognition, or analysis.

[0083] Through the processes of these three steps S301 - S303, the extraction accuracy of texture information can be effectively improved, and the detail loss rate during the image reconstruction process can be reduced. The feature extraction module can pay more attention to details, which helps to improve the texture restoration quality of the image, especially for the image processing of complex textures. Using the latent space representation module to map the image information to a lower - dimensional space effectively reduces the computational amount. The step - by - step processing of the image reconstruction and feature extraction modules also helps to optimize the computing resources, improve the operation efficiency, and reduce unnecessary computational overhead.

[0084] S40. Process the semantic information and the texture information based on the latent layer representation diffusion model to obtain the fused information.

[0085] Fusing the semantic information and the texture information based on the latent layer representation diffusion model can efficiently combine the information in the latent layer space of the model, thereby being able to capture higher - level image features and improve the generation effect of the image. For example, the latent layer representation diffusion model can effectively process complex image features and details to generate smoother and more natural images.

[0086] In some embodiments, processing the semantic information and the texture information based on the latent layer representation diffusion model to obtain the fused information may include: through the latent layer representation diffusion model, based on the visual knowledge of the latent layer representation diffusion model, fuse the semantic information and the texture information to obtain the fused information. In this way, by integrating different types of information (such as visual knowledge, semantic information, and texture information) into the same representation space through the latent layer representation diffusion model, this multi - modal fusion method can provide a more rich and comprehensive input for the model, enhance the model's understanding ability of the image and its semantic level, and thus improve the accuracy and reliability of the fused information.

[0087] S50. Process the fused information through the image generation module to obtain the decoded image.

[0088] Through the image generation module, the fused information is further processed and the decoded image is generated. By using the fused semantic information and texture information, the decoded image can be optimized in both semantic consistency and visual details.

[0089] In some embodiments, the image generation module includes an image generator and an enhancement network module connected in sequence; processing the fusion information through the image generation module to obtain a decoded image may include:

[0090] S501. Process the fusion information through the image generator to obtain an initially generated image;

[0091] S502. Remove block artifacts from the initially generated image through the enhancement network module to obtain a decoded image.

[0092] During the image generation process, block artifacts (also known as mosaic effects) usually appear during the image decoding process, especially in the restoration of low-resolution images or high-compression ratio images. By using the enhancement network module to remove block artifacts from the initially generated image, this kind of artifact can be effectively eliminated, making the final image smoother and more natural. After removing the block artifacts, the visual quality of the image will be significantly improved, especially in the details, and more texture and structure information can be retained, making the decoded image more clear and realistic.

[0093] In some embodiments, the encoding end includes a text-to-image large model, a latent space representation extraction model, a first encoder connected to the text-to-image large model, and a second encoder connected to the latent space representation extraction model; the text code stream is obtained by encoding first semantic information through the first encoder; the first semantic information is the semantic information of the original image extracted through the text-to-image large model.

[0094] The image processing method of this embodiment can effectively reduce information loss during the decoding process through a multi-stage processing method, can effectively improve the quality, detail performance and visual effect of the decoded image, and make the finally generated decoded image closer to the original image.

[0095] The image processing method provided by the embodiments of this application can obtain a relatively clear image. In the case of a high compression ratio, it can reconstruct an image that meets the subjective effect of the human eye, greatly reducing the probability of artifacts such as image blurring and block artifacts, and the processing effect is good.

[0096] Another embodiment of this application provides an image processing method, which is applied to the encoding end. The encoding end includes a text-to-image model, a latent space representation extraction model, a first encoder connected to the text-to-image model, and a second encoder connected to the latent space representation extraction model. Refer to Figure 3 As shown, this method may include steps 1-step 5:

[0097] Step 1. Process the original image through the text-to-image model to obtain semantic information.

[0098] The image - to - text model is a model that converts images into text or semantic descriptions, and it is used to extract semantic information (such as objects, scenes, actions, etc.) in the image through the analysis of the image. The original image is first input into the image - to - text model. The image - to - text model recognizes various objects, relationships, and contexts in the original image and converts this information into semantic information that can be further processed. This semantic information is a high - level understanding of the original image, including information such as the relationships between various elements in the image, object categories, scene types, etc.

[0099] Step 2: Process the original image through the latent space representation extraction model to obtain texture information.

[0100] The latent space representation extraction model can extract low - level texture patterns, colors, shapes, details, and other feature information from the original image through a deep learning model (such as a convolutional neural network, etc.). These feature information do not directly express the object content of the image (different from the semantic information in Step 1), but express the texture features of the original image, such as the texture pattern of the image, detailed structure, color distribution, etc.

[0101] Step 3: Encode the semantic information using the first encoder to obtain a text code stream.

[0102] The text code stream can be a format that converts semantic information through an encoding process into a format that is convenient for transmission and processing. This is usually a compressed symbol sequence or feature vector, containing a simplified representation of all the semantic information in the image.

[0103] Step 4: Encode the texture information using the second encoder to obtain a feature code stream.

[0104] The second encoder is used to encode the texture information extracted in Step 2. It can convert visual details and texture features into a feature code stream. This process compresses the detail information and converts it into another format suitable for storage and processing.

[0105] Step 5: Send both the text code stream and the feature code stream to the decoding end so that the decoding end executes the image - processing method applied to the decoding end in any of the above - mentioned embodiments.

[0106] Send the text code stream and the feature code stream to the decoding end for the decoding end to decode these two code streams respectively and restore the complete information of the image. The decoding process will use semantic decoding and texture decoding strategies respectively to combine semantic and texture information to obtain the decoded image.

[0107] Exemplarily, for an input image, it can be first segmented to obtain multiple segmented images (for example, an input image can be segmented into 4 segmented images). Each segmented image is input into the encoding end for processing, and finally, the decoded images corresponding to each segmented image output at the decoding end are stitched together to obtain a complete image corresponding to the input image.

[0108] Another embodiment of the present application provides an image processing method. The semantic information of an image is generated by a text-to-image large model, the texture information of the image is modeled through the latent space representation of the image, and the semantic information and texture information of the image are fused based on the visual knowledge of the generative diffusion large model, thereby achieving efficient compression of the image.

[0109] The generative encoding and decoding framework of this embodiment is as Figure 4 shown. The generative encoding and decoding framework includes a text-to-image large model, a latent space representation extraction model, two encoders, two decoders, and a generative diffusion model. The input image is respectively input into the text-to-image large model and the latent space representation extraction model. The text-to-image large model and the latent space representation extraction model are respectively connected to an encoder. The encoder connected to the text-to-image large model outputs a text bitstream, and this text bitstream is input into a decoder for decoding to obtain semantic information. The encoder connected to the latent space representation extraction model outputs a feature bitstream, and this feature bitstream is input into another decoder for decoding to obtain texture information. Both the semantic information and the texture information are input into the generative diffusion model for processing to obtain a reconstructed image.

[0110] The encoding end framework of this embodiment is as Figure 5 shown. The encoding end framework includes a text-to-image model, a latent space representation generation module, and two AEs (autoencoders, abbreviated as AE). The text-to-image model is connected to one AE, and the latent space representation generation module is connected to the other AE.

[0111] The decoding end framework is as Figure 6As shown, the decoding - end framework includes two ADs (Auto Decoders), an image reconstruction module, and a latent - representation diffusion model, an image generator, and an enhancement module connected in sequence. The input end of the latent - representation diffusion model is connected to a feature extractor, and the image reconstruction module is connected to the latent - representation diffusion model through another feature extractor. One AD decodes the text bitstream to obtain the semantic information of the decoded text, and this semantic information is input into the latent - representation diffusion model; another AD decodes the feature bitstream to obtain the latent - space representation, and then the latent - space representation is input into the image reconstruction module for processing. The information output by the image reconstruction module is input into the feature extractor connected to the image reconstruction module for processing, and the output texture information is input into the latent - representation diffusion model; another feature extractor inputs the visual information into the latent - representation diffusion model; the latent - representation diffusion model processes the received semantic information, visual information, and texture information, and inputs the output processing result into the image generator for processing. The image generator inputs the output processing result into the enhancement module for processing, and the enhancement module outputs the decoded image.

[0112] Exemplarily, first, the text and the latent - space representation are decoded from the bitstream, and then a feature decoder is used to reconstruct a low - quality image using the latent - space representation. Then, the enhanced image and text are input into LAControlNet with pre - trained SD2.1 to sequentially generate a denoised hidden representation and the original pixels. During the training process, LAControlNet is trained and other parameters are fixed. Finally, the two representations are input into the diffusion model, and the diffusion model iteratively denoises the noisy image content. The image generator uses DiffBIR, a blind image restoration model built based on Stable Diffusion2.1.

[0113] To obtain compact and accurate text, the method of the embodiments of the present application respectively uses the 1.5 version of mPLUG - Owl and StableDiffusion as the image - to - text and text - to - image models. The decoding end stitches the segmented images to generate a complete image, and uses an enhancement network to enhance the quality of the image to remove the block effect.

[0114] Four image datasets are selected for training and testing, including the OpenImages, DiffusionDB, Kodak, and AVS datasets. OpenImages is the training dataset, and the DiffusionDB, Kodak, and AVS datasets are the test datasets. DiffusionDB includes 100 AIGC images. AVS is obtained by taking the first frame of each test sequence. The bit - depth of the datasets is 8 bits. The detailed introduction is shown in the following table.

[0115]

[0116] Experiments and comparisons were conducted on the above dataset. The methods compared with the method of this embodiment include JPEG, HEVC (HM-16.0), and VVC (VTM-11.0). The visualization diagram of the comparison results is as shown in Figure 7 shown below. Figure 4 In [figure], NIGM represents the curve corresponding to the method of this embodiment. According to Figure 7 the results shown, it can be determined that the method proposed in this embodiment has a better processing effect on images than the other three methods. The method proposed in this embodiment is far superior to other coding methods in terms of subjective human eye indicators. This embodiment adopts generative coding for image data and realizes efficient image coding by using the text and image compact representations generated by the text-to-image large model and the rich visual information of the generative diffusion model.

[0117] This embodiment will provide semantic information in text representation and texture structure information in latent space feature representation, and adopts a diffusion model based on text-latent space representation to achieve efficient image compression. This embodiment uses the text and image compact representations generated by the text-to-image large model and the visual information ability of the generative extension model to achieve efficient image coding, and exceeds the effects of related technologies in terms of subjective human eye effects, which provides strong support for the applicability of generative coding.

[0118] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated in this article.

[0119] Refer to Figure 8 shown below. Another embodiment of the present application provides an image processing device, which is applied to the decoding end. The decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the device includes:

[0120] A first decoding module, configured to decode the text bitstream from the encoding end through the first decoder to obtain the semantic information of the original image;

[0121] A second decoding module, configured to decode the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image;

[0122] A texture information acquisition module, configured to process the feature decoding information through the texture information acquisition module to obtain the texture information of the original image;

[0123] A fusion information acquisition module, configured to process the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information;

[0124] An image generation module, configured to process the fusion information through the image generation module to obtain a decoded image.

[0125] Exemplarily, the texture information acquisition module includes a latent space representation module, an image reconstruction module, and a feature extractor connected in sequence; the texture information acquisition module includes:

[0126] A latent space representation information acquisition unit, configured to input the feature decoding information into the latent space representation module for processing to obtain latent space representation information;

[0127] An image reconstruction information acquisition unit, configured to process the latent space representation information through the image reconstruction module to obtain image reconstruction information;

[0128] A texture information acquisition unit, configured to process the image reconstruction information through the feature extractor to obtain the texture information of the original image.

[0129] Exemplarily, the fusion information acquisition module is further specifically configured to fuse the semantic information and the texture information through the hidden layer representation diffusion model based on the visual knowledge of the hidden layer representation diffusion model to obtain fusion information.

[0130] Exemplarily, the image generation module includes an image generator and an enhancement network module connected in sequence; the process of processing the fusion information through the image generation module to obtain a decoded image includes:

[0131] Processing the fusion information through the image generator to obtain an initial generated image;

[0132] Removing block artifacts from the initial generated image through the enhancement network module to obtain a decoded image.

[0133] Exemplarily, the encoding end includes a text-to-image large model, a latent space representation extraction model, a first encoder connected to the text-to-image large model, and a second encoder connected to the latent space representation extraction model; the text bitstream is encoded by the first encoder for the first semantic information; the first semantic information is the semantic information of the original image extracted by the text-to-image large model.

[0134] The image processing device provided by the embodiment of the present application decodes the text bitstream from the encoding end through the first decoder to obtain the semantic information of the original image, and decodes the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image. The texture information acquisition module processes the feature decoding information to obtain the texture information of the original image. The fused information is obtained by processing the semantic information and the texture information based on the hidden layer representation diffusion model. The decoded image is obtained by processing the fused information through the image generation module, and a relatively clear image can be obtained, significantly reducing the situation of image blurring, and the processing effect is good.

[0135] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0136] Refer to Figure 9 As shown, another embodiment of the present application provides an image processing device, which is applied to the encoding end. The encoding end includes a text-to-image model, a hidden space representation extraction model, a first encoder connected to the text-to-image model, and a second encoder connected to the hidden space representation extraction model. The device includes:

[0137] A semantic information acquisition module, configured to process the original image through the text-to-image model to obtain semantic information;

[0138] A texture information acquisition module, configured to process the original image through the hidden space representation extraction model to obtain texture information;

[0139] A text bitstream acquisition module, configured to encode the semantic information by using the first encoder to obtain a text bitstream;

[0140] A feature bitstream acquisition module, configured to encode the texture information by using the second encoder to obtain a feature bitstream;

[0141] A sending module, configured to send both the text bitstream and the feature bitstream to the decoding end, so that the decoding end executes the image processing method applied to the decoding end in any implementation manner of the present application.

[0142] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0143] Another embodiment of the present application provides an image encoding and decoding system, including an encoder and a decoder. The encoder is configured to execute the method applied to the encoder in any implementation manner of the present application, and the decoder is configured to execute the image processing method applied to the decoding end in any implementation manner of the present application.

[0144] The decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the image processing method applied to the decoding end includes: decoding the text bitstream from the encoding end through the first decoder to obtain the semantic information of the original image; decoding the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image; processing the feature decoding information through the texture information acquisition module to obtain the texture information of the original image; processing the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fused information; and processing the fused information through the image generation module to obtain the decoded image.

[0145] The encoding end includes an image-to-text model, a latent space representation extraction model, a first encoder connected to the image-to-text model, and a second encoder connected to the latent space representation extraction model. The image processing method applied to the encoding end includes: processing the original image through the image-to-text model to obtain semantic information; processing the original image through the latent space representation extraction model to obtain texture information; encoding the semantic information using the first encoder to obtain a text bitstream; encoding the texture information using the second encoder to obtain a feature bitstream; and sending both the text bitstream and the feature bitstream to the decoding end, so that the decoding end executes the image processing method applied to the decoding end according to any embodiment of this application.

[0146] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities can be referred to each other. For the sake of brevity, they are not elaborated herein.

[0147] Another embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the image processing method according to any of the above embodiments.

[0148] Refer to Figure 10 As shown, the electronic device 10 may include: a processor 100, a memory 101, a bus 102, and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102; a computer program executable on the processor 100 is stored in the memory 101, and when the processor 100 runs the computer program, it executes the method provided in any of the foregoing embodiments of this application.

[0149] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0150] The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 101 is used to store a program. After receiving an execution instruction, the processor 100 executes the program. Any implementation manner of the method disclosed in any embodiment of the present application can be applied to the processor 100 or implemented by the processor 100.

[0151] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 100 or by instructions in software form. The above-mentioned processor 100 can be a general-purpose processor, which can include a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the above method.

[0152] The electronic device provided by the embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.

[0153] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image processing method according to any of the above embodiments can be implemented. Refer to Figure 11 As shown, the computer-readable storage medium shown is an optical disc 20, on which a computer program (i.e., a program product) is stored. When the computer program runs on a processor, it will execute the method provided by any of the foregoing embodiments.

[0154] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0155] The computer-readable storage medium provided in the above embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored thereon.

[0156] It should be noted that:

[0157] The term "module" is not intended to be limited to a specific physical form. Depending on the specific application, a module can be implemented as hardware, firmware, software, and / or a combination thereof. In addition, different modules can share common components or even be implemented by the same components. There may or may not be a clear boundary between different modules.

[0158] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the examples based herein. Based on the above description, the structure required to construct such a device is obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of a specific language above is to disclose the best implementation mode of the present application.

[0159] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0160] The above embodiments only represent the implementation manners of the present application, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image processing method, characterized in that, Applied to the decoding end, the decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the texture information acquisition module includes a hidden space representation module, an image reconstruction module, and a feature extractor connected in sequence; the method includes: Decode the text bitstream from the encoding end through the first decoder to obtain the semantic information of the original image; Decode the feature bitstream from the encoding end through the second decoder to obtain the feature decoding information of the original image; Process the feature decoding information through the texture information acquisition module to obtain the texture information of the original image; Based on the hidden layer representation diffusion model, process the semantic information and the texture information to obtain fusion information; Process the fusion information through the image generation module to obtain the decoded image; The process of obtaining the texture information of the original image by processing the feature decoding information through the texture information acquisition module includes: Input the feature decoding information into the hidden space representation module for processing to obtain hidden space representation information; Process the hidden space representation information through the image reconstruction module to obtain image reconstruction information; Process the image reconstruction information through the feature extractor to obtain the texture information of the original image.

2. The method according to claim 1, characterized in that, The process of obtaining fusion information by processing the semantic information and the texture information based on the hidden layer representation diffusion model includes: Through the hidden layer representation diffusion model, based on the visual knowledge of the hidden layer representation diffusion model, fuse the semantic information and the texture information to obtain fusion information.

3. The method according to claim 1, wherein The image generation module includes an image generator and an enhancement network module connected in sequence; the process of obtaining the decoded image by processing the fusion information through the image generation module includes: Process the fusion information through the image generator to obtain an initial generated image; Perform block effect removal on the initial generated image through the enhancement network module to obtain the decoded image.

4. The method according to claim 1, wherein The encoding end includes a text-to-image large model, a hidden space representation extraction model, a first encoder connected to the text-to-image large model, and a second encoder connected to the hidden space representation extraction model; The text bitstream is encoded by the first encoder for the first semantic information; the first semantic information is the semantic information of the original image extracted by the text-to-image large model.

5. An image processing method, characterized in that, Applied to the encoding end, the encoding end includes a text-to-image model, a hidden space representation extraction model, a first encoder connected to the text-to-image model, and a second encoder connected to the hidden space representation extraction model, the method includes: Process the original image through the text-to-image model to obtain semantic information; Process the original image through the hidden space representation extraction model to obtain texture information; Encode the semantic information using the first encoder to obtain a text bitstream; Encode the texture information using the second encoder to obtain a feature bitstream; Send both the text bitstream and the feature bitstream to a decoding end, so that the decoding end executes the image processing method described in any one of claims 1-4.

6. An image processing apparatus, characterized in that, Applied to a decoding end, the decoding end includes a first decoder, a second decoder, and a generative diffusion model; the generative diffusion model includes a texture information acquisition module, a hidden layer representation diffusion model, and an image generation module connected in sequence; the device includes: A first decoding module, configured to decode the text bitstream from an encoding end through the first decoder to obtain semantic information of an original image; A second decoding module, configured to decode the feature bitstream from the encoding end through the second decoder to obtain feature decoding information of the original image; A texture information acquisition module, configured to process the feature decoding information through the texture information acquisition module to obtain texture information of the original image; A fusion information acquisition module, configured to process the semantic information and the texture information based on the hidden layer representation diffusion model to obtain fusion information; An image generation module, configured to process the fusion information through the image generation module to obtain a decoded image; The texture information acquisition module includes a hidden space representation module, an image reconstruction module, and a feature extractor connected in sequence; the texture information acquisition module is further configured to: Input the feature decoding information into the hidden space representation module for processing to obtain hidden space representation information; Process the hidden space representation information through the image reconstruction module to obtain image reconstruction information; Process the image reconstruction information through the feature extractor to obtain the texture information of the original image.

7. An image processing apparatus, characterized in that, Applied to an encoding end, the encoding end includes an image-to-text model, a hidden space representation extraction model, a first encoder connected to the image-to-text model, and a second encoder connected to the hidden space representation extraction model; the device includes: A semantic information acquisition module, configured to process an original image through the image-to-text model to obtain semantic information; A texture information acquisition module, configured to process the original image through the hidden space representation extraction model to obtain texture information; A text bitstream acquisition module, configured to encode the semantic information by using the first encoder to obtain a text bitstream; A feature bitstream acquisition module, configured to encode the texture information by using the second encoder to obtain a feature bitstream; A sending module, configured to send both the text bitstream and the feature bitstream to a decoding end, so that the decoding end executes the image processing method described in any one of claims 1-4.

8. An image encoding and decoding system, characterized in that, Includes an encoder and a decoder, the decoder is configured to execute the image processing method described in any one of claims 1-4, and the encoder is configured to execute the image processing method described in claim 5.

9. An electronic device, characterized in that, Includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the method described in any one of claims 1-5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to implement the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Image coding, decoding, reconstruction and analysis method and system, and electronic equipment

    CN113660486A