Picture processing methods, apparatuses and system, and electronic device and storage medium
By extracting and encoding the one-dimensional feature vectors and multi-dimensional feature maps of image blocks, two-layer code streams are formed, which solves the problem of large information loss in the existing technology at low code rates, and achieves better image reconstruction effect and visual experience.
Patent Information
- Application Number
- PCT/CN2024/136446
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-26
AI Technical Summary
Existing image coding standards and neural network coding technologies have large information losses at low code rates, and poor image reconstruction effect and visual experience.
The one-dimensional feature vector of the original image block is extracted, and transformed into a multi-dimensional feature map through the vector, and quantized encoding and discrete encoding are performed respectively to form a two-layer code stream to achieve efficient image compression.
Through the encoding method of two-layer code streams, information loss can be reduced at low code rates, and the visual effect and visual experience of image reconstruction can be improved.
Smart Images

Figure CN2024136446_26062025_PF_FP_ABST
Abstract
Description
Image processing method, device, system, electronic device and storage medium Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, device, system, electronic device and storage medium. Background Art
[0002] With the development of artificial intelligence, traditional image coding standards (such as JPEG (Joint Photographic Experts Group), MPEG (Moving Picture Expert Group), AVS (Audio Video Coding Standard), and BPG (Better Portable Graphics)) are pixel-level optimizations for human vision and cannot effectively support various machine vision requirements. Attempts are underway to incorporate neural networks into image coding to achieve more intelligent and efficient image coding. Therefore, deep learning-based neural network coding has become a research focus. However, both traditional image coding standards and neural network coding technologies use a single bitstream, which results in significant information loss at low bitrates, poor image reconstruction, and a poor visual experience. Summary of the Invention
[0003] The purpose of this application is to propose an image processing method, device, electronic device and storage medium to address the deficiencies of the above-mentioned prior art, and this purpose is achieved through the following technical solutions.
[0004] A first aspect of the present application provides an image processing method, applied to an encoding end, the method comprising:
[0005] Extract the one-dimensional feature vector of the original image block;
[0006] transforming the original image block into a multidimensional feature map based on the one-dimensional feature vector;
[0007] Performing quantization encoding on the one-dimensional feature vector to obtain a first code stream;
[0008] Performing discrete encoding on the multidimensional feature map to obtain a second code stream;
[0009] The first code stream and the second code stream are sent to a decoding end.
[0010] Based on the image processing method described in the first aspect above, this application has at least the following beneficial effects or advantages:
[0011] By extracting a one-dimensional feature vector from the original image block, a low-dimensional, spatially independent vector that contains the image's texture structure information, the original image block is then transformed into a multidimensional feature map based on this one-dimensional feature vector. This multidimensional feature map contains not only the image's spatially related information but also its high-frequency information. Considering that the spatially independent vector and the multidimensional feature map each contain information at different levels of the image, efficient compression of the spatially independent vector and the multidimensional feature map is achieved by quantizing and encoding the one-dimensional feature vector into a first bitstream and discretely encoding the multidimensional feature map into a second bitstream. Since the encoded bitstream consists of two layers containing information from different layers of the image, even at low bitrates, little information is lost. Therefore, the image reconstructed from the two layers can improve visual effects and enhance the visual experience.
[0012] Optionally, transforming the original image block into a multidimensional feature map based on the one-dimensional feature vector includes: performing multiple iterative noise additions on the original image block using the one-dimensional feature vector through a preset diffusion model to obtain a multidimensional feature map.
[0013] Optionally, the diffusion model includes a first noise prediction network and a noise addition layer; the preset diffusion model is used to perform multiple iterative noise additions on the original image block using the one-dimensional feature vector, including: for each noise addition process, predicting the current noise based on the previous noise addition result and the one-dimensional feature vector through the first noise prediction network; adding the noise output by the first noise prediction network to the previous noise addition result through the noise addition layer; wherein, for the first noise addition process, the previous noise addition result is the original image block.
[0014] Optionally, the noise output by the first noise prediction network is added to the last noise addition result through the noise addition layer, including: obtaining the noise level coefficient used last and the noise level coefficient used currently; and obtaining the current noise addition result based on the noise level coefficient used last, the noise level coefficient used currently, the noise output by the first noise prediction network, and the last noise addition result.
[0015] Optionally, the performing quantization encoding on the one-dimensional feature vector to obtain the first code stream includes: performing quantization encoding on the one-dimensional feature vector using a preset quantization encoding model to obtain the first code stream.
[0016] Optionally, the quantization coding model includes a feature coding network, a quantization layer, and an entropy estimation network; the quantization coding of the one-dimensional feature vector by the preset quantization coding model to obtain a first code stream includes: downsampling the one-dimensional feature vector by the feature coding network to obtain downsampled features; converting the feature values of the downsampled features into integer values by the quantization layer to obtain quantized features; and encoding the quantized features into a first code stream by the entropy estimation network.
[0017] Optionally, the discretely encoding the multidimensional feature map to obtain the second code stream includes: discretely encoding the multidimensional feature map using a preset discrete coding model to obtain the second code stream.
[0018] Optionally, the discrete coding model includes a feature coding network, a discrete representation layer, and a code stream conversion layer; the discrete encoding of the multidimensional feature map by a preset discrete coding model to obtain a second code stream includes: downsampling the multidimensional feature map by the feature coding network to obtain downsampled features; discretely representing the downsampled features by the discrete representation layer to obtain an index matrix; and converting the index matrix into a second code stream by the code stream conversion layer.
[0019] Optionally, the downsampled features are discretely represented by the discrete representation layer to obtain an index matrix, including: matching the most similar vector for each dimension of the feature vector of the downsampled features in a preset feature dictionary; and obtaining an index matrix using the index of the vector matching the feature vector of each dimension.
[0020] A second aspect of the present application provides an image processing method, which is applied to a decoding end. The method includes:
[0021] receiving a first code stream and a second code stream from an encoding end;
[0022] Performing inverse quantization decoding on the first code stream to obtain a one-dimensional feature vector;
[0023] Performing inverse discrete decoding on the second code stream to obtain a multi-dimensional feature map;
[0024] Perform image reconstruction on the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block.
[0025] Based on the image processing method described in the second aspect above, this application has at least the following beneficial effects or advantages:
[0026] Since the received first and second bitstreams contain information at different levels of the image, the first bitstream contains texture and style information of the image, and the second bitstream contains high-frequency information and structural information of the image, the reconstructed image obtained by decoding the first and second bitstreams can improve the image visual effect and does not lose much information even under bandwidth constraints.
[0027] Optionally, performing inverse quantization decoding on the first bitstream to obtain a one-dimensional feature vector includes: performing inverse quantization decoding on the first bitstream using a preset inverse quantization decoding model to obtain a one-dimensional feature vector;
[0028] The inverse quantization decoding model includes an entropy decoding network, an inverse quantization layer, and a feature decoding network; the inverse quantization decoding of the first code stream by the preset inverse quantization decoding model to obtain a one-dimensional feature vector includes: decoding the first code stream into quantized features by the entropy decoding network; converting the quantized features into floating-point values by the inverse quantization layer to obtain inverse quantized features; and upsampling the inverse quantized features by the feature decoding network to obtain a one-dimensional feature vector.
[0029] Optionally, performing inverse discrete decoding on the second code stream to obtain a multidimensional feature map includes: performing inverse discrete decoding on the second code stream using a preset discrete decoding model to obtain a multidimensional feature map;
[0030] The discrete decoding model includes a code stream decoding layer, an inverse discrete representation layer, and a feature decoding network; the inverse discrete decoding of the second code stream by a preset discrete decoding model to obtain a multidimensional feature map includes: decoding the second code stream into an index matrix by the code stream decoding layer; converting the index matrix into continuous features based on a preset feature dictionary by the inverse discrete representation layer; and upsampling the continuous features by the feature decoding network to obtain a multidimensional feature map.
[0031] Optionally, reconstructing the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block includes: performing multiple iterative denoising operations on the multidimensional feature map using the one-dimensional feature vector through a preset diffusion model to obtain a reconstructed image block.
[0032] Optionally, the diffusion model includes a second noise prediction network and a denoising layer; the multidimensional feature map is iteratively denoised multiple times using the one-dimensional feature vector through the preset diffusion model to obtain a reconstructed image block, including: for each denoising process, predicting the current noise based on the previous denoising result and the one-dimensional feature vector through the second noise prediction network; denoising the previous denoising result through the denoising layer using the noise output by the second noise prediction network; wherein, for the first denoising process, the previous denoising result is the multidimensional feature map.
[0033] Optionally, the denoising layer uses the noise output by the second noise prediction network to denoise the previous denoising result, including: obtaining the noise level coefficient used last time and the noise level coefficient used currently; and obtaining the current denoising result based on the noise level coefficient used last time, the noise level coefficient used currently, the noise output by the second noise prediction network, and the previous denoising result.
[0034] A third aspect of the present application provides an image processing device, applied to an encoding end, comprising:
[0035] A first extraction module is used to extract a one-dimensional feature vector of the original image block;
[0036] A second extraction module, configured to transform the original image block into a multidimensional feature map based on the one-dimensional feature vector;
[0037] A first encoding module, configured to perform quantization encoding on the one-dimensional feature vector to obtain a first code stream;
[0038] A second encoding module, configured to perform discrete encoding on the multidimensional feature map to obtain a second code stream;
[0039] The sending module is used to send the first code stream and the second code stream to the decoding end.
[0040] A fourth aspect of the present application provides an image processing device, applied to a decoding end, comprising:
[0041] A receiving module, receiving a first code stream and a second code stream from an encoding end;
[0042] A first decoding module, configured to perform inverse quantization decoding on the first bit stream to obtain a one-dimensional feature vector;
[0043] A second decoding module, configured to perform inverse discrete decoding on the second code stream to obtain a multidimensional feature map;
[0044] A reconstruction module is used to perform image reconstruction on the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block.
[0045] A fifth aspect of the present application provides an image processing system, comprising:
[0046] An encoding end, configured to execute the image processing method described in the first aspect;
[0047] The decoding end is used to execute the image processing method described in the second aspect.
[0048] The sixth aspect of the present application proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first or second aspect above.
[0049] The seventh aspect of the present application proposes a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method as described in the first or second aspect above.
[0050] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0052] FIG1 is a flow chart of an embodiment of an image processing method according to an exemplary embodiment;
[0053] FIG2 is a schematic structural diagram of a diffusion model according to an exemplary embodiment;
[0054] FIG3 is a schematic structural diagram of a quantization coding model according to an exemplary embodiment;
[0055] FIG4 is a schematic structural diagram of a discrete coding model according to an exemplary embodiment;
[0056] FIG5 is a flow chart of another image processing method according to an exemplary embodiment;
[0057] FIG6 is a schematic structural diagram of a quantization decoding model according to an exemplary embodiment;
[0058] FIG7 is a schematic structural diagram of a discrete decoding model according to an exemplary embodiment;
[0059] 8 to 11 are schematic diagrams showing how to implement image processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0060] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0061] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0062] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0063] Example 1:
[0064] FIG1 is a flowchart of an embodiment of an image processing method according to an exemplary embodiment. The image processing method is applied to an encoding end, which is used to compress and encode the original image into a code stream. As shown in FIG1 , the image processing method includes the following steps:
[0065] Step 101: extracting a one-dimensional feature vector of the original image block.
[0066] In this step, the original image block can be one of the blocks obtained by dividing the image to be encoded according to a preset block size, and its size is the preset block size, for example, 32*32. The visual content representation of this one-dimensional feature vector includes the image texture structure. Due to the relatively low dimension of the feature, the spatial domain correlation of the image represented by it is relatively weak, and the image is spatially unrelated.
[0067] In an optional embodiment, a preset convolutional encoder can be used to extract features from the original image block to obtain a one-dimensional feature vector. Specifically, the feature extraction formula is as follows: t =Enc(x0)
[0068] Among them, Enc() represents the convolutional encoder, and x0 represents the input original image block.
[0069] Step 102: transforming the original image block into a multi-dimensional feature map based on the one-dimensional feature vector.
[0070] In this step, the multidimensional feature map contains the visual content representation of the image high-frequency information and image structure information in the original image block. It is worth noting that the size of the multidimensional feature map is consistent with the size of the original image block.
[0071] In an optional embodiment, the one-dimensional feature vector may be used to perform multiple iterations of noise addition on the original image block using a preset diffusion model to obtain a multi-dimensional feature map.
[0072] Among them, the diffusion model includes a forward diffusion process and a backward process. In the forward diffusion process, the original image block is iteratively denoised, and the resulting multidimensional feature map is a noise map. Since the diffusion model has the ability to retain the semantic structure of the data and uses a one-dimensional feature vector to modulate the characteristic distribution of the noise instead of adding randomly sampled Gaussian noise each time, the noise map with an approximate normal distribution obtained after multiple iterative denoising of the original image contains the image's high-frequency information and image structure information, while also maintaining the image's spatial resolution, and there is spatial correlation between features.
[0073] In a specific embodiment, the diffusion model shown in Figure 2 includes a first noise prediction network and a noise addition layer in the forward diffusion process. During multiple iterative noise addition processes on the original image block, for each noise addition process, the first noise prediction network predicts the current noise based on the previous noise addition result and the one-dimensional feature vector, and then the noise output by the first noise prediction network is added to the previous noise addition result through the noise addition layer.
[0074] For the first noisy process, the result of the previous noisy process is the original image block. By replacing randomly sampled Gaussian noise with deterministic noise estimated by a noise prediction network in the forward diffusion process, the traditional random diffusion process is transformed into a deterministic encoding process. Since the noise prediction introduces a one-dimensional feature vector and the conditions of the previous noisy result to modulate the feature distribution of the noise prediction network, the predicted noise can retain spatially relevant information and rich structural information. As a result, the noise map obtained after multiple iterative noisy processes has spatial correlation and also contains high-frequency information and image structure information.
[0075] Specifically, the specific noise addition implementation process for the noise addition layer includes: obtaining the noise level coefficient used last time and the noise level coefficient used currently, and then obtaining the current noise addition result based on the noise level coefficient used last time, the noise level coefficient used currently, the noise output by the first noise prediction network, and the previous noise addition result.
[0076] The noise addition formula can be:
[0077] Among them, x t Indicates the current noise addition result, x t-1 Indicates the last noise addition result, ε θ (x t-1 , t-1, z t ) represents the noise output by the first noise prediction network, z t represents the one-dimensional feature vector, α t Indicates the current noise level coefficient, α t-1 Indicates the previous noise level coefficient. It is worth noting that the noise level coefficient indicates the degree of noise, ranging from 0 to 1, and the noise level coefficient used increases gradually, that is, the current noise level coefficient is greater than the previous noise level coefficient.
[0078] The above-mentioned noise addition formula is only an exemplary description. Other formulas that can achieve noise addition results are also within the protection scope of this application. Therefore, this application does not limit the specific form of the loading formula.
[0079] It should be noted that, as can be seen from Figure 2, the reverse process of the diffusion model (i.e., the denoising process) also uses a noise prediction network for noise prediction for denoising. This reverse process is described in detail when describing image reconstruction. To better achieve image reconstruction and maintain consistency between denoising and denoising, the first noise prediction network can use the noise prediction network used in the reverse process. In other words, the first noise prediction network is the same as the second noise prediction network. Of course, the first noise prediction network can also use a different network structure from the second noise prediction network.
[0080] Based on the above description, by using the diffusion model to drive the efficient representation of image visual content, since the diffusion model does not need to consider the scene content of the image in the process of generating visual content representation, it can achieve efficient representation of any image scene content. Compared with the existing generative image coding technology, the learning ability of complex data distribution is limited, and the model parameters need to be optimized for specific image scenes, which cannot be applied to complex and changeable content scenes. The embodiment of the present application uses the diffusion model for efficient representation of image visual content, which is more adaptable.
[0081] In addition, since the diffusion model is used to drive the image visual content representation, there is no need to consider the image scene content. The image can be divided into multiple image blocks and encoded separately. Therefore, even images with high-resolution content can be efficiently encoded. Compared with the existing generative image coding technology, the training of high-resolution content is limited by resources, algorithm design and other issues, and cannot be applied to high-resolution image content coding. The embodiment of the present application uses a diffusion model to implement coding, which can achieve efficient coding of content of any resolution.
[0082] Step 103: quantize and encode the one-dimensional feature vector to obtain a first bit stream.
[0083] In this step, since the spatial correlation of the one-dimensional feature vector is very weak, but the local correlation is relatively strong, a coding method based on local correlation can be used to quantize and encode the feature vector in the feature dimension, such as an end-to-end image coding method, a block-based video intra-frame coding method, etc.
[0084] In an optional embodiment, the one-dimensional feature vector may be quantized and encoded using a preset quantization encoding model to obtain a first code stream, thereby achieving end-to-end image encoding.
[0085] In one embodiment, the quantization coding model shown in FIG3 includes a feature coding network, a quantization layer, and an entropy estimation network. The encoding process includes: downsampling a one-dimensional feature vector through the feature coding network to obtain a downsampled feature; then converting the feature values of the downsampled feature into integer values through the quantization layer to obtain a quantized feature; and finally, encoding the quantized feature into a first bitstream through the entropy estimation network.
[0086] Among them, the values of the integer eigenvalues in the quantized features are diverse, so it is necessary to use the probability table of the entropy estimation network to encode the integer eigenvalues in the quantized features into a binary first code stream.
[0087] As mentioned above, the one-dimensional feature vector contains image texture information, so the first bitstream obtained by quantizing and encoding the one-dimensional feature vector will carry the image texture information, which is convenient for the decoding end to reconstruct the image texture.
[0088] Step 104: discretely encode the multi-dimensional feature map to obtain a second code stream.
[0089] In this step, since the local correlation of the multidimensional feature map is relatively weak, while the spatial domain correlation is very strong, the encoding method based on local correlation is not applicable, but the encoding method based on discrete representation is applicable.
[0090] In an optional embodiment, the multi-dimensional feature map may be discretely encoded using a preset discrete encoding model to obtain a second code stream, thereby implementing discrete representation image encoding.
[0091] In one embodiment, the discrete coding model shown in FIG4 includes a feature coding network, a discrete representation layer, and a bitstream conversion layer. The coding process includes downsampling the multidimensional feature map through the feature coding network to obtain downsampled features, then discretely representing the downsampled features through the discrete representation layer to obtain an index matrix, and finally converting the index matrix into a second bitstream through the bitstream conversion layer.
[0092] Among them, the index values in the index matrix come from the indices of the vectors contained in the feature dictionary used in the discrete representation layer. Therefore, the value range of the index values is fixed. A conversion function can be designed in the code stream conversion layer to convert the index values in the index matrix into a binary second code stream.
[0093] As mentioned above, the multidimensional feature map contains high-frequency information and structural information of the image. Therefore, the second code stream obtained by discretely encoding the multidimensional feature map will carry high-frequency information and structural information of the image, which is convenient for the decoding end to reconstruct the image.
[0094] Specifically, for the discrete representation layer, during the training process, a feature dictionary is learned by clustering the multidimensional feature graph. The feature dictionary defines a discrete latent variable space, also called embedding space, whose size is K*D, where K represents the number of embedded vectors and D represents the length of each embedded vector.
[0095] Based on this, the discrete representation process includes: matching the most similar vector to each dimension of the downsampled feature vector in a preset feature dictionary, thereby obtaining an index matrix using the index of the vector matching each dimension of the feature vector. The feature vector can be matched to similar vectors in the feature dictionary by selecting the vector with the closest Euclidean distance.
[0096] Step 105: Send the first code stream and the second code stream to the decoding end.
[0097] At this point, the image processing flow shown in Figure 1 is completed. By extracting the one-dimensional feature vector of the original image block, which is a low-dimensional spatially independent vector that contains the texture structure information of the image, the original image block is then transformed into a multidimensional feature map based on this one-dimensional feature vector. The multidimensional feature map not only contains the spatially related information of the image, but also contains the high-frequency information of the image. Considering that the spatially independent vector and the multidimensional feature map respectively contain information at different levels of the image, the one-dimensional feature vector is quantized and encoded into the first code stream, while the multidimensional feature map is discretely encoded into the second code stream to achieve efficient compression of the spatially independent vector and the multidimensional feature map. Since the encoded code stream consists of two layers of code streams, containing information from different layers of the image, even at low bit rates, not much information is lost. Therefore, the image reconstructed from the two layers of code streams can improve the visual effect and enhance the visual experience.
[0098] Example 2:
[0099] FIG5 is a flow chart of another image processing method according to an exemplary embodiment. The embodiment shown in FIG1 above provides an image processing solution for the encoding end. On this basis, the image processing method provided in this embodiment is applied to the decoding end. As shown in FIG5 , the image processing method includes the following steps:
[0100] Step 501: Receive a first code stream and a second code stream from an encoding end.
[0101] Step 502: Perform inverse quantization decoding on the first code stream to obtain a one-dimensional feature vector.
[0102] In this step, as mentioned above, the first code stream is obtained by encoding using a quantization coding model. Accordingly, the decoding process of the first code stream can be to perform inverse quantization decoding on the first code stream through a preset inverse quantization decoding model to obtain a one-dimensional feature vector, thereby realizing the decoding of the first code stream.
[0103] In one feasible implementation, the inverse quantization decoding model shown in Figure 6 includes an entropy decoding network, an inverse quantization layer, and a feature decoding network. The decoding process includes: decoding the first code stream into quantized features through the entropy decoding network, then converting the quantized features into floating-point values through the inverse quantization layer to obtain inverse quantized features, and finally upsampling the inverse quantized features through the feature decoding network to obtain a one-dimensional feature vector.
[0104] Step 503: Perform inverse discrete decoding on the second code stream to obtain a multi-dimensional feature map.
[0105] In this step, as mentioned above, the second code stream is obtained by encoding using a discrete coding model. Accordingly, the decoding process of the second code stream can be performed by performing inverse discrete decoding on the second code stream through a preset discrete decoding model to obtain a multi-dimensional feature map; thereby realizing the decoding of the second code stream.
[0106] In one feasible implementation, the discrete decoding model shown in FIG7 includes a code stream decoding layer, an inverse discrete representation layer, and a feature decoding network. The decoding process includes: decoding the second code stream into an index matrix through the code stream decoding layer, then converting the index matrix into continuous features based on a preset feature dictionary through the inverse discrete representation layer, and finally upsampling the continuous features through the feature decoding network to obtain a multidimensional feature map.
[0107] The feature dictionary used by the inverse discrete representation layer is the same as that used by the discrete representation layer to facilitate decoding the index matrix. Based on the discrete representation process described above, the process of converting the index matrix into continuous features can be to represent each index in the index matrix using the embedding vector corresponding to the index in the feature dictionary, thereby obtaining a continuous feature.
[0108] Step 504: reconstruct the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block.
[0109] As mentioned above, since the multidimensional feature map is obtained by the forward diffusion process of the diffusion model, the realization of image generation from the multidimensional feature map is obtained by the backward process of the diffusion model, that is, denoising decoding is performed by the backward process.
[0110] Based on this, the multidimensional feature map is iteratively denoised multiple times using a one-dimensional feature vector through a preset diffusion model to obtain a reconstructed image block.
[0111] In one feasible embodiment, as shown in FIG2 , the backward process of the diffusion model includes a second noise prediction network and a denoising layer. During multiple iterative denoising processes of the multidimensional feature map, for each denoising process, the second noise prediction network predicts the current noise based on the previous denoising result and the one-dimensional feature vector. Then, the denoising layer denoises the previous denoising result using the noise output by the second noise prediction network.
[0112] For the first denoising process, the previous denoising result is a multidimensional feature map. Specifically, the specific denoising implementation process for the denoising layer includes: obtaining the noise level coefficient used last and the current noise level coefficient, and then obtaining the current denoising result based on the noise level coefficient used last, the current noise level coefficient, the noise output by the second noise prediction network, and the previous denoising result.
[0113] Based on the above-mentioned denoising formula, the corresponding denoising formula can be selected as:
[0114] Among them, x t Indicates the last denoising result, x t-1 Represents the current denoising result, ε θ (x t ,t,z t ) represents the noise output by the second noise prediction network, z t represents the one-dimensional feature vector, α t Indicates the last noise level coefficient, α t-1 Indicates the current noise level coefficient.
[0115] At this point, the image processing flow shown in Figure 5 is complete. Since the received first and second bitstreams contain information at different levels of the image, the first bitstream contains texture and style information of the image, and the second bitstream contains high-frequency information and structural information of the image, the reconstructed image obtained by decoding the first and second bitstreams can improve the image visual effect and does not lose much information even under bandwidth constraints.
[0116] The execution subject of the embodiment of the present application can be an application, service, instance, functional module in software form, virtual machine (VM), container or cloud server, etc., or a hardware device with data processing function (such as a server or terminal device) or hardware chip (such as CPU, GPU, FPGA, NPU, AI acceleration card or DPU), etc. The device for implementing image processing can be deployed on the computing device of the application party that provides the corresponding service or on a cloud computing platform that provides computing power, storage and network resources. The mode in which the cloud computing platform provides services to the outside world can be IaaS (Infrastructure as a Service), PaaS (Platform as a Service), SaaS (Software as a Service) or DaaS (Data as a Service). Taking the platform providing SaaS software as a service (Software as a Service) as an example, the cloud computing platform can use its own computing resources to provide training of image processing models or functional execution of image processing modules, and the specific application architecture can be built according to service requirements. For example, the platform can provide construction services based on the above model to applications or individuals using platform resources, and further call the above model and implement online or offline image processing functions based on image processing requests submitted by relevant clients or servers and other devices.
[0117] Based on the embodiments shown in Figures 1 and 5 above, the present application also proposes an image processing system, which includes an encoding end and a decoding end, wherein the encoding end is used to execute the image processing method of the embodiment shown in Figure 1 above, and the decoding end is used to execute the image processing method of the embodiment shown in Figure 5 above.
[0118] The encoding end and the decoding end are combined below to explain the image processing method proposed in this application in detail.
[0119] At the encoding end, first, as shown in Figure 8, the convolutional encoder is used to extract the spatial domain-independent feature vector Z of the original image block x0. t (i.e. one-dimensional eigenvector), and based on the diffusion process of the diffusion model, the spatial domain independent eigenvector Z is used tAdd noise to the original image block x0 one by one and output the noise map X T (ie, a multidimensional feature map); then, as shown in FIG9 , the spatial domain independent feature vector Z is quantized by the coding model. t Perform feature coding, quantization and entropy estimation in sequence to obtain the first bitstream; as shown in FIG10 , the noise image X is coded by the discrete coding model. T Perform feature encoding and discrete representation in sequence to obtain an index matrix, and convert the index matrix into a second code stream;
[0120] At the decoding end, first, as shown in FIG9 , the first bit stream is dequantized and feature decoded by the dequantization decoding model to obtain the decoded spatial domain independent feature vector As shown in Figure 10, the second code stream is sequentially de-discretely represented and feature decoded through the discrete decoding model to obtain the reconstructed noise map. Then, as shown in Figure 8, based on the backward process of the diffusion model, the reconstructed spatial domain independent feature vector The reconstructed noise map Perform successive denoising and output the reconstructed image blocks.
[0121] Based on the above description, as shown in Figure 11, using the same input image, the original image is encoded and reconstructed by the HiFiC scheme to obtain Figure a, the original image is encoded and reconstructed by the JPEG scheme to obtain Figure b, and the original image is encoded and reconstructed using the scheme of the present application to obtain Figure c. By comparing the screenshots of the same areas in the original image, Figure a, Figure b, and Figure c with the screenshots of the original image, it can be found that the reconstructed effects obtained by using the HiFiC scheme and the JPEG scheme both have artifacts, while the reconstructed effect obtained by using the scheme of the present application has no artifacts and is almost the same as the original image effect.
[0122] Corresponding to the embodiment of the aforementioned image processing method, the present application further provides an embodiment of an image processing device, which is used to execute the image processing method provided in the embodiment shown in FIG. 1 and is applied to an encoding end. The image processing device includes:
[0123] A first extraction module is used to extract a one-dimensional feature vector of the original image block;
[0124] A second extraction module, configured to transform the original image block into a multidimensional feature map based on the one-dimensional feature vector;
[0125] A first encoding module, configured to perform quantization encoding on the one-dimensional feature vector to obtain a first code stream;
[0126] A second encoding module, configured to perform discrete encoding on the multidimensional feature map to obtain a second code stream;
[0127] The sending module is used to send the first code stream and the second code stream to the decoding end.
[0128] The converting the original image block into a multi-dimensional feature map based on the one-dimensional feature vector includes:
[0129] In an optional implementation, the second extraction module is specifically configured to perform multiple iterations of noise addition on the original image block using the one-dimensional feature vector through a preset diffusion model to obtain a multi-dimensional feature map.
[0130] In an optional implementation, the diffusion model includes a first noise prediction network and a noise addition layer; the second extraction module is specifically used to predict the current noise according to the previous noise addition result and the one-dimensional feature vector through the first noise prediction network for each noise addition process; and add the noise output by the first noise prediction network to the previous noise addition result through the noise addition layer; wherein, for the first noise addition process, the previous noise addition result is the original image block.
[0131] In an optional implementation, the second extraction module is specifically used to obtain the noise level coefficient used last and the noise level coefficient used currently in the process of adding the noise output by the first noise prediction network to the noise addition result last time through the noise addition layer; and obtain the noise addition result currently based on the noise level coefficient used last, the noise level coefficient currently, the noise output by the first noise prediction network, and the noise addition result last time.
[0132] In an optional implementation, the first encoding module is specifically configured to perform quantization encoding on the one-dimensional feature vector using a preset quantization encoding model to obtain a first code stream.
[0133] In an optional implementation, the quantization coding model includes a feature coding network, a quantization layer, and an entropy estimation network; the first coding module is specifically used to downsample the one-dimensional feature vector through the feature coding network to obtain downsampled features; convert the feature values of the downsampled features into integer values through the quantization layer to obtain quantized features; and encode the quantized features into a first code stream through the entropy estimation network.
[0134] In an optional implementation, the second encoding module is specifically configured to perform discrete encoding on the multidimensional feature map using a preset discrete encoding model to obtain a second code stream.
[0135] In an optional implementation, the discrete coding model includes a feature coding network, a discrete representation layer, and a code stream conversion layer; the second coding module is specifically used to downsample the multidimensional feature map through the feature coding network to obtain downsampled features; discretely represent the downsampled features through the discrete representation layer to obtain an index matrix; and encode the index matrix into a second code stream through the code stream conversion layer.
[0136] In an optional implementation, the second encoding module is specifically used to match the most similar vector for each dimension of the feature vector of the downsampled feature in a preset feature dictionary during the process of discretely representing the downsampled feature through the discrete representation layer to obtain an index matrix; and obtain the index matrix using the index of the vector matching the feature vector of each dimension.
[0137] Corresponding to the embodiment of the aforementioned image processing method, the present application also provides another embodiment of an image processing device, which is used to execute the image processing method provided in the embodiment shown in FIG. 5 and is applied to a decoding end. The image processing device includes:
[0138] A receiving module, receiving a first code stream and a second code stream from an encoding end;
[0139] A first decoding module, configured to perform inverse quantization decoding on the first bit stream to obtain a one-dimensional feature vector;
[0140] A second decoding module, configured to perform inverse discrete decoding on the second code stream to obtain a multidimensional feature map;
[0141] A reconstruction module is used to perform image reconstruction on the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block.
[0142] In an optional implementation, the first decoding module is specifically configured to perform inverse quantization decoding on the first bitstream using a preset inverse quantization decoding model to obtain a one-dimensional feature vector;
[0143] The inverse quantization decoding model includes an entropy decoding network, an inverse quantization layer and a feature decoding network; the first decoding module is specifically used to decode the first code stream into quantized features through the entropy decoding network; convert the quantized features into floating-point values through the inverse quantization layer to obtain inverse quantized features; and upsample the inverse quantized features through the feature decoding network to obtain a one-dimensional feature vector.
[0144] In an optional implementation, the second decoding module is specifically configured to perform inverse discrete decoding on the second bitstream using a preset discrete decoding model to obtain a multidimensional feature map;
[0145] The discrete decoding model includes a code stream decoding layer, an inverse discrete representation layer and a feature decoding network; the second decoding module is specifically used to decode the second code stream into an index matrix through the code stream decoding layer; convert the index matrix into continuous features based on a preset feature dictionary through the inverse discrete representation layer; and upsample the continuous features through the feature decoding network to obtain a multidimensional feature map.
[0146] In an optional implementation, the reconstruction module is specifically configured to perform multiple iterative denoising operations on the multidimensional feature map using the one-dimensional feature vector through a preset diffusion model to obtain a reconstructed image block.
[0147] In an optional implementation, the diffusion model includes a second noise prediction network and a denoising layer; the reconstruction module is specifically used to predict the current noise according to the previous denoising result and the one-dimensional feature vector through the second noise prediction network for each denoising process; and denoise the previous denoising result using the noise output by the second noise prediction network through the denoising layer; wherein, for the first denoising process, the previous denoising result is the multidimensional feature map.
[0148] In an optional implementation, the reconstruction module is specifically used to obtain the noise level coefficient used last and the noise level coefficient used currently in the process of denoising the last denoising result using the noise output by the second noise prediction network through the denoising layer; and obtain the current denoising result based on the noise level coefficient used last, the current noise level coefficient, the noise output by the second noise prediction network, and the last denoising result.
[0149] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0150] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0151] The present application also provides an electronic device corresponding to the image processing method provided in the aforementioned embodiment, for executing the aforementioned image processing method. The electronic device includes: a communication interface, a processor, a memory, and a bus; wherein the communication interface, the processor, and the memory communicate with each other via the bus. The processor can execute the image processing method described above by reading and executing machine-executable instructions corresponding to the control logic of the image processing method in the memory. The specific content of this method is described in the aforementioned embodiment and will not be repeated here.
[0152] The memory mentioned in this application can be any electronic, magnetic, optical, or other physical storage device, and can contain stored information, such as executable instructions, data, etc. Specifically, the memory can be RAM (Random Access Memory), flash memory, a storage drive (such as a hard disk drive), any type of storage disk (such as an optical disk, DVD, etc.), or similar storage media, or a combination thereof. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface (which can be wired or wireless), which can use the Internet, wide area network, local area network, metropolitan area network, etc.
[0153] The bus may be an ISA bus, a PCI bus or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0154] The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or an instruction in the form of software. The above-mentioned processor can be a general-purpose processor, including a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be combined to execute.
[0155] The electronic device provided in the embodiment of the present application and the image processing method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.
[0156] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0157] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0158] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: Applied to the encoding end, the method includes: Extract the one-dimensional feature vector of the original image block; Transforming the original image block into a multi-dimensional feature map based on the one-dimensional feature vector; Performing quantization encoding on the one-dimensional feature vector to obtain a first code stream; Discretely encode the multidimensional feature map to obtain a second code stream; The first code stream and the second code stream are sent to a decoding end.
2. The method according to claim 1, characterized in that: The step of transforming the original image block into a multi-dimensional feature map based on the one-dimensional feature vector comprises: The original image block is iteratively denoised multiple times using the one-dimensional feature vector through a preset diffusion model to obtain a multi-dimensional feature map.
3. The method according to claim 2, characterized in that The diffusion model includes a first noise prediction network and a noise adding layer; the one-dimensional feature vector is used to perform multiple iterations of noise adding on the original image block through the preset diffusion model, including: For each noise adding process, the first noise prediction network predicts the current noise according to the previous noise adding result and the one-dimensional feature vector; Adding the noise output by the first noise prediction network to the last noise addition result through the noise addition layer; Wherein, for the first denoising process, the last denoising result is the original image block.
4. The method according to claim 3, characterized in that Adding the noise output by the first noise prediction network to the last noise addition result through the noise addition layer includes: Get the noise level coefficient used last time and the current noise level coefficient; The current noise addition result is obtained according to the noise level coefficient used last time, the current noise level coefficient, the noise output by the first noise prediction network, and the previous noise addition result.
5. The method according to claim 1, characterized in that: The step of quantizing and encoding the one-dimensional feature vector to obtain a first bit stream includes: The one-dimensional feature vector is quantized and encoded using a preset quantization encoding model to obtain a first code stream.
6. The method according to claim 5, characterized in that The quantization coding model includes a feature coding network, a quantization layer, and an entropy estimation network; The step of quantizing and encoding the one-dimensional feature vector by a preset quantization coding model to obtain a first bit stream includes: Downsampling the one-dimensional feature vector through the feature encoding network to obtain a downsampled feature; The quantization layer converts the feature value of the downsampled feature into an integer value to obtain a quantized feature; The quantized features are encoded into a first code stream through the entropy estimation network.
7. The method according to claim 1, characterized in that The discrete encoding of the multi-dimensional feature map to obtain a second code stream includes: The multi-dimensional feature map is discretely encoded using a preset discrete encoding model to obtain a second code stream.
8. The method according to claim 7, characterized in that The discrete coding model includes a feature coding network, a discrete representation layer, and a code stream conversion layer; The discretely encoding the multidimensional feature map by using a preset discrete encoding model to obtain a second code stream includes: Downsampling the multidimensional feature map through the feature encoding network to obtain downsampled features; Discretely represent the downsampled features through the discrete representation layer to obtain an index matrix; The index matrix is converted into a second code stream through the code stream conversion layer.
9. The method according to claim 8, characterized in that The discrete representation of the downsampled features by the discrete representation layer to obtain an index matrix includes: In a preset feature dictionary, matching the most similar vector for each dimension of the downsampled feature vector; An index matrix is obtained using the index of the vector that matches the feature vector of each dimension.
10. An image processing method, characterized in that: Applied to a decoding end, the method comprises: Receiving a first code stream and a second code stream from an encoding end; Performing inverse quantization decoding on the first code stream to obtain a one-dimensional feature vector; Performing inverse discrete decoding on the second code stream to obtain a multi-dimensional feature map; The multi-dimensional feature map is reconstructed based on the one-dimensional feature vector to obtain a reconstructed image block.
11. The method according to claim 10, characterized in that The step of performing inverse quantization decoding on the first bitstream to obtain a one-dimensional feature vector includes: Dequantizing and decoding the first bitstream using a preset dequantization decoding model to obtain a one-dimensional feature vector; The inverse quantization decoding model includes an entropy decoding network, an inverse quantization layer, and a feature decoding network; the inverse quantization decoding of the first bitstream is performed by a preset inverse quantization decoding model to obtain a one-dimensional feature vector, including: Decoding the first bitstream into quantized features through the entropy decoding network; The quantized features are converted into floating point values by the dequantization layer to obtain dequantized features; The dequantized features are upsampled by the feature decoding network to obtain a one-dimensional feature vector.
12. The method according to claim 10, characterized in that The performing inverse discrete decoding on the second code stream to obtain a multi-dimensional feature map includes: Performing inverse discrete decoding on the second code stream through a preset discrete decoding model to obtain a multi-dimensional feature map; The discrete decoding model includes a code stream decoding layer, an anti-discrete representation layer and a feature decoding network; The step of performing inverse discrete decoding on the second bitstream by using a preset discrete decoding model to obtain a multi-dimensional feature map includes: Decoding the second code stream into an index matrix through the code stream decoding layer; The index matrix is converted into continuous features based on a preset feature dictionary by the anti-discrete representation layer; The continuous features are upsampled through the feature decoding network to obtain a multi-dimensional feature map.
13. The method according to claim 10, characterized in that The step of reconstructing the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block includes: The multi-dimensional feature map is iteratively denoised multiple times using the one-dimensional feature vector through a preset diffusion model to obtain a reconstructed image block.
14. The method according to claim 13, characterized in that The diffusion model includes a second noise prediction network and a denoising layer; the multi-dimensional feature map is iteratively denoised multiple times using the one-dimensional feature vector through the preset diffusion model to obtain a reconstructed image block, including: For each denoising process, predicting the current noise according to the previous denoising result and the one-dimensional feature vector through the second noise prediction network; De-noising the last de-noising result by using the noise output by the second noise prediction network through the de-noising layer; Among them, for the first denoising process, the last denoising result is the multi-dimensional feature map.
15. The method according to claim 14, characterized in that De-noising the last de-noising result by using the de-noising layer and the noise output by the second noise prediction network, including: Get the noise level coefficient used last time and the current noise level coefficient; The current denoising result is obtained according to the noise level coefficient used last time, the current noise level coefficient, the noise output by the second noise prediction network, and the previous denoising result.
16. An image processing device, characterized in that: Applied to the encoding end, the device includes: A first extraction module, used to extract a one-dimensional feature vector of an original image block; A second extraction module, configured to transform the original image block into a multi-dimensional feature map based on the one-dimensional feature vector; A first encoding module, used for performing quantization encoding on the one-dimensional feature vector to obtain a first code stream; A second encoding module, used for discretely encoding the multidimensional feature map to obtain a second code stream; The sending module is used to send the first code stream and the second code stream to a decoding end.
17. An image processing device, characterized in that: Applied to a decoding end, the device comprises: A receiving module, receiving a first code stream and a second code stream from an encoding end; A first decoding module, used for performing inverse quantization decoding on the first bit stream to obtain a one-dimensional feature vector; A second decoding module, used for performing inverse discrete decoding on the second code stream to obtain a multi-dimensional feature map; A reconstruction module is used to perform image reconstruction on the multidimensional feature map based on the one-dimensional feature vector to obtain a reconstructed image block.
18. An image processing system, characterized in that: The system comprises: An encoding end, used to execute the image processing method according to any one of claims 1 to 9; A decoding end, used to execute the image processing method described in any one of claims 10-15.
19. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the program to implement the method according to any one of claims 1 to 15.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Coding and decoding network training method and device, equipment and storage medium
CN112929666A
Natural adversarial patch generation method, and target detection model training method and device
CN116631043A
Image transformation method, electronic equipment and storage medium
CN117170560A
Image processing method, device and system, electronic equipment and storage medium
CN117459727A
System, method, and computer program for recommending items using a direct neural network structure
US20210150337A1