Image target characteristic compression and reconstruction method based on generative diffusion super-division

Through the image target characteristic compression and reconstruction method based on the generative diffusion super-score, combined with the technical means of massive information learning embedding, self-attention enhancement and generative diffusion, the problem of limited image compression and reconstruction performance in the prior art is solved, and efficient image data compression purification and reconstruction tasks are realized, and image quality and detail features are improved.

CN120107065APending Publication Date: 2025-06-06CHINA ACAD OF AEROSPACE SCI & TECH INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064091.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing CNN-based image compression reconstruction model has shortcomings in local detail recovery and global semantic information processing, resulting in limited compression reconstruction performance and quantization operations will lead to blurred image quality.

Method used

The image target characteristic compression and reconstruction method based on generative diffusion super-score is adopted. The embedded image semantic encoding compression and purification model is used to learn massive information to convert the original image into low-dimensional feature encoding, and then a binary code stream is formed after quantization; then the binary code stream is decoded using the self-attention-enhanced image reconstruction model to obtain a preliminary reconstruction image; finally, the image super-resolution model based on generative diffusion is enhanced through the pixel-perceptual cross-attention model to repair the details and texture features of the reconstructed image.

Benefits of technology

It effectively overcomes the shortcomings of the prior art in local detail recovery and global semantic information processing, improves the performance and quality of image compression and reconstruction, and ensures the clarity of image quality and the retention of detailed features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107065A_ABST
    Figure CN120107065A_ABST
Patent Text Reader

Abstract

The invention discloses an image target characteristic compression and reconstruction method based on generative diffusion super-division, which comprises the following steps of: inputting an original image to be compressed into an image semantic coding compression and purification model embedded based on mass information learning for processing to obtain a binary code stream file; inputting the binary code stream file into a self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image; and inputting the preliminary reconstructed image into an image super-resolution model based on generative diffusion for processing to obtain a high-resolution image. According to the method, image data compression purification and reconstruction tasks are efficiently completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image compression and reconstruction, and in particular relates to an image target characteristic compression and reconstruction method based on generative diffusion super-resolution. Background Art

[0002] In recent years, variational autoencoders (VAEs) have achieved better rate-distortion performance than traditional compression methods in terms of peak signal-to-noise ratio and structural similarity metrics, showing great compression potential. For encoding, VAE-based image compression and reconstruction methods use linear and nonlinear parameter analysis transformations to map images to latent space. After quantization, the entropy estimation module predicts the distribution of the latent vector, and then the delay is compressed into the bitstream based on lossless context-adaptive binary arithmetic coding (CABAC). At the same time, hyper-prior, autoregressive prior, and Gaussian mixture models allow the entropy estimation module to more accurately predict the distribution of delays and achieve better rate-distortion (RD) performance. For decoding, the lossless CABAC or decoder decompresses the bitstream and then maps the decompressed latent vector to the reconstructed image through transformation. Combined with the above sequential units, end-to-end training can be performed. At present, the mainstream method of CNN-based learning image compression is also to carry out unit and module optimization work for VAE-based structures, and has made some progress.

[0003] Image compression and reconstruction based on deep learning networks have the following difficulties:

[0004] (a) The core problem of CNN-based models is that the original convolutional layers are designed for high-level global feature extraction rather than low-level local detail recovery. Therefore, CNN models are still affected by weak local detail learning capabilities, limiting the performance improvement of compression and reconstruction tasks.

[0005] (b) The global semantic information in image compression and reconstruction tasks is not as effective as that in other computer vision tasks. On the contrary, spatially adjacent elements have stronger correlations.

[0006] (c) The image compression and reconstruction process requires quantization to reduce the number of bits required for image data in order to achieve the desired locking purpose. However, since the quantization operation will lose some feature information, the generated reconstructed image will have blurred image quality. Summary of the invention

[0007] The technical problem solved by the present invention is: to overcome the deficiencies of the prior art, to provide an image target characteristic compression and reconstruction method based on generative diffusion super-resolution, and to efficiently complete the image data compression, purification and reconstruction tasks.

[0008] The purpose of the present invention is achieved through the following technical solutions: a method for compressing and reconstructing image target characteristics based on generative diffusion super-resolution, comprising: inputting the original image to be compressed into an image semantic coding compression purification model based on massive information learning embedding for processing to obtain a binary code stream file; inputting the binary code stream file into a self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image; inputting the preliminary reconstructed image into an image super-resolution model based on generative diffusion for processing to obtain a high-resolution image.

[0009] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the image semantic coding compression and purification model based on massive information learning embedding includes a first convolutional layer, a first generalized divisive normalization, a second convolutional layer, a second generalized divisive normalization, a third convolutional layer, a third generalized divisive normalization, a fourth convolutional layer, a channel autoregressive entropy model, a quantizer and a bitstream encoder.

[0010] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the original image to be compressed is input into the image semantic coding compression purification model based on massive information learning embedding for processing to obtain a binary code stream file, including: the original image to be compressed is successively passed through the first convolution layer, the first generalized split normalization, the second convolution layer, the second generalized split normalization, the third convolution layer, the third generalized split normalization and the fourth convolution layer to complete feature extraction, and the extracted features enter the channel autoregressive entropy model and the quantizer respectively; the quantizer processes the features to obtain discrete features, and transmits the discrete features to the code stream encoder; the channel autoregressive entropy model processes the features to obtain super prior information, and transmits the super prior information to the code stream encoder; the code stream encoder uses the super prior information to perform code stream encoding on the discrete features to obtain a binary code stream containing discrete features and super prior information.

[0011] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the self-attention enhanced image reconstruction model includes a bitstream decoder, a second channel autoregressive entropy model, a first window attention module, a fifth convolutional layer, a first inverse generalized split normalization, a sixth convolutional layer, a second window attention module, a second inverse generalized split normalization, a seventh convolutional layer, a third inverse generalized split normalization, and an eighth convolutional layer.

[0012] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the binary code stream file is input into the self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image, including: the binary code stream file is input into the code stream decoder and the second channel autoregressive entropy model respectively; the second channel autoregressive entropy model obtains entropy model information according to the binary code stream file, and transmits the entropy model information to the code stream decoder; the code stream decoder decodes the binary code stream file according to the entropy model information to obtain a feature map; the feature map is processed in sequence through the first window attention module, the fifth convolutional layer, the first inverse generalized split normalization, the sixth convolutional layer, the second window attention module, the second inverse generalized split normalization, the seventh convolutional layer, the third inverse generalized split normalization and the eighth convolutional layer to obtain a preliminary reconstructed image.

[0013] In the above-mentioned image target characteristic compression and reconstruction method based on generative diffusion super-resolution, the image super-resolution model based on generative diffusion includes an autoencoder and a diffusion super-resolution model.

[0014] In the above-mentioned image target characteristic compression and reconstruction method based on generative diffusion super-resolution, the preliminary reconstructed image is input into the image super-resolution model based on generative diffusion for processing to obtain a high-resolution image, including: inputting the preliminary reconstructed image into the encoder of the autoencoder to form a second feature map, and inputting the second feature map into the diffusion super-resolution model; the diffusion super-resolution model processes the second feature map through a diffusion process to obtain a high-resolution feature map; and inputting the high-resolution feature map into the decoder of the autoencoder to output a high-resolution image.

[0015] In the above-mentioned image target characteristic compression and reconstruction method based on generative diffusion super-resolution, the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer are the same, all including N convolution kernels, a size of 5 and a convolution step of 2, where N is a positive integer.

[0016] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the channel autoregressive entropy model separates the features into multiple slices, performs entropy coding on the first slice to obtain the first entropy model information; when entropy coding the second slice, it will be combined with the first entropy model information to perform entropy coding to obtain the second entropy model information; and so on, until the entropy model information of the last slice is obtained.

[0017] In the above-mentioned image target feature compression and reconstruction method based on generative diffusion super-resolution, the diffusion super-resolution model includes a pixel-aware cross-attention module and a degradation removal module, wherein the pixel-aware cross-attention module calculates the attention of the features so that the diffusion process can perceive the local structure of the image at the pixel level, and at the same time uses the degradation removal module to extract features that are insensitive to degradation to guide the diffusion process.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] (1) The present invention converts the original image into a low-dimensional feature code by learning an embedded image semantic coding compression purification model based on massive information, and then discretizes the feature code by a quantizer to reduce the size of the data and form a binary code stream;

[0020] (2) The present invention uses a self-attention enhanced image reconstruction model, based on self-attention enhancement, combined with super-prior information analysis to recover the reconstructed image from the decoded binary code stream;

[0021] (3) The present invention uses a generative diffusion-based image super-resolution model and a pixel-aware cross-attention model to perform diffusion super-resolution information enhancement processing, thereby repairing the details and texture features of the reconstructed image and improving the quality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0023] Figure 1 It is a flow chart of an image target characteristic compression and reconstruction method based on generative diffusion super-resolution provided by an embodiment of the present invention;

[0024] Figure 2 It is a schematic diagram of an image semantic coding compression and purification model based on massive information learning embedding provided by an embodiment of the present invention;

[0025] Figure 3 is a schematic diagram of a self-attention enhanced image reconstruction model provided by an embodiment of the present invention;

[0026] Figure 4 is a schematic diagram of a window attention module provided by an embodiment of the present invention;

[0027] Figure 5 Schematic diagram of an image super-resolution model based on generative diffusion provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to be able to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0029] Figure 1 : is a flow chart of an image target characteristic compression and reconstruction method based on generative diffusion super-resolution provided by an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0030] The original image to be compressed is input into the image semantic coding compression purification model based on massive information learning embedding for processing to obtain a binary code stream file;

[0031] Input the binary code stream file into the self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image;

[0032] The preliminary reconstructed image is input into the image super-resolution model based on generative diffusion for processing to obtain a high-resolution image.

[0033] Design an image semantic coding compression purification model based on massive information learning embedding. Based on massive information learning embedding, extract the image binary code stream through the convolution model and the channel autoregressive entropy model to complete the compression process. Specifically, design an image semantic coding compression purification model based on massive information learning embedding. The original image to be compressed is first encoded using the convolutional neural network model, and then quantized using the channel autoregressive entropy model. Feature representation and feature discretization are completed in succession to obtain a compact binary code stream file.

[0034] A self-attention enhanced image reconstruction model is designed. Based on self-attention enhancement, the binary file is decoded to obtain the reconstructed image by combining super-prior information analysis to complete the reconstruction process. Specifically, a self-attention enhanced image reconstruction model is designed. The binary code stream file is analyzed by super-prior information using the channel autoregressive entropy model. The code stream file is decoded by the convolutional network and the window attention module based on the analysis results to obtain a preliminary reconstructed image.

[0035] Specifically, an image super-resolution model based on generative diffusion is designed, and the diffusion super-resolution information is enhanced through the pixel-aware cross-attention model to improve the quality of the reconstructed image. Specifically, an image super-resolution model based on generative diffusion is designed, and the diffusion super-resolution information is enhanced through the pixel-aware cross-attention module to supplement the pixel-level image information and further improve the quality of the reconstructed image.

[0036] like Figure 2 As shown, the image semantic coding compression and purification model based on massive information learning embedding includes a first convolution layer, a first generalized split normalization, a second convolution layer, a second generalized split normalization, a third convolution layer, a third generalized split normalization, a fourth convolution layer, a channel autoregressive entropy model, a quantizer and a bitstream encoder.

[0037] The original image to be compressed is input into the image semantic coding compression purification model based on massive information learning embedding for processing to obtain a binary code stream file, including: the original image to be compressed is sequentially passed through the first convolution layer, the first generalized split normalization, the second convolution layer, the second generalized split normalization, the third convolution layer, the third generalized split normalization and the fourth convolution layer to complete feature extraction, and the extracted features enter the channel autoregressive entropy model and the quantizer respectively; the quantizer processes the features to obtain discrete features, and transmits the discrete features to the code stream encoder; the channel autoregressive entropy model processes the features to obtain super prior information, and transmits the super prior information to the code stream encoder; the code stream encoder uses the super prior information to perform code stream encoding on the discrete features to obtain a binary code stream containing discrete features and super prior information.

[0038] The image semantic coding compression and purification model based on massive information learning embedding has a network structure as follows Figure 2 As shown in the figure. The image is input into the network, and it passes through the convolution layer and generalized split normalization alternately. After feature extraction is completed, the features enter two branches. The first branch discretizes the features, maps the continuous values ​​to a set of discrete values ​​through a quantizer, obtains discrete features, and reduces the number of bits required for the feature map. In the second branch, the features are passed through the channel autoregressive entropy model, combining the spatial grouping hybrid method with the autoregressive modeling along the channel dimension to extract the super-prior information of the features, that is, predict the probability distribution of the features. Finally, the discrete features are encoded using the super-prior information to obtain a binary code stream containing discrete features and super-prior information, completing image compression.

[0039] The image semantic coding compression and purification model based on massive information learning embedding mainly includes convolutional layers, generalized split normalization, and channel self-review entropy model.

[0040] The convolution layer has the same parameter scale, using N convolution kernels with a size of 5 and a convolution step of 2. The downsampling of the feature map is achieved through convolution operations, where N is a positive integer.

[0041] Generalized split normalization is a normalization scheme and nonlinear activation function for image compression algorithms, which aims to Gaussianize the local joint statistics of natural images so that they can efficiently capture the statistical characteristics of image data. The core formula is:

[0042]

[0043] Among them, p i is the i-th channel component of the output feature, x i is the i-th channel component of the input feature, β i To prevent the denominator from being zero in the division operation of the i-th channel component, γ i is the weight coefficient of the i-th channel, and i represents the i-th channel.

[0044] Generalized split normalization performs spatially adaptive normalization on the entire feature map and introduces nonlinearity into the neural network, which helps to eliminate spatial redundancy of the convolutional network and capture spatial structure.

[0045] The overall architecture of the channel autoregressive entropy model is to separate the generated features at the channel level. For example, assuming that the feature map size of the generated features is (batch, width, height, channels), it is manually separated into C_1,...C_n slices, where the size of each slice is batch, width, height, channels / n). First, entropy encoding is performed on the first slice to obtain the entropy model information (μ 1 ,σ 1 ), then, when entropy encoding the second slice, the entropy model information of the first slice (μ 1 ,σ 1 ) is entropy encoded to obtain entropy model information (μ 2 ,σ 2 ). And so on, until the entropy model information of the last slice is obtained (μ n ,σ n ). In this way, better parallel processing and higher coding efficiency can be achieved. n is the information mean of the nth slice entropy model, σ n is the information variance of the nth slice entropy model, and n is the number of slices.

[0046] like Figure 3As shown, the self-attention enhanced image reconstruction model includes a bitstream decoder, a second channel autoregressive entropy model, a first window attention module, a fifth convolutional layer, a first inverse generalized split normalization, a sixth convolutional layer, a second window attention module, a second inverse generalized split normalization, a seventh convolutional layer, a third inverse generalized split normalization, and an eighth convolutional layer.

[0047] Inputting the binary code stream file into the self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image includes: inputting the binary code stream file into the code stream decoder and the second channel autoregressive entropy model respectively; the second channel autoregressive entropy model obtains entropy model information according to the binary code stream file, and transmits the entropy model information to the code stream decoder; the code stream decoder decodes the binary code stream file according to the entropy model information to obtain a feature map; the feature map is processed in sequence through the first window attention module, the fifth convolutional layer, the first inverse generalized split normalization, the sixth convolutional layer, the second window attention module, the second inverse generalized split normalization, the seventh convolutional layer, the third inverse generalized split normalization and the eighth convolutional layer to obtain a preliminary reconstructed image.

[0048] The self-attention enhanced image reconstruction model has a network structure as follows Figure 3 As shown in the figure. The binary code stream file containing discrete features and super prior information is input into the self-attention enhanced image reconstruction model. The super prior information uses the channel autoregressive entropy model to obtain entropy model information, and combines this information to decode the code stream of discrete features to complete the conversion from binary code stream to feature map. The feature map is input into the window attention module. By generating an attention mask based on spatial adjacent elements, the performance of compression distortion rate can be improved at a low computational cost. Subsequently, the feature map continues to pass through a series of network module components such as convolutional layers, inverse generalized split normalization, and window attention modules to output the reconstructed image and complete the image information reconstruction task.

[0049] The window attention module is an important network module component of the network. It solves the problem of increased computational complexity caused by generating attention masks based on the global receptive field in the traditional attention mechanism, and global semantic information is not necessarily practical. This component aims to effectively model and focus on spatially adjacent elements. It designs a window-based attention, divides the feature map into M×M windows in a non-overlapping manner, and then calculates the attention map of each window separately. After that, the feature information is enhanced through the window attention module, such as Figure 4 shown.

[0050] The window attention module is divided into three branches. The first branch passes through three residual modules after window differentiation. Each residual module consists of 1×1 and 3×3 convolutional layers. Then, the 1×1 convolutional layer is used to adjust the number of channels and calculate the Sigmoid function value as the weighting coefficient. In the second branch, after the feature map passes through three residual modules, it is multiplied by the weighting coefficient of the first branch to obtain a weighted feature map. The third branch is a skip connection (shortcut connection), which adds the weighted feature map and the original feature map to obtain the output. This allows the window attention module to learn the residual function, which helps improve the gradient flow and makes the model learn more effectively.

[0051] Inverse generalized divisive normalization is the reverse of generalized divisive normalization, which is used to recover the original image from the normalized representation.

[0052] like Figure 5 As shown, the image super-resolution model based on generative diffusion includes an autoencoder and a diffusion super-resolution model.

[0053] Inputting the preliminary reconstructed image into the image super-resolution model based on generative diffusion for processing to obtain a high-resolution image includes: inputting the preliminary reconstructed image into the encoder of the autoencoder to form a second feature map, and inputting the second feature map into the diffusion super-resolution model; the diffusion super-resolution model processes the second feature map through a diffusion process to obtain a high-resolution feature map; and inputting the high-resolution feature map into the decoder of the autoencoder to output a high-resolution image.

[0054] like Figure 5 As shown in Figure 1, the image super-resolution model based on generative diffusion consists of two main parts, the autoencoder and the diffusion super-resolution model. The low-resolution image is first input into the encoder of the autoencoder to form a feature map, and then enters the diffusion process to complete the process of generating a blurred image into a clear image, and finally outputs a high-resolution image through the decoder.

[0055] The diffusion process is an important part of the image super-resolution model, which maps the source image x to the target image y∈R through a random iterative refinement process. d , learn the parameter approximation of the conditional transfer distribution p(y|x). This method will generate the target image y_0 from the pure noise image y_T in T refinement steps. Starting from y_T~N(0,1), through continuous iterations (y_T-1, y_T-2,...,y_0), according to the learned conditional transfer distribution p θ (y_T-1|y_T,x), so that y_0~p(y|x). Among them, R dis a d-dimensional vector space, d represents d dimensions, p(y|x) is the conditional transfer distribution of the target image y under the condition of the source image x, T represents a total of T refinement steps, y_T is the noise image of the T-th refinement step, y_T-1 is the noise image of the T-1th refinement step, y_T-2 is the noise image of the T-2th refinement step, p θ (y_T-1|y_T,x) is the conditional transfer distribution of the noise image y_T and the noise image y_T-1 under the condition of the source image x.

[0056] In each forward diffusion, the low-resolution image is included as a priori condition. Specifically, a low-resolution image is added to each x_t, and fused through concatenation (concat) to give the diffusion model a clear generation direction, so that the final generated y_0 is a clear image. The size of the concat low-resolution image is consistent with the final high-resolution image, usually by linearly interpolating the low-resolution image to the same resolution as the target.

[0057] A pixel-aware cross-attention module is introduced into the diffusion process. There is no need to add additional shortcut connections to maintain the original pixel-level image structure. Instead, the diffusion process can perceive the local structure of the image at the pixel level by calculating the attention of the features. At the same time, the degradation removal module is used to extract features that are insensitive to degradation to guide the diffusion process.

[0058] The first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, the fifth convolution layer, the sixth convolution layer, the seventh convolution layer and the eighth convolution layer are the same, all including N convolution kernels, a size of 5 and a convolution step of 2, where N is a positive integer.

[0059] The channel autoregressive entropy model separates the features into multiple slices, performs entropy coding on the first slice to obtain the first entropy model information; when entropy coding the second slice, it combines the first entropy model information to perform entropy coding to obtain the second entropy model information; and so on, until the entropy model information of the last slice is obtained.

[0060] The diffusion super-resolution model includes a pixel-aware cross-attention module and a degradation removal module. The pixel-aware cross-attention module enables the diffusion process to perceive the local structure of the image at the pixel level by calculating the attention of the features. At the same time, the degradation removal module is used to extract features that are insensitive to degradation to guide the diffusion process.

[0061] The image semantic coding compression and purification model based on massive information learning embedding in the first step is trained together with the self-attention enhanced image reconstruction model in the second step. The learning and training method is as follows:

[0062] 1) Training target construction and loss function

[0063] The convolutional layer E in the image semantic coding compression and purification model maps a given image X to a latent vector Y. After the quantization operation Q, is the discrete representation of the latent vector Y. Then, the window attention module and convolutional layer D in the self-attention enhanced image reconstruction model are used to transform Mapping back to the reconstructed image The main process is as follows:

[0064] Y=E(X;φ)

[0065]

[0066] Where Y is the potential vector in the neural network, E(X; φ) represents the convolution operation of a given image X with a convolution layer E with a trainable parameter φ, X is a given image, φ is a trainable parameter of the convolution layer E, is the discrete representation of the potential vector Y, Q(Y) represents the quantization operation of the potential vector Y, To reconstruct the image, Denotes the discrete representation of the convolution layer D with parameter weight θ Perform a convolution operation, θ is the parameter weight of the convolution layer D.

[0067] Introducing the i-th channel super-prior information element Each element Modeled as a single Gaussian distribution with standard deviation Ψ i and mean Ω i . Distribution Modeled by the channel autoregressive entropy model:

[0068]

[0069] in, Represents the super-prior information element of the i-th channel The i-th channel element of the discrete representation of the latent vector Y under the condition The conditional probability distribution of Expressed in Ω i is the mean value with Ψ i is a Gaussian distribution with standard deviation, is the super-prior information element of the i-th channel, is the ith channel element of the discrete representation of the latent vector Y, Ψ i is the standard deviation corresponding to the i-th channel, Ω i is the mean corresponding to the i-th channel, and i represents the i-th channel.

[0070] The common loss function of the image semantic coding compression purification model and the self-attention enhanced image reconstruction model is:

[0071]

[0072] Among them, λ is a parameter that controls the trade-off between compression rate R and distortion rate K. To obey p X The mean of the distribution X, To represent super prior information Discrete representation of the latent vector Y under the condition The conditional probability distribution of is the conditional probability distribution of the super prior information, is the super prior information, represents the Euclidean distance between X and .

[0073] 2) Training experiment settings

[0074] The training set can be selected from large-scale datasets such as OpenImages and ImageNet. The Adam optimizer is used to train the image semantic coding compression purification model and the self-attention enhanced image reconstruction model. The initial learning rate is set to 1.0×10 -4 , and decays with the training process.

[0075] The learning and training method of the image super-resolution model based on generative diffusion in the third step is as follows:

[0076] 1) Training target construction and loss function

[0077] The diffusion super-resolution model receives two inputs: the source image G and the low-resolution target image in is defined as:

[0078]

[0079] ∈ is a noise vector sampled from a standard normal distribution, γ is the weight coefficient, y 0 is a high-resolution image, and N(0,1) is a standard Gaussian distribution. The goal of the model is to try to A high-resolution image V is generated from 0 .

[0080] Input and training objectives of the diffusion super-resolution model:

[0081] In addition to the source image G and the low-resolution image In addition, the model also accepts an additional input, which represents the variance of the noise. This allows the model to know the level of noise. The training goal of the model is to predict the noise vector ∈. For this goal, the following objective function is proposed

[0082]

[0083] in, Represents the source image G and the low-resolution image The joint expectation, E (∈,γ) represents the joint expectation of the noise vector ∈ and the noise variance γ, f θ represents a network model with θ as parameter, p can be 1 or 2, indicating the use of L1 or L2 norm. At the same time, γ is sampled from a distribution p(γ).

[0084] This embodiment converts the original image to be compressed into a more compact low-dimensional feature code, retains the main information of the image, and reduces the storage space requirement; at the same time, it can recover the reconstructed image from the binary code stream, and then through the generative diffusion process, repair the details and texture features of the reconstructed image, thereby enhancing the information quality of the reconstructed image.

[0085] This embodiment converts the original image into a low-dimensional feature code through an image semantic coding compression purification model embedded based on massive information learning, and then discretizes the feature code by a quantizer to reduce the size of the data and form a binary code stream; this embodiment uses a self-attention enhanced image reconstruction model, based on self-attention enhancement, combined with super-prior information analysis to recover the reconstructed image from the binary code stream; this embodiment uses an image super-resolution model based on generative diffusion, and performs diffusion super-resolution information enhancement processing through a pixel-aware cross-attention model, so as to repair the details and texture features of the reconstructed image and improve the quality of the reconstructed image.

[0086] Although the present invention has been disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for image target characteristic compression and reconstruction based on generative diffusion super-resolution, characterized in that include: The original image to be compressed is input into the image semantic coding compression purification model based on massive information learning embedding for processing to obtain a binary code stream file; Input the binary code stream file into the self-attention enhanced image reconstruction model for processing to obtain a preliminary reconstructed image; The preliminary reconstructed image is input into the image super-resolution model based on generative diffusion for processing to obtain a high-resolution image.

2. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 1 is characterized in that: The image semantic coding compression and purification model based on massive information learning embedding includes a first convolution layer, a first generalized split normalization, a second convolution layer, a second generalized split normalization, a third convolution layer, a third generalized split normalization, a fourth convolution layer, a channel autoregressive entropy model, a quantizer and a bitstream encoder.

3. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 2 is characterized in that: The original image to be compressed is input into the image semantic coding compression purification model based on massive information learning embedding to obtain a binary code stream file including: The original image to be compressed is sequentially passed through the first convolution layer, the first generalized split normalization, the second convolution layer, the second generalized split normalization, the third convolution layer, the third generalized split normalization and the fourth convolution layer to complete feature extraction, and the extracted features enter the channel autoregressive entropy model and the quantizer respectively; The quantizer processes the features to obtain discrete features, and transmits the discrete features to the bitstream encoder; The channel autoregressive entropy model processes the features to obtain super-prior information, and transmits the super-prior information to the bitstream encoder; The code stream encoder uses the super prior information to encode the discrete features and obtain a binary code stream containing the discrete features and the super prior information.

4. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 1 is characterized in that: The self-attention enhanced image reconstruction model includes a bitstream decoder, a second channel autoregressive entropy model, a first window attention module, a fifth convolutional layer, a first inverse generalized split normalization, a sixth convolutional layer, a second window attention module, a second inverse generalized split normalization, a seventh convolutional layer, a third inverse generalized split normalization, and an eighth convolutional layer.

5. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 4 is characterized in that: Input the binary code stream file into the self-attention enhanced image reconstruction model for processing to obtain the preliminary reconstructed image including: Inputting the binary code stream file into the code stream decoder and the second channel autoregressive entropy model respectively; The second channel autoregressive entropy model obtains entropy model information according to the binary code stream file, and transmits the entropy model information to the code stream decoder; The bitstream decoder decodes the binary bitstream file according to the entropy model information to obtain a feature map; The feature map is processed in sequence by the first window attention module, the fifth convolutional layer, the first inverse generalized split normalization, the sixth convolutional layer, the second window attention module, the second inverse generalized split normalization, the seventh convolutional layer, the third inverse generalized split normalization and the eighth convolutional layer to obtain a preliminary reconstructed image.

6. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 1, characterized in that: The image super-resolution model based on generative diffusion includes an autoencoder and a diffusion super-resolution model.

7. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 6 is characterized by: The preliminary reconstructed image is input into the image super-resolution model based on generative diffusion for processing to obtain a high-resolution image including: Input the preliminary reconstructed image into the encoder of the autoencoder to form a second feature map, and input the second feature map into the diffusion super-resolution model; The diffusion super-resolution model processes the second feature map through a diffusion process to obtain a high-resolution feature map; the high-resolution feature map is input into the decoder of the autoencoder to output a high-resolution image.

8. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 2 is characterized by: The first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer are the same, all including N convolution kernels, a size of 5 and a convolution step of 2, where N is a positive integer.

9. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 2 or 3, characterized in that: The channel autoregressive entropy model separates the features into multiple slices, performs entropy coding on the first slice to obtain the first entropy model information; when entropy coding the second slice, it combines the first entropy model information to perform entropy coding to obtain the second entropy model information; and so on, until the entropy model information of the last slice is obtained.

10. The image target characteristic compression and reconstruction method based on generative diffusion super-resolution according to claim 6 or 7, characterized in that: The diffusion super-resolution model includes a pixel-aware cross-attention module and a degradation removal module, wherein the pixel-aware cross-attention module enables the diffusion process to perceive the local structure of the image at the pixel level by calculating the attention of the features, and the degradation removal module is used to extract features that are insensitive to degradation to guide the diffusion process.