Enhancement model training method, enhancement method, apparatus and device, and computer product
By constructing the training data set of the image enhancement model and using multi-scale vector quantization variational autoencoder to convert image features, the problems of slow model inference speed and room for repair effect improvement in the prior art are solved, and more efficient image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510300407.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Among the existing image repair and enhancement technologies, the model's inference speed is slower and there is still room for improvement in the repair effect.
By constructing a training dataset for image enhancement models, a multi-scale vector quantization variational autoencoder is used to convert low-quality images into features of multiple different scales, and iteratively updates through image enhancement models to improve inference speed and enhancement effects.
It has achieved the improvement of the model's inference speed and image enhancement effect, verified the technical potential of the autoregressive paradigm in the field of image super-scoring and repair, and provided a new technological evolution direction for generative visual tasks.
Smart Images

Figure CN120219203A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method for training an image enhancement model, an image enhancement method, a device for training an image enhancement model, an image enhancement device, an electronic device, and a computer program product. Background Art
[0002] Image restoration and enhancement refer to a processing technology for reconstructing a low-quality image (Low-Resolution Image, LR) into a high-quality image (High-Resolution Image, HR). This technology can not only achieve resolution improvement but also cover the restoration of the image quality and enhancement of details of the input low-quality image. In the processes of video data acquisition, compression and transmission, and digitalization of historical images, problems such as resolution loss, detail blurring, and noise interference are common. By using the image restoration and enhancement technology, the user's visual experience can be effectively improved.
[0003] However, the current image restoration and enhancement technologies still have problems such as slow inference speed of the model and there is still room for improvement in the restoration effect.
[0004] In view of this, there is an urgent need in the art for a method that can improve the inference speed and enhancement effect of the model.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a method for training an image enhancement model, an image enhancement method, a device for training an image enhancement model, an image enhancement device, an electronic device, and a computer program product, so as to at least to some extent improve the inference speed and enhancement effect of the model.
[0007] According to a first aspect of the present disclosure, there is provided a method for training an image enhancement model, including:
[0008] Construct a training data set for the image enhancement model, where the training data set includes multiple groups of training data pairs composed of a first-quality sample image and a second-quality sample image, and among them, the image quality of the second-quality sample image is higher than that of the first-quality sample image;
[0009] Input the first-quality sample image and the second-quality sample image into a multi-scale vector quantization variational autoencoder respectively, to obtain multiple first image features of different scales of the first-quality sample image, and multiple second image features of different scales of the second-quality sample image;
[0010] Obtain the semantic features of the first quality sample image, use the semantic features as the starting point of inference, start from the first image features at the smallest scale, and input them into the image enhancement model in sequence. Through the inference of the image enhancement model, obtain the predicted values of the first image features at the next scale;
[0011] Obtain the enhanced sample image corresponding to the first quality sample image according to the predicted value of the first image features at the largest scale, and obtain the first loss according to the enhanced sample image and the second quality sample image;
[0012] Obtain the second loss according to the predicted values of the first image features at each scale and the second image features, obtain the overall loss according to the first loss and the second loss, and iteratively update the model parameters in the image enhancement model based on the overall loss.
[0013] In an exemplary embodiment of the present disclosure, constructing the training dataset of the image enhancement model includes:
[0014] Obtain a set of second quality sample images, and the set of second quality sample images contains multiple second quality sample images;
[0015] Obtain multiple different degradation paths according to various different degradation operations, and perform degradation processing on the second quality sample images according to the degradation paths to obtain the first quality sample images corresponding to the second quality sample images;
[0016] Construct training data pairs based on the second quality sample images and the corresponding first quality sample images, and obtain the training dataset of the image enhancement model according to multiple groups of the training data pairs.
[0017] In an exemplary embodiment of the present disclosure, inputting the first quality sample image into a multi-scale vector quantization variational autoencoder to obtain first image features of multiple different scales of the first quality sample image includes:
[0018] Input the first quality sample image into a multi-scale vector quantization variational autoencoder, and through the multi-scale vector quantization variational autoencoder, decompose the continuous latent space vector corresponding to the first quality sample image into residual latent space vectors of different scales;
[0019] Obtain first image features of multiple different scales of the first quality sample image according to the residual latent space vectors of different scales.
[0020] In an exemplary embodiment of the present disclosure, the obtaining first image features of multiple different scales of the first quality sample image according to the residual latent space vectors of different scales includes:
[0021] Quantize the residual latent space vectors of different scales through a quantizer, and map the residual latent space vectors into a codebook;
[0022] According to the mapping results of the residual latent space vectors of different scales in the codebook, obtain the first image features of multiple different scales of the first quality sample image, wherein the residual latent space vectors of different scales share the codebook.
[0023] In an exemplary embodiment of the present disclosure, obtaining the semantic features of the first quality sample image includes:
[0024] Input the first quality sample image into a semantic feature extractor to obtain the semantic features of the first quality sample image.
[0025] In an exemplary embodiment of the present disclosure, the semantic feature extractor includes a weighted language image pre-training model and a multi-layer perceptron, and the method further includes:
[0026] Iteratively update the model parameters in the multi-layer perceptron of the semantic feature extractor based on the overall loss.
[0027] In an exemplary embodiment of the present disclosure, taking the semantic features as the starting point of inference, starting from the first image features of the smallest scale, inputting them into the image enhancement model in sequence, and inferring the predicted value of the first image features of the next scale through the image enhancement model includes:
[0028] Take the semantic features as the predicted value of the first image features of the smallest scale;
[0029] According to the semantic features, the first image features of the current scale, and the predicted value of the first image features of the previous scale, obtain the input features of the current scale;
[0030] Input the input features of the current scale into the image enhancement model, and infer the predicted value of the first image features of the next scale through the image enhancement model.
[0031] In an exemplary embodiment of the present disclosure, the image enhancement model includes multiple transformation modules, and inputting the input features of the current scale into the image enhancement model, and inferring the predicted value of the first image features of the next scale through the image enhancement model includes:
[0032] Add the first image features of the current scale and the predicted value of the first image features of the previous scale, and input them into a plurality of the transformation modules connected in series;
[0033] After normalizing the semantic features through a linear layer and an adaptive layer, input them into each of the conversion modules respectively, and perform offset and numerical scaling processing on the feature layers in the conversion modules;
[0034] Upsample the output of the last conversion module to obtain the sampling feature corresponding to the current scale, and add the sampling feature corresponding to the current scale to the sampling features corresponding to all scales before the current scale to obtain the total sampling feature;
[0035] Downsample the total sampling feature according to the target size of the current scale to obtain the predicted value of the first image feature of the next scale.
[0036] In an exemplary embodiment of the present disclosure, the upsampling the output of the last conversion module to obtain the sampling feature corresponding to the current scale includes:
[0037] Determine the maximum size of upsampling according to the input size of the first quality sample image and the sampling multiple of the multi-scale vector quantization variational autoencoder in the spatial dimension;
[0038] Upsample the output of the last conversion module according to the maximum size to obtain the sampling feature corresponding to the current scale.
[0039] According to a second aspect of the present disclosure, there is provided an image enhancement method, including:
[0040] Input a first quality image into a multi-scale vector quantization variational autoencoder to obtain first image features of multiple different scales of the first quality image;
[0041] Obtain the semantic features of the first quality image, use the semantic features as the starting point of reasoning, start from the first image feature of the smallest scale, and input them into the image enhancement model in sequence, and obtain the predicted value of the first image feature of the next scale through reasoning of the image enhancement model;
[0042] Obtain the restored and enhanced image corresponding to the first quality image according to the predicted value of the first image feature of the largest scale;
[0043] Wherein, the image enhancement model is obtained by the training method of the image enhancement model as described above.
[0044] According to a third aspect of the present disclosure, there is provided a training device for an image enhancement model, including:
[0045] A training data construction module, configured to construct a training data set for an image enhancement model, where the training data set includes multiple pairs of training data composed of a first-quality sample image and a second-quality sample image, and the image quality of the second-quality sample image is higher than that of the first-quality sample image;
[0046] An image feature determination module, configured to input the first-quality sample image and the second-quality sample image into a multi-scale vector quantization variational autoencoder respectively to obtain multiple first image features of different scales of the first-quality sample image and multiple second image features of different scales of the second-quality sample image;
[0047] An image feature prediction module, configured to obtain the semantic feature of the first-quality sample image, use the semantic feature as the starting point of inference, start from the first image feature of the smallest scale, and input it into the image enhancement model in sequence, and obtain the predicted value of the first image feature of the next scale through the inference of the image enhancement model;
[0048] A sample image enhancement module, configured to obtain the enhanced sample image corresponding to the first-quality sample image according to the predicted value of the first image feature of the largest scale, and obtain the first loss according to the enhanced sample image and the second-quality sample image;
[0049] A model parameter update module, configured to obtain a second loss according to the predicted values of the first image features of each scale and the second image features, obtain an overall loss according to the first loss and the second loss, and iteratively update the model parameters in the image enhancement model based on the overall loss.
[0050] In an exemplary embodiment of the present disclosure, the training data construction module includes:
[0051] A second-quality sample image acquisition unit, configured to acquire a set of second-quality sample images, and the set of second-quality sample images contains multiple second-quality sample images;
[0052] An image degradation processing unit, configured to obtain multiple different degradation paths according to multiple different degradation operations, and perform degradation processing on the second-quality sample image according to the degradation paths to obtain the first-quality sample image corresponding to the second-quality sample image;
[0053] A training data pair construction unit, configured to construct a training data pair based on the second-quality sample image and the corresponding first-quality sample image, and obtain a training data set for the image enhancement model according to multiple groups of the training data pairs.
[0054] In an exemplary embodiment of the present disclosure, the image feature determination module includes:
[0055] A residual latent space vector generation unit, configured to input the first quality sample image into a multi-scale vector quantization variational autoencoder, and decompose the continuous latent space vector corresponding to the first quality sample image into residual latent space vectors of different scales through the multi-scale vector quantization variational autoencoder;
[0056] A first image feature determination unit, configured to obtain multiple first image features of different scales of the first quality sample image according to the residual latent space vectors of different scales.
[0057] In an exemplary embodiment of the present disclosure, the first image feature determination unit includes:
[0058] A vector mapping unit, configured to perform quantization processing on the residual latent space vectors of different scales through a quantizer, and map the residual latent space vectors into a codebook;
[0059] A mapping result determination unit, configured to obtain multiple first image features of different scales of the first quality sample image according to the mapping results of the residual latent space vectors of different scales in the codebook, where the residual latent space vectors of different scales share the codebook.
[0060] In an exemplary embodiment of the present disclosure, the image feature prediction module includes:
[0061] A semantic feature acquisition unit, configured to input the first quality sample image into a semantic feature extractor to obtain the semantic feature of the first quality sample image.
[0062] In an exemplary embodiment of the present disclosure, the model parameter update module includes:
[0063] A multi-layer perceptron parameter update unit, configured to iteratively update the model parameters in the multi-layer perceptron of the semantic feature extractor based on the overall loss.
[0064] In an exemplary embodiment of the present disclosure, the image feature prediction module further includes:
[0065] An initial prediction value determination unit, configured to use the semantic feature as the prediction value of the first image feature of the smallest scale;
[0066] An input feature determination unit, configured to obtain the input feature of the current scale according to the semantic feature, the first image feature of the current scale, and the prediction value of the first image feature of the previous scale;
[0067] The next feature prediction unit is configured to input the input features of the current scale into the image enhancement model, and obtain the predicted value of the first image features of the next scale through inference by the image enhancement model.
[0068] In an exemplary embodiment of the present disclosure, the next feature prediction unit includes:
[0069] The image feature input unit is configured to add the first image features of the current scale and the predicted value of the first image features of the previous scale, and input the sum into a plurality of the conversion modules connected in series;
[0070] The semantic feature input unit is configured to normalize the semantic features through a linear layer and an adaptive layer, and then input them into each of the conversion modules respectively to perform offset and numerical scaling processing on the feature layers in the conversion modules;
[0071] The sampling feature determination unit is configured to upsample the output of the last conversion module to obtain the sampling features corresponding to the current scale, and add the sampling features corresponding to the current scale and the sampling features corresponding to all scales before the current scale to obtain the total sampling features;
[0072] The downsampling unit is configured to downsample the total sampling features according to the target size of the current scale to obtain the predicted value of the first image features of the next scale.
[0073] In an exemplary embodiment of the present disclosure, the sampling feature determination unit includes:
[0074] The maximum size determination unit is configured to determine the maximum size of the upsampling according to the input size of the first quality sample image and the sampling multiple of the multi-scale vector quantization variational autoencoder in the spatial dimension;
[0075] The upsampling unit is configured to upsample the output of the last conversion module according to the maximum size to obtain the sampling features corresponding to the current scale.
[0076] According to the fourth aspect of the present disclosure, there is provided an image enhancement device, including:
[0077] The first feature determination module is configured to input the first quality image into the multi-scale vector quantization variational autoencoder to obtain the first image features of the first quality image at multiple different scales;
[0078] The first feature prediction module is configured to execute obtaining the semantic features of the first quality image, use the semantic features as the starting point of reasoning, start from the first image features at the smallest scale, and sequentially input them into the image enhancement model, and obtain the predicted value of the first image features at the next scale through the reasoning of the image enhancement model;
[0079] The image restoration and enhancement module is configured to execute obtaining the restored and enhanced image corresponding to the first quality image according to the predicted value of the first image features at the largest scale;
[0080] Wherein, the image enhancement model is obtained by the training method of the image enhancement model as described above.
[0081] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the training method of the image enhancement model described in any one of the above.
[0082] According to a sixth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the training method of the image enhancement model described in any one of the above when executed by a processor.
[0083] The exemplary embodiments of the present disclosure may have the following beneficial effects:
[0084] In the training method of the image enhancement model of the exemplary embodiment of the present disclosure, multiple image features at different scales are obtained through a multi-scale vector quantization variational autoencoder, and are sequentially input into the image enhancement model for reasoning, and the restored image is obtained according to the final reasoning result; in the training method of the image enhancement model in the exemplary embodiment of the present disclosure, on the one hand, the visual autoregressive modeling technology is applied to the image restoration task, and by deconstructing the image into features at different resolution scales for feature prediction, efficient image super-resolution and restoration can be achieved, and image generation can be realized through sequence prediction, which can improve the reasoning speed and enhancement effect of the model; on the other hand, the technical potential of the autoregressive paradigm in the field of image super-resolution and restoration is verified, providing a new technical evolution direction for generative visual tasks.
[0085] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0087] Figure 1 A flowchart showing a method for training an image enhancement model according to an exemplary embodiment of the present disclosure;
[0088] Figure 2 A flowchart showing a process of constructing a training dataset for an image enhancement model according to an exemplary embodiment of the present disclosure;
[0089] Figure 3 A flowchart schematically showing an image degradation process according to a specific embodiment of the present disclosure;
[0090] Figure 4 A flowchart schematically showing an inference process of a vector quantization variational autoencoder according to a specific embodiment of the present disclosure;
[0091] Figure 5 A flowchart showing a process of obtaining first image features of multiple different scales according to an exemplary embodiment of the present disclosure;
[0092] Figure 6 A flowchart schematically showing an inference process of a multi-scale vector quantization variational autoencoder according to a specific embodiment of the present disclosure;
[0093] Figure 7 A flowchart schematically showing a process of obtaining first image features and second image features of multiple different scales according to a specific embodiment of the present disclosure;
[0094] Figure 8 A schematic structural diagram showing a semantic feature extractor according to a specific embodiment of the present disclosure;
[0095] Figure 9 A flowchart showing a process of image enhancement model inference according to an exemplary embodiment of the present disclosure;
[0096] Figure 10 A flowchart schematically showing an inference process of an image enhancement model according to a specific embodiment of the present disclosure;
[0097] Figure 11 A flowchart showing a process of obtaining a predicted value of first image features of the next scale through image enhancement model inference according to an exemplary embodiment of the present disclosure;
[0098] Figure 12 Schematically shows a schematic diagram of the structure of an image enhancement autoregressive model according to a specific embodiment of the present disclosure;
[0099] Figure 13 Schematically shows a schematic diagram of the feature sampling process during the inference process of an image enhancement model according to a specific embodiment of the present disclosure;
[0100] Figure 14 Shows a schematic flowchart of an image enhancement method according to an exemplary embodiment of the present disclosure;
[0101] Figure 15 Schematically shows a schematic diagram of a low-quality image encoded into multi-scale image features by a multi-scale vector quantization variational autoencoder according to a specific embodiment of the present disclosure;
[0102] Figure 16 Schematically shows a schematic diagram of an image enhancement repair result according to a specific embodiment of the present disclosure;
[0103] Figure 17 Shows a block diagram of a training device for an image enhancement model according to an exemplary embodiment of the present disclosure;
[0104] Figure 18 Shows a block diagram of an image enhancement device according to an exemplary embodiment of the present disclosure;
[0105] Figure 19 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Specific Embodiments
[0106] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0107] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein.
[0108] The following exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will recognize that one or more of the specific details may be omitted in practicing the technical solutions of this disclosure, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0109] In addition, the accompanying drawings are only schematic illustrations of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0110] In some related embodiments, image super-resolution algorithms based on generative adversarial networks (GANs) and diffusion models have significantly improved the performance ceiling of image restoration techniques. Among them, the image restoration framework based on diffusion models generally adopts a technical solution of coupling a control network with a text-to-image base model, that is, using the text-to-image model as the base, adding a control branch network to its network architecture, and guiding the generation process through low-quality image condition control to achieve high-fidelity image reconstruction. Its technical solutions may include SeeSR (Semantic-Aware Super-Resolution), XPSR (Cross-modal Priors Super-Resolution), CasSR (Cascade Super-Resolution), etc. The above solutions achieve an effective balance between the richness of generated details and visual fidelity through innovative designs such as prompt engineering, reference condition introduction, and image semantic alignment.
[0111] In some other related embodiments, a "diffusion model + control network" architecture can be adopted. Although this method has made remarkable progress, there is still room for improvement as follows: (1) Inference efficiency issue: This solution is based on a multi-step iterative denoising process, usually requiring 20 - 50 iterations of processing, resulting in a relatively long processing time for a single image. (2) Limitations in multimodal adaptability: Large Language Models (LLMs) typically adopt the next token prediction paradigm (such as GPT-3, etc.), while diffusion models generally use Transformer, such as Diffusion Transformer (DiT) or U-Net, such as the iterative inference method of the network base of Stable Diffusion v1.5 (SDv1.5). The differences in their inference architectures limit the application expansion of multimodal large models in image restoration tasks.
[0112] In the field of generative pre-training, the generative pre-training technology of large language models, such as Generative Pretrained Transformers (GPT), has significantly improved the ability of natural language processing and understanding. Its core technology is trained based on the autoregressive paradigm, which has become the basic architecture of large language models. It is particularly worth noting that autoregressive inference constructs the technical path of sequence generation through the mechanism of gradually predicting future tokens, that is, Next token prediction (gradually predicting the next token).
[0113] Furthermore, in the visual field, Visual Auto-Regressive Modeling (VAR) draws on the technical concept of LLMs and makes further expansion and improvement. The paradigm of VAR redefines image autoregressive learning as "next scale prediction" or "next resolution prediction" from coarse to fine. VAR demonstrates performance advantages over GANs and diffusion models in image generation tasks. VAR surpasses diffusion transformers in multiple dimensions such as image quality, inference speed (20-fold improvement), data efficiency, and scalability. However, the VAR technology is only applicable to class image generation. For example, given one of the 1000 classes in ImageNet, the VAR model generates the corresponding image.
[0114] On this basis, the present exemplary embodiment first provides a training method for an image enhancement model. Refer to Figure 1 As shown, the above training method for the image enhancement model may include the following steps:
[0115] Step S110. Construct a training dataset for the image enhancement model. The training dataset includes multiple pairs of training data composed of a first-quality sample image and a second-quality sample image, where the image quality of the second-quality sample image is higher than that of the first-quality sample image.
[0116] Step S120. Input the first-quality sample image and the second-quality sample image into the multi-scale vector quantization variational autoencoder respectively to obtain multiple first image features of different scales of the first-quality sample image and multiple second image features of different scales of the second-quality sample image.
[0117] Step S130. Obtain the semantic features of the first-quality sample image, use the semantic features as the starting point of inference, start from the first image feature of the smallest scale, and input them into the image enhancement model in sequence. Through the inference of the image enhancement model, obtain the predicted value of the first image feature of the next scale.
[0118] Step S140. Obtain the enhanced sample image corresponding to the first-quality sample image according to the predicted value of the first image feature of the largest scale, and obtain the first loss according to the enhanced sample image and the second-quality sample image.
[0119] Step S150. Obtain the second loss according to the predicted values of the first image features of each scale and the second image features, obtain the overall loss according to the first loss and the second loss, and iteratively update the model parameters in the image enhancement model based on the overall loss.
[0120] In the training method of the image enhancement model according to the exemplary embodiment of the present disclosure, multiple image features of different scales are obtained through the multi-scale vector quantization variational autoencoder and input into the image enhancement model for inference in sequence, and the repaired image is obtained according to the final inference result. In the training method of the image enhancement model according to the exemplary embodiment of the present disclosure, on the one hand, the visual autoregressive modeling technology is applied to the image restoration task. By decomposing the image into features of different resolution scales for feature prediction, efficient image super-resolution and restoration can be achieved, and image generation can be realized through sequence prediction, which can improve the inference speed and enhancement effect of the model. On the other hand, the technical potential of the autoregressive paradigm in the field of image super-resolution and restoration is verified, providing a new technical evolution direction for generative visual tasks.
[0121] Next, in combination with Figures 2 to 13 The above steps of this exemplary embodiment will be described in more detail.
[0122] In step S110, a training dataset for the image enhancement model is constructed. The training dataset includes multiple pairs of training data composed of a first-quality sample image and a second-quality sample image, where the image quality of the second-quality sample image is higher than that of the first-quality sample image.
[0123] In this exemplary embodiment, the image enhancement model is a machine learning model for image super-resolution and enhancement repair tasks, such as IE-VAR (Image-Enhanced VAR, Image-Enhanced Visual Autoregressive Model). This model combines image data and an autoregressive model, and enhances the performance of the autoregressive model by introducing image information. It can be used to process sequence generation tasks related to images.
[0124] In this exemplary embodiment, a training dataset for training the image enhancement model can be constructed. The training dataset includes multiple pairs of training data composed of a first-quality sample image and a second-quality sample image. The first-quality sample image refers to a low-quality sample image with low resolution and quality, and the second-quality sample image refers to a high-quality sample image with high resolution and quality. By constructing a training data pair of a high-quality image and a low-quality image, the training dataset of the image enhancement model is obtained.
[0125] In this exemplary embodiment, as Figure 2 shown, the method for constructing the training dataset of the image enhancement model can specifically include the following steps:
[0126] Step S210. Obtain a second-quality sample image set, which contains multiple second-quality sample images.
[0127] During the process of constructing the training dataset, data collection needs to be carried out first. By integrating public datasets and internal data resources, high-quality sample images (High-resolution Images, HR) are obtained to construct a high-quality high-resolution image library. When collecting high-quality images, it can be controlled that the images meet the benchmark requirements of a resolution greater than or equal to 1024×1024 and no compression artifacts.
[0128] Step S220. Obtain multiple different degradation paths according to various different degradation operations, and perform degradation processing on the second-quality sample images according to the degradation paths to obtain the first-quality sample images corresponding to the second-quality sample images.
[0129] After obtaining the high-quality image set, the corresponding low-quality images can be obtained through image quality degradation operations. Specifically, a multi-dimensional degradation operation chain can be implemented on the high-quality images, such as Figure 3As shown, the degradation operations may include random Gaussian / motion blur kernel convolution, upsampling and downsampling, mixed noise injection (such as Gaussian noise and Poisson noise), dynamic range JPEG compression (the quality factor can be randomly distributed in [30 - 95]), etc. Among them, JPEG (Joint Photographic Experts Group) is a widely used image compression standard, mainly used for compressing and storing static images.
[0130] The degradation path can be determined by various specific degradation operations and random parameters. For example, a random parameter between 0 and 1 can be generated for each degradation operation respectively, and a parameter threshold, such as 0.7, can be set. When the random parameter corresponding to a certain degradation operation is greater than or equal to the parameter threshold, it means that this degradation operation can be added to the degradation path, so as to generate various random degradation paths. In addition, the degradation path can also be generated by combining independent random parameters. For example, different sampling resolutions can be set for the upsampling and downsampling operations, and different compression factors can be set for the JPEG compression operation.
[0131] Step S230. Construct a training data pair based on the second quality sample image and the corresponding first quality sample image, and obtain the training dataset of the image enhancement model according to multiple groups of training data pairs.
[0132] Each high-quality image can generate multiple groups of differentiated low-quality images through multiple different degradation paths and multi-level degradation, which can ensure that a single high-quality sample image is expanded into multiple groups of high-quality - low-quality training pairs, achieving the effect of expanding the training data.
[0133] In step S120, the first quality sample image and the second quality sample image are respectively input into the multi-scale vector quantization variational autoencoder to obtain multiple first image features of different scales of the first quality sample image, and multiple second image features of different scales of the second quality sample image.
[0134] The vector quantization variational autoencoder (VQ-VAE, Vector Quantized-Variational Autoencoder) can map the input image to a continuous latent space vector (Latent embedding), and then quantize the output continuous latent representation into a discrete codebook vector.
[0135] Figure 4Schematically shown is the inference flow chart of a vector quantization variational autoencoder according to a specific embodiment of the present disclosure. The vector quantization variational autoencoder includes an encoder for mapping an input image to a continuous latent space vector; a quantizer for quantizing the continuous latent representation output by the encoder into a discrete codebook vector, and the quantization process usually adopts nearest neighbor search to map each continuous vector to the nearest vector in the codebook; a decoder for reconstructing the quantized discrete latent representation into data in the original data space; and a codebook for storing a set of fixed discrete vectors obtained after training, which can be used in the quantization process, and each discrete vector represents a feature vector of the image latent space.
[0136] In this exemplary embodiment, based on the VQ-VAE, a multi-scale vector quantization variational autoencoder (Multi-scale VQ-VAE, Multi-scale Vector Quantized Variational Autoencoder) is proposed, which can convert an image into multi-scale image tokens.
[0137] In this exemplary embodiment, as Figure 5 shown, inputting the first quality sample image into the multi-scale vector quantization variational autoencoder to obtain first image features of the first quality sample image at multiple different scales, which specifically may include the following steps:
[0138] Step S510. Input the first quality sample image into the multi-scale vector quantization variational autoencoder, and decompose the continuous latent space vector corresponding to the first quality sample image into residual latent space vectors at different scales through the multi-scale vector quantization variational autoencoder.
[0139] Figure 6 Schematically shown is the inference flow chart of a multi-scale vector quantization variational autoencoder according to a specific embodiment of the present disclosure. The multi-scale vector quantization variational autoencoder can decompose a continuous latent space vector into residual latent space vectors at different image scales.
[0140] For example, the encoder of the VQ-VAE downsamples by 16 times in the spatial dimension. If the size of the input first-quality sample image is 512x512, then the size of the latent space vector is 32x32. In the original VAR, the 32x32 latent space vector is resampled and decomposed into feature residual maps of different sizes, such as (1x1, 2x2, 4x4, 8x8), etc. In the embodiment of this example, consistent with VAR, the resolutions of different-scale features are (1, 2, 3, 4, 5, 6, 8, 10, 13, 16, 32) respectively. The final latent space vector is obtained by superimposing the residual latent space vectors of all scales.
[0141] Specifically, first, downsample the 32x32 latent space vector to a 1x1 residual latent space vector, then upsample the 1x1 residual latent space vector to 32x32, subtract it from the 32x32 latent space vector, and then downsample again to obtain a 2x2 residual latent space vector, and so on. Finally, all the obtained residual latent space vectors are upsampled to 32x32 and then added together to obtain the original latent space vector.
[0142] Step S520. Obtain multiple first image features of different scales of the first-quality sample image according to the residual latent space vectors of different scales.
[0143] Further discretize the residual latent space vectors of different scales, and map them through a codebook to obtain multi-scale image tokens, that is, the first image features, denoted as r1, r2, r3, …, r k 。
[0144] In the embodiment of this example, the residual latent space vectors of different scales can be quantized through a quantizer, mapped into the codebook, and then according to the mapping results of the residual latent space vectors of different scales in the codebook, multiple first image features of different scales of the first-quality sample image are obtained, where the residual latent space vectors of different scales share the codebook.
[0145] For each input first-quality sample image, a series of tokens can be obtained through multi-scale VQ-VAE. Among them, each scale has a given resolution value. The processing method of the second-quality sample image is similar.
[0146] In the embodiment of this example, the HR image and the LR image can be encoded into latent space vectors f HR and f LR 。 Further, multi-scale VQ-VAE encodes and quantizes both into a series of image tokens, denoted as and Among them, k represents the number of multi - scales. For example, the parameter k can be set to 11. The width and height of each scale are the same, which are (1, 2, 3, 4, 5, 6, 8, 10, 13, 16, 32) respectively, and the number of tokens for each scale is (1, 4, 9, 16, 25, 36, 64, 100, 169, 256, 1024).
[0147] Figure 7 Schematically shows a flowchart of obtaining first - image features and second - image features of multiple different scales according to a specific embodiment of the present disclosure. The high - quality sample image and the low - quality sample image can be encoded into corresponding multi - scale image tokens through the multi - scale VQ - VAE, that is, the sequence tokenization of the image is completed, and the original image is serialized into a series of tokens.
[0148] In step S130, obtain the semantic features of the first - quality sample image. Taking the semantic features as the starting point of reasoning, starting from the first - image features of the smallest scale, input them into the image enhancement model in sequence, and obtain the predicted value of the first - image features of the next scale through the reasoning of the image enhancement model.
[0149] In this exemplary embodiment, the first - quality sample image can be input into the semantic feature extractor to obtain the semantic features of the first - quality sample image. The semantic feature extractor is composed of a SigLIP model (Sigmoid - Weighted Language - Image Pretraining Model) and an MLP (Multilayer Perceptron) module. After the image is input into SigLIP, a 1x768 feature vector is obtained, and after the feature vector passes through the MLP, a feature vector of a given dimension 1xD is obtained. For example, in this exemplary embodiment, the output dimension can be 1x1024.
[0150] Figure 8 Schematically shows a structural diagram of a semantic feature extractor according to a specific embodiment of the present disclosure. The semantic feature extractor includes a weighted language - image pre - training model SigLIP and a multi - layer perceptron MLP. The first - quality sample image can pass through the semantic feature extractor to obtain corresponding semantic features, denoted as [s], for subsequent reasoning processes.
[0151] In this exemplary embodiment, as Figure 9 shown, taking the semantic features as the starting point of reasoning, starting from the first - image features of the smallest scale, input them into the image enhancement model in sequence, and obtain the predicted value of the first - image features of the next scale through the reasoning of the image enhancement model. Specifically, it can include the following steps:
[0152] Step S910. Use the semantic feature as the predicted value of the first image feature at the smallest scale.
[0153] During the inference process of the image enhancement model, first use the semantic feature token[s] of the first quality sample image as the starting point of inference, denoted as r0, as the initial predicted value of the first image feature at the smallest scale.
[0154] Step S920. Obtain the input feature at the current scale according to the semantic feature, the first image feature at the current scale, and the predicted value of the first image feature at the previous scale.
[0155] The input at each scale is the inference result of the previous step, the token of the first quality sample image at the current corresponding scale, and the semantic token of the first quality sample image.
[0156] The first quality sample image extracts a semantic feature vector through the SigLIP model, with a dimension of 1x768, and then obtains a 1x1024 feature vector through an MLP, denoted as [s] (representing start). This vector token and the token of the first scale of the residual latent space vector of the first quality sample image Together serve as the starting input for autoregressive inference, and then start predicting at each resolution.
[0157] Step S930. Input the input feature at the current scale into the image enhancement model, and obtain the predicted value of the first image feature at the next scale through the inference of the image enhancement model.
[0158] Take the image semantic token[s], the token of the smallest scale of the first quality sample image And r0 (r0 = [s]) as three parameters and input them into the image enhancement autoregressive model to infer the image token r1 at the next scale; then take the token of the next scale of the LR image And the previous inference result r1, as well as the image semantic token[s] as inputs, and input them into the image enhancement autoregressive model again to obtain the image token r2 at the next scale; iterate k times in sequence to obtain the final output r k .
[0159] Figure 10 Schematically shows a flowchart of the inference process of the image enhancement model in a specific embodiment according to the present disclosure. The first quality sample image is encoded into a series of multi-scale tokens through the encoder Encoder of the multi-scale VQ-VAE, with a total of k scales, denoted respectively as The first quality sample image passes through a semantic feature extractor to obtain the semantic tokens of the image, denoted as [s]. Then, the image semantic token [s] is used as the starting point of the inference, which can also be denoted as r0. The image semantic token [s], the token of the smallest scale of the first quality sample image and r0 (r0 = [s]) are used as three inputs and input into the image enhancement autoregressive model to infer the image token r1 of the next scale; then the token of the next scale of the first quality sample image and the previous inference result r1, as well as the image semantic token [s] are used as inputs and input into the image enhancement autoregressive model again to obtain the image token r2 of the next scale; iterate k times in sequence to obtain the final output r k . Where the value of k is the number of multi-scales.
[0160] In the embodiment of this example, as Figure 11 shown, the input features of the current scale are input into the image enhancement model, and the predicted value of the first image feature of the next scale is inferred through the image enhancement model. Specifically, it may include the following steps:
[0161] Step S1110. Add the first image feature of the current scale and the predicted value of the first image feature of the previous scale, and input them into a series of multiple transformation modules.
[0162] Figure 12 Schematically shows a structural diagram of an image enhancement autoregressive model according to a specific embodiment of the present disclosure. The main body of the image enhancement autoregressive model consists of multiple standard transformer blocks (transformation modules). When inferring to the nth scale, the input is the inference result r of the previous step n-1 , and the token corresponding to the nth scale of the first quality sample image The number of tokens of both is the same, both are h n-1 ×w n-1 . Before inputting into the Transformer, an add operation is performed on both of them.
[0163] Step S1120. After normalizing the semantic features through a linear layer and an adaptive layer, input them into each transformation module respectively to perform offset and numerical scaling processing on the feature layer in the transformation module.
[0164] Meanwhile, the model input also includes image semantic tokens [s]. The semantic tokens are added to each Transformer module through a Linear layer and Adaptive Layer Normalization (AdaLN), which performs shifting (shift) and scaling (scale) on the feature layer.
[0165] Step S1130. Upsample the output of the last transformation module to obtain the sampled features corresponding to the current scale, and add the sampled features corresponding to the current scale to the sampled features corresponding to all scales before the current scale to obtain the total sampled features.
[0166] In this exemplary embodiment, the maximum size of the upsampling can be determined according to the input size of the first quality sample image and the sampling multiple of the multi-scale vector quantization variational autoencoder in the spatial dimension. Then, the output of the last transformation module is upsampled according to the maximum size to obtain the sampled features corresponding to the current scale.
[0167] The number of tokens output by the model is the same as the input, both being h n-1 ×w n-1 tokens. Upsample the tokens at this scale to the maximum size H×W to obtain the sampled features f n-1 corresponding to the current scale. For example, if the input is 512x512, since the VQ-VAE latent space is downsampled by 16 times, the maximum size is 32x32. Then add f n-1 to the previously inferred f1, f2, f3,... to obtain the total sampled feature f.
[0168] Step S1140. Downsample the total sampled features according to the target size of the current scale to obtain the predicted value of the first image feature at the next scale.
[0169] Finally, downsample the total sampled feature f to the target size at the nth scale, that is, h n ×w n . Assume k = 11, and the width and height of each scale are the same, which are (1, 2, 3, 4, 5, 6, 8, 10, 13, 16, 32) respectively. If it is necessary to infer the size result of n = 4, then the input is 3x3 and the output is 4x4.
[0170] Figure 13 Schematically shows a schematic diagram of the feature sampling process in the inference process of the image enhancement model according to a specific embodiment of the present disclosure. For each autoregressive unit r k , it depends on the tokens at all previous resolutions, including those of the low-resolution input image and the tokens predicted at each level (r1, r2,..., r k-1), starting from an initial token of size 1×1 and autoregressively predicting at successively higher resolutions (r1, r2, …, r k-1 ). The sequence generation process can be represented as the multiplication of K conditional probabilities, denoted as:
[0171]
[0172] where the spatial resolution at each scale is denoted as Here, V represents the VQ-VAE codebook, h and w represent the spatial resolution at each scale respectively, the square of the resolution at each level is the number of tokens at that level, and the codebook V is shared across all scales.
[0173] In step S140, an enhanced sample image corresponding to the first quality sample image is obtained based on the predicted value of the first image feature at the largest scale, and a first loss is obtained based on the enhanced sample image and the second quality sample image.
[0174] The predicted value r of the first image feature at the largest scale k is input into the decoder Decoder of the VQ-VAE, and the final output result, i.e., the restored and enhanced image, can be obtained.
[0175] In this exemplary embodiment, taking the restoration of a 512x512 image as an example, the input is a LR low-quality image, and the output is the restored and enhanced image with the same size as the input image. If the resolution of the LR low-quality image is less than 512x512, it is upsampled to 512x512.
[0176] After obtaining the enhanced sample image, an L1 loss is calculated between the enhanced sample image and the corresponding second quality sample image as the first loss of the model.
[0177] In step S150, a second loss is obtained based on the predicted values of the first image features at each scale and the second image features, the overall loss is obtained based on the first loss and the second loss, and the model parameters in the image enhancement model are iteratively updated based on the overall loss.
[0178] In this exemplary embodiment, for the results (r1, r2, …, r k ) obtained by inference prediction and the image of the high-quality image, a cross-entropy loss is calculated to obtain the second loss of the model, then the first loss and the second loss are added to obtain the overall loss, and the model parameters in the image enhancement model are iteratively updated based on the overall loss.
[0179] In this exemplary embodiment, in addition to the model parameters in the image enhancement model, the model parameters in the multi-layer perceptron of the semantic feature extractor can also be iteratively updated based on the overall loss.
[0180] In addition, this exemplary embodiment also provides an image enhancement method. Refer to Figure 14 As shown, the above image enhancement method may include the following steps:
[0181] Step S1410. Input the first-quality image into the multi-scale vector quantization variational autoencoder to obtain first image features of the first-quality image at multiple different scales.
[0182] Input the first-quality image into the multi-scale vector quantization variational autoencoder. The multi-scale vector quantization variational autoencoder decomposes the continuous latent space vector corresponding to the first-quality image into residual latent space vectors at different scales, and then obtains first image features of the first-quality image at multiple different scales according to the residual latent space vectors at different scales.
[0183] Figure 15 Schematically shows a schematic diagram of encoding a low-quality image into multi-scale image features in a specific embodiment according to the present disclosure. The low-quality image can be encoded into corresponding multi-scale image tokens through the multi-scale VQ-VAE, thereby serializing the original image into a series of tokens at different scales.
[0184] Step S1420. Obtain the semantic feature of the first-quality image, use the semantic feature as the starting point of reasoning, start from the first image feature at the smallest scale, and sequentially input it into the image enhancement model, and obtain the predicted value of the first image feature at the next scale through the inference of the image enhancement model.
[0185] Among them, the image enhancement model can be obtained through the above training method of the image enhancement model.
[0186] By inputting the first-quality image into the semantic feature extractor, the corresponding semantic feature is obtained, and then the semantic feature is used as the starting point of reasoning. Starting from the first image feature at the smallest scale, it is sequentially input into the image enhancement model. The input at each scale is the predicted value of the first image feature obtained by the previous step of reasoning, as well as the first image feature and semantic feature at the current scale. The predicted value of the first image feature at the next scale is obtained through the inference of the image enhancement model. The inference process of the image enhancement model is the same as the inference process in the training process and will not be elaborated here.
[0187] Step S1430. Obtain the restored and enhanced image corresponding to the first-quality image according to the predicted value of the first image feature at the largest scale.
[0188] Finally, by passing the predicted value of the first image feature at the largest scale obtained through inference through the decoder, the restored and enhanced image corresponding to the first-quality image can be obtained.
[0189] Figure 16 Schematically shows a schematic diagram of the image enhancement and restoration result in a specific embodiment according to the present disclosure. Among them, Restoration Result 1 and Restoration Result 2 are obtained based on different inference sampling coefficients. It can be seen that the image quality after restoration has been improved, the clarity has been enhanced, and at the same time, a high fidelity and consistency with the original image have been maintained.
[0190] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0191] Furthermore, the present disclosure also provides a training device for an image enhancement model. Referring to Figure 17 As shown, the training device for the image enhancement model may include a training data construction module 1710, an image feature determination module 1720, an image feature prediction module 1730, a sample image enhancement module 1740, and a model parameter update module 1750. Among them:
[0192] The training data construction module 1710 is configured to construct a training dataset for the image enhancement model. The training dataset includes multiple pairs of training data composed of first-quality sample images and second-quality sample images, where the image quality of the second-quality sample images is higher than that of the first-quality sample images;
[0193] The image feature determination module 1720 is configured to input the first-quality sample image and the second-quality sample image into a multi-scale vector quantization variational autoencoder respectively to obtain multiple first image features of different scales of the first-quality sample image, and multiple second image features of different scales of the second-quality sample image;
[0194] The image feature prediction module 1730 is configured to obtain the semantic feature of the first-quality sample image, use the semantic feature as the starting point of inference, start from the first image feature at the smallest scale, and input it into the image enhancement model in sequence, and infer the predicted value of the first image feature at the next scale through the image enhancement model;
[0195] The sample image enhancement module 1740 is configured to execute to obtain an enhanced sample image corresponding to the first quality sample image according to the predicted value of the first image feature at the maximum scale, and obtain a first loss according to the enhanced sample image and the second quality sample image;
[0196] The model parameter update module 1750 is configured to execute to obtain a second loss according to the predicted values of the first image features at each scale and the second image feature, obtain an overall loss according to the first loss and the second loss, and iteratively update the model parameters in the image enhancement model based on the overall loss.
[0197] In some exemplary embodiments of the present disclosure, the training data construction module 1710 may include a second quality sample image acquisition unit, an image degradation processing unit, and a training data pair construction unit.
[0198] Wherein:
[0199] The second quality sample image acquisition unit is configured to execute to acquire a second quality sample image set, and the second quality sample image set includes a plurality of second quality sample images;
[0200] The image degradation processing unit is configured to execute to obtain a plurality of different degradation paths according to a plurality of different degradation operations, and perform degradation processing on the second quality sample image according to the degradation path to obtain a first quality sample image corresponding to the second quality sample image;
[0201] The training data pair construction unit is configured to execute to construct a training data pair based on the second quality sample image and the corresponding first quality sample image, and obtain a training data set of the image enhancement model according to multiple groups of training data pairs.
[0202] In some exemplary embodiments of the present disclosure, the image feature determination module 1720 may include a residual latent space vector generation unit and a first image feature determination unit. Wherein:
[0203] The residual latent space vector generation unit is configured to execute to input the first quality sample image into a multi-scale vector quantization variational autoencoder, and decompose the continuous latent space vector corresponding to the first quality sample image into residual latent space vectors of different scales through the multi-scale vector quantization variational autoencoder;
[0204] The first image feature determination unit is configured to execute to obtain multiple first image features of different scales of the first quality sample image according to the residual latent space vectors of different scales.
[0205] In some exemplary embodiments of the present disclosure, the first image feature determination unit may include a vector mapping unit and a mapping result determination unit. Wherein:
[0206] A vector mapping unit, configured to perform quantization processing on residual latent space vectors of different scales through a quantizer, and map the residual latent space vectors into a codebook;
[0207] A mapping result determination unit, configured to obtain first image features of multiple different scales of a first quality sample image according to mapping results of residual latent space vectors of different scales in the codebook, where the residual latent space vectors of different scales share the codebook.
[0208] In some exemplary embodiments of the present disclosure, the image feature prediction module 1730 may include a semantic feature acquisition unit, configured to input the first quality sample image into a semantic feature extractor to obtain the semantic features of the first quality sample image.
[0209] In some exemplary embodiments of the present disclosure, the model parameter update module 1750 may include a multi-layer perceptron parameter update unit, configured to perform iterative update on model parameters in the multi-layer perceptron of the semantic feature extractor based on the overall loss.
[0210] In some exemplary embodiments of the present disclosure, the image feature prediction module 1730 may further include an initial prediction value determination unit, an input feature determination unit, and a next feature prediction unit. Among them:
[0211] The initial prediction value determination unit is configured to use the semantic features as the prediction value of the first image features at the smallest scale;
[0212] The input feature determination unit is configured to obtain the input features at the current scale according to the semantic features, the first image features at the current scale, and the prediction value of the first image features at the previous scale;
[0213] The next feature prediction unit is configured to input the input features at the current scale into an image enhancement model, and infer the prediction value of the first image features at the next scale through the image enhancement model.
[0214] In some exemplary embodiments of the present disclosure, the next feature prediction unit may include an image feature input unit, a semantic feature input unit, a sampling feature determination unit, and a downsampling unit. Among them:
[0215] The image feature input unit is configured to add the first image features at the current scale and the prediction value of the first image features at the previous scale, and input them into a plurality of cascaded conversion modules;
[0216] The semantic feature input unit is configured to perform normalization processing on the semantic features through a linear layer and an adaptive layer, and respectively input them into each conversion module to perform offset and numerical scaling processing on the feature layers in the conversion modules;
[0217] The sampling feature determination unit is configured to perform upsampling on the output of the last conversion module to obtain the sampling features corresponding to the current scale, and add the sampling features corresponding to the current scale to the sampling features corresponding to all scales before the current scale to obtain the total sampling features;
[0218] The downsampling unit is configured to perform downsampling on the total sampling features according to the target size of the current scale to obtain the predicted value of the first image feature of the next scale.
[0219] In some exemplary embodiments of the present disclosure, the sampling feature determination unit may include a maximum size determination unit and an upsampling unit. Among them:
[0220] The maximum size determination unit is configured to determine the maximum size of upsampling according to the input size of the first quality sample image and the sampling multiple of the multi-scale vector quantization variational autoencoder in the spatial dimension;
[0221] The upsampling unit is configured to perform upsampling on the output of the last conversion module according to the maximum size to obtain the sampling features corresponding to the current scale.
[0222] Furthermore, the present disclosure also provides an image enhancement device. Referring to Figure 18 As shown, the image enhancement device may include a first feature determination module 1810, a first feature prediction module 1820, and an image restoration and enhancement module 1830. Among them:
[0223] The first feature determination module 1810 is configured to input the first quality image into the multi-scale vector quantization variational autoencoder to obtain the first image features of multiple different scales of the first quality image;
[0224] The first feature prediction module 1820 is configured to obtain the semantic features of the first quality image, use the semantic features as the starting point of reasoning, start from the first image features of the smallest scale, and sequentially input them into the image enhancement model, and infer the predicted value of the first image features of the next scale through the image enhancement model;
[0225] The image restoration and enhancement module 1830 is configured to obtain the restored and enhanced image corresponding to the first quality image according to the predicted value of the first image features of the largest scale;
[0226] Among them, the image enhancement model is obtained by the training method of the above image enhancement model.
[0227] The specific details of each module / unit in the above image enhancement model training device and image enhancement device have been described in detail in the corresponding method embodiment part, and will not be repeated here.
[0228] Figure 19 The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure.
[0229] It should be noted that Figure 19 the computer system 1900 of the shown electronic device is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0230] As Figure 19 shown, the computer system 1900 includes a central processing unit (CPU) 1901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1902 or the program loaded from the storage section 1908 into the random access memory (RAM) 1903. In the RAM 1903, various programs and data required for system operation are also stored. The CPU 1901, ROM 1902, and RAM 1903 are connected to each other via a bus 1904. The input / output (I / O) interface 1905 is also connected to the bus 1904.
[0231] The following components are connected to the I / O interface 1905: an input section 1906 including a keyboard, a mouse, etc.; an output section 1907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1908 including a hard disk, etc.; and a communication section 1909 including a network interface card such as a LAN card, a modem, etc. The communication section 1909 performs communication processing via a network such as the Internet. A drive 1910 is also connected to the I / O interface 1905 as required. A removable medium 1911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1910 as required, so that the computer program read from it can be installed into the storage section 1908 as required.
[0232] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1909, and / or installed from the removable medium 1911. When the computer program is executed by the central processing unit (CPU) 1901, various functions defined in the system of the present disclosure are executed.
[0233] The exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned method for training an image enhancement model is implemented.
[0234] In one embodiment, a computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.
[0235] In one embodiment, a computer program product may be an intangible product containing a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as an executable file storing the computer program, digital files such as installation packages.
[0236] The code of the computer program can be written in one or more programming languages. Programming languages such as C, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).
[0237] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, can cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as the training method of the above image enhancement model can be executed.
[0238] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0239] It should be noted that although several modules of devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to embodiments of the present disclosure, the features and functions of two or more of the above-described modules may be embodied in one module. Conversely, the features and functions of one module described above may be further divided and embodied by multiple modules.
[0240] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure.
[0241] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for an image enhancement model, characterized in that: include: Constructing a training data set for an image enhancement model, wherein the training data set includes a plurality of training data pairs consisting of a first quality sample image and a second quality sample image, wherein the image quality of the second quality sample image is higher than that of the first quality sample image; Inputting the first quality sample image and the second quality sample image into a multi-scale vector quantization variational autoencoder respectively, to obtain a plurality of first image features of different scales of the first quality sample image, and a plurality of second image features of different scales of the second quality sample image; Acquire semantic features of the first quality sample image, use the semantic features as a reasoning starting point, start from the first image feature of the smallest scale, sequentially input them into the image enhancement model, and obtain predicted values of the first image features of the next scale through reasoning by the image enhancement model; Obtaining an enhanced sample image corresponding to the first quality sample image according to the predicted value of the first image feature of the largest scale, and obtaining a first loss according to the enhanced sample image and the second quality sample image; A second loss is obtained according to the predicted values of the first image features at each scale and the second image features, an overall loss is obtained according to the first loss and the second loss, and model parameters in the image enhancement model are iteratively updated based on the overall loss.
2. The method for training an image enhancement model according to claim 1, characterized in that: The training data set for constructing the image enhancement model includes: Acquire a second quality sample image set, where the second quality sample image set includes a plurality of second quality sample images; Obtaining a plurality of different degradation paths according to a plurality of different degradation operations, and performing degradation processing on the second quality sample image according to the degradation paths to obtain a first quality sample image corresponding to the second quality sample image; A training data pair is constructed based on the second quality sample image and the corresponding first quality sample image, and a training data set of the image enhancement model is obtained according to a plurality of sets of the training data pairs.
3. The training method of the image enhancement model according to claim 1, characterized in that: Inputting the first quality sample image into a multi-scale vector quantization variational autoencoder to obtain a plurality of first image features of different scales of the first quality sample image, including: Inputting the first quality sample image into a multi-scale vector quantization variational autoencoder, and decomposing the continuous latent space vector corresponding to the first quality sample image into residual latent space vectors of different scales through the multi-scale vector quantization variational autoencoder; A plurality of first image features of different scales of the first quality sample image are obtained according to the residual latent space vectors of different scales.
4. The method for training an image enhancement model according to claim 3, characterized in that: The obtaining first image features of a plurality of different scales of the first quality sample image according to the residual latent space vectors of different scales includes: quantizing the residual latent space vectors of different scales by a quantizer, and mapping the residual latent space vectors into a codebook; A plurality of first image features of different scales of the first quality sample image are obtained according to the mapping results of the residual latent space vectors of different scales in the codebook, wherein the residual latent space vectors of different scales share a codebook.
5. The method for training an image enhancement model according to claim 1, characterized in that: The acquiring the semantic feature of the first quality sample image includes: The first quality sample image is input into a semantic feature extractor to obtain the semantic features of the first quality sample image.
6. The method for training an image enhancement model according to claim 5, characterized in that: The semantic feature extractor includes a weighted language image pre-training model and a multi-layer perceptron, and the method further includes: The model parameters in the multilayer perceptron of the semantic feature extractor are iteratively updated based on the overall loss.
7. The method for training an image enhancement model according to claim 1, characterized in that: The method of using the semantic feature as the inference starting point, starting from the first image feature of the smallest scale, sequentially inputting the features into the image enhancement model, and obtaining the predicted value of the first image feature of the next scale through inference by the image enhancement model includes: Using the semantic feature as a predicted value of the first image feature of the minimum scale; Obtaining an input feature of the current scale according to the semantic feature, the first image feature of the current scale, and the predicted value of the first image feature of the previous scale; The input feature of the current scale is input into the image enhancement model, and the predicted value of the first image feature of the next scale is obtained by inference through the image enhancement model.
8. The method for training an image enhancement model according to claim 7, characterized in that: The image enhancement model includes a plurality of conversion modules, and the input feature of the current scale is input into the image enhancement model, and the predicted value of the first image feature of the next scale is obtained by inference through the image enhancement model, including: Add the predicted values of the first image feature of the current scale and the first image feature of the previous scale, and input them into a plurality of the conversion modules connected in series; After the semantic features are normalized by the linear layer and the adaptive layer, they are respectively input into each of the conversion modules, and the feature layers in the conversion modules are offset and numerically scaled; Upsampling the output of the last conversion module to obtain sampling features corresponding to the current scale, and adding the sampling features corresponding to the current scale and the sampling features corresponding to all scales before the current scale to obtain a total sampling feature; The total sampled features are downsampled according to the target size of the current scale to obtain a predicted value of the first image feature of the next scale.
9. The method for training an image enhancement model according to claim 8, characterized in that: The up-sampling the output of the last conversion module to obtain the sampling features corresponding to the current scale includes: Determine a maximum size of upsampling according to an input size of the first quality sample image and a sampling multiple of the multi-scale vector quantization variational autoencoder in a spatial dimension; The output of the last conversion module is upsampled according to the maximum size to obtain sampling features corresponding to the current scale.
10. An image enhancement method, characterized in that: include: Inputting the first quality image into a multi-scale vector quantization variational autoencoder to obtain first image features of multiple different scales of the first quality image; Acquire semantic features of the first quality image, use the semantic features as a reasoning starting point, start from the first image features of the smallest scale, sequentially input them into the image enhancement model, and obtain predicted values of the first image features of the next scale through reasoning by the image enhancement model; Obtaining a repaired and enhanced image corresponding to the first quality image according to a predicted value of the first image feature at a maximum scale; Wherein, the image enhancement model is obtained by the training method of the image enhancement model as described in any one of claims 1 to 9.
11. A training device for an image enhancement model, characterized in that: include: A training data construction module is configured to execute construction of a training data set for an image enhancement model, wherein the training data set includes a plurality of training data pairs consisting of a first quality sample image and a second quality sample image, wherein the image quality of the second quality sample image is higher than that of the first quality sample image; an image feature determination module, configured to input the first quality sample image and the second quality sample image into a multi-scale vector quantization variational autoencoder respectively, to obtain a plurality of first image features of different scales of the first quality sample image, and a plurality of second image features of different scales of the second quality sample image; An image feature prediction module is configured to obtain semantic features of the first quality sample image, use the semantic features as a reasoning starting point, start from the first image feature of the smallest scale, sequentially input them into the image enhancement model, and obtain a predicted value of the first image feature of the next scale through reasoning by the image enhancement model; a sample image enhancement module, configured to obtain an enhanced sample image corresponding to the first quality sample image according to a predicted value of the first image feature of the maximum scale, and obtain a first loss according to the enhanced sample image and the second quality sample image; The model parameter updating module is configured to execute the second loss obtained according to the predicted values of the first image features at each scale and the second image features, obtain the overall loss according to the first loss and the second loss, and iteratively update the model parameters in the image enhancement model based on the overall loss.
12. An image enhancement device, characterized in that: include: A first feature determination module is configured to input a first quality image into a multi-scale vector quantization variational autoencoder to obtain first image features of multiple different scales of the first quality image; A first feature prediction module is configured to execute the acquisition of semantic features of the first quality image, take the semantic features as the inference starting point, start from the first image features of the smallest scale, sequentially input them into the image enhancement model, and obtain the predicted value of the first image features of the next scale through the inference of the image enhancement model; An image restoration and enhancement module is configured to obtain a restoration and enhancement image corresponding to the first quality image according to a prediction value of the first image feature at a maximum scale; Wherein, the image enhancement model is obtained by the training method of the image enhancement model as described in any one of claims 1 to 9.
13. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of the image enhancement model as described in any one of claims 1 to 9 or the image enhancement method as described in claim 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the image enhancement model as described in any one of claims 1 to 9 or the image enhancement method as described in claim 10.
Citation Information
Cited By
VLA model pre-training method based on Internet video multi-scale decoupling
CN122416349A
A VLA model pre-training method based on multi-scale decoupling of internet video.
CN122416349B