Image quality enhancement method, device, electronic device and storage medium
By generating code rate vectors and adjusting parameters, the problem of retraining the model after the video code rate changes is solved, adaptive image quality enhancement is achieved, and the R&D cycle is shortened.
Patent Information
- Application Number
- CN202410083522.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-01-19
AI Technical Summary
The existing technology requires retraining the image quality enhancement model after the video bit rate changes, resulting in cumbersome and lengthy processes.
By obtaining the code rate of the video stream, the code rate vector is generated, and the pre-trained image processing model generates adjustment parameters based on the code rate vector, and adjusts the feature map extracted by the image processing model during the image quality enhancement process to achieve a high correlation between the image enhancement effect and the code rate.
When the encoding rate changes, there is no need to retrain the model. The image processing model can be adaptively adjusted, shortening the model's R&D cycle and ensuring the image quality enhancement effect.
Smart Images

Figure CN117857879B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image quality enhancement, and in particular, to an image quality enhancement method, apparatus, electronic device, and storage medium. Background Art
[0002] In a live broadcast scenario, the broadband cost accounts for a relatively large proportion of the enterprise's operating cost. To save bandwidth costs, the encoding bitrate of the video stream is usually set relatively low, and in subsequent optimizations, the bitrate may be further reduced to control costs. The video bitrate and video image quality are usually highly correlated. For videos with the same image quality, there will be a large difference in image quality between the encoding results with a high bitrate and those with a low bitrate. In this context, when applying an image quality enhancement model, it is necessary to consider the impact of the bitrate on the images input to the image quality enhancement model.
[0003] Currently, the common practice in the business is that after the video bitrate changes, test whether the enhancement effect of the previous image quality enhancement model meets the requirements. If the effect requirements are not met, it is necessary to reconstruct the training data based on the images corresponding to the new video bitrate and retrain the model, and the entire process is cumbersome and lengthy. Summary of the Invention
[0004] In view of this, an object of the present invention is to provide an image quality enhancement method, apparatus, electronic device, and storage medium to solve the problem in the prior art that when the video bitrate changes, it is necessary to retrain the model, resulting in a cumbersome and lengthy entire process.
[0005] To achieve the above object, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, the present invention provides an image quality enhancement method, and the method includes:
[0007] Obtain a bitrate vector according to the encoding bitrate of the video stream;
[0008] Input the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model;
[0009] Generate adjustment parameters based on the bitrate vector through the image processing model, and perform image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted by the image processing model during the image quality enhancement processing.
[0010] In an alternative embodiment, the image processing model includes an image quality enhancement network and a bitrate modulation network; the method of generating adjustment parameters based on the bitrate vector through the image processing model and performing image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed includes:
[0011] Generating adjustment parameters based on the bitrate vector through the bitrate modulation network and outputting the adjustment parameters to the image quality enhancement network;
[0012] Performing image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed.
[0013] In an alternative embodiment, the bitrate modulation network includes a feature extraction module, a first adjustment parameter generation module, and a second adjustment parameter generation module; the method of generating adjustment parameters based on the bitrate vector through the bitrate modulation network and outputting the adjustment parameters to the image quality enhancement network includes:
[0014] Extracting features from the bitrate vector through the feature extraction module and respectively outputting the obtained bitrate feature vectors to the first adjustment parameter generation module and the second adjustment parameter generation module;
[0015] Extracting features from the bitrate feature vector through the first adjustment parameter generation module and generating first adjustment parameters to be output to the image quality enhancement network;
[0016] Extracting features from the bitrate feature vector through the second adjustment parameter generation module and generating second adjustment parameters to be output to the image quality enhancement network.
[0017] In an alternative embodiment, the image quality enhancement network includes an encoder, a decoder, and a residual connection module. The input end of the encoder and the output end of the decoder are both connected to the residual connection module, and the output end of the encoder is connected to the input end of the decoder; the method of performing image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed includes:
[0018] Extracting features from the image to be processed through the encoder, adjusting the data distribution of the depth feature map extracted in the encoder according to the first adjustment parameter, and finally outputting an encoder feature map to the decoder;
[0019] The encoder feature map is subjected to feature extraction by the decoder, and the data distribution of the depth feature map extracted in the decoder is adjusted according to the second adjustment parameter, and finally a decoder feature map is output to the residual connection module;
[0020] The decoder feature map and the image to be processed are added together by the residual connection module to obtain a high-quality image corresponding to the image to be processed.
[0021] In an alternative embodiment, the adjusted depth feature map in the encoder is where F C represents the depth feature map before adjustment in the encoder, F' C represents the depth feature map after adjustment in the encoder, μ(F C ) and σ(F C ) represent the mean and standard deviation corresponding to F C respectively, and σ C and μ C are obtained according to the first adjustment parameter;
[0022] The adjusted depth feature map in the decoder is where F D represents the depth feature map before adjustment in the decoder, F' D represents the depth feature map after adjustment in the decoder, μ(F D ) and σ(F D ) represent the mean and standard deviation corresponding to F D respectively, and σ D and μ D are obtained according to the second adjustment parameter.
[0023] In an alternative embodiment, the obtaining of the code rate vector according to the coding bit rate of the video stream includes:
[0024] When there is a preset bit rate in multiple preset bit rates that is the same as the coding bit rate of the video stream, the coding bit rate of the video stream is quantized to obtain a code rate vector;
[0025] When there is no preset bit rate in multiple preset bit rates that is the same as the coding bit rate of the video stream, a target bit rate closest to the coding bit rate of the video stream among the multiple preset bit rates is determined, and the target bit rate is quantized to obtain a code rate vector.
[0026] In an alternative embodiment, the image processing model is trained through the following steps:
[0027] Select sample images and high-quality sample images with the same content as the sample images from a pre-built training dataset; the training dataset includes image datasets corresponding to multiple preset bitrates, and the sample images in each image dataset are extracted from a sample video corresponding to the preset bitrate, and different sample videos are obtained by encoding the same video using different preset bitrates;
[0028] Quantize the preset bitrate corresponding to the sample image to obtain a corresponding preset bitrate vector;
[0029] Use the high-quality sample image as the label of the sample image, and input the sample image with the label and the preset bitrate vector into a pre-built image processing model to obtain the output result of the image processing model;
[0030] Calculate a loss value based on the output result of the image processing model and the label of the sample image;
[0031] Iteratively update the parameters of the image processing model according to the loss value, and finally obtain a trained image processing model.
[0032] In a second aspect, the present invention provides an image quality enhancement device, and the device includes:
[0033] A quantization module for obtaining a bitrate vector according to the encoding bitrate of a video stream;
[0034] An input module for inputting the bitrate vector and each frame of image to be processed in the video stream into a pre-trained image processing model;
[0035] A processing module for generating adjustment parameters based on the bitrate vector through the image processing model, and performing image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted by the image processing model during the image quality enhancement processing.
[0036] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the image quality enhancement method according to any one of the foregoing embodiments are implemented.
[0037] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the image quality enhancement method according to any one of the foregoing embodiments are implemented.
[0038] The image quality enhancement method, device, electronic device, and storage medium provided by the embodiments of the present invention. The method includes: obtaining a bitrate vector according to the encoding bitrate of the video stream, inputting the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model, generating adjustment parameters based on the bitrate vector through the image processing model, and performing image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted during the image quality enhancement processing by the image processing model. Since the image processing model can adjust the extracted feature maps based on the bitrate vector corresponding to the encoding bitrate, the image enhancement effect is highly correlated with the encoding bitrate. When the encoding bitrate changes, it can adapt to the changed encoding bitrate and ensure the image quality enhancement effect at this encoding bitrate without retraining the model, thus greatly shortening the R & D cycle of the model.
[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0041] Figure 1 Shows a schematic flowchart of an image quality enhancement method provided by an embodiment of the present invention;
[0042] Figure 2 Shows a schematic diagram of an image processing model;
[0043] Figure 3 Shows another schematic diagram of an image processing model;
[0044] Figure 4 Shows a corresponding relationship diagram between a preset bitrate and an image data set;
[0045] Figure 5 Shows a functional module diagram of an image quality enhancement device provided by an embodiment of the present invention;
[0046] Figure 6 Shows a schematic block diagram of an electronic device provided by an embodiment of the present invention.
[0047] Icons: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 600 - Image quality enhancement device; 610 - Quantization module; 620 - Input module; 630 - Processing module. Detailed implementation
[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0049] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0050] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0051] In the prior art, in order to ensure the image quality enhancement effect after the video bitrate changes, after the video bitrate changes, it is necessary to test whether the enhancement effect of the previous image quality enhancement model meets the requirements. If the effect requirements are not met, it is necessary to reconstruct the training data based on the image corresponding to the new video bitrate and retrain the model, and the whole process is cumbersome and lengthy.
[0052] Based on this, an image quality enhancement method, device, electronic device, and storage medium are proposed in the embodiments of the present invention. By generating adjustment parameters from the code rate vector corresponding to the encoding bit rate, the feature maps extracted by the image processing model during the image quality enhancement process are adjusted, achieving that the image enhancement effect is highly correlated with the encoding bit rate. When the encoding bit rate changes, the image processing model can adapt to the changed encoding bit rate and ensure the image quality enhancement effect at this encoding bit rate without retraining the model, greatly shortening the model research and development cycle.
[0053] Next, each embodiment of the present invention will be described in detail with reference to the accompanying drawings.
[0054] Please refer to Figure 1 , which is a schematic flowchart of an image quality enhancement method provided by an embodiment of the present invention. It should be noted that the image quality enhancement method of the present invention is not limited to Figure 1 the following specific order. It should be understood that in other embodiments, the order of some steps of the image quality enhancement method of the present invention can be interchanged according to actual needs, or some of the steps can also be omitted or deleted. This image quality enhancement method can be applied to electronic devices such as servers and PCs (Personal Computers). Next, the Figure 1 specific process shown will be elaborated in detail.
[0055] Step S101, obtain a code rate vector according to the encoding bit rate of the video stream.
[0056] In this embodiment, after the electronic device obtains the encoding bit rate of the video stream, it is necessary to convert the encoding bit rate of the video stream into a vector that can be processed by the image processing model, that is, the code rate vector.
[0057] Step S102, input the code rate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model.
[0058] In this embodiment, after the electronic device obtains the code rate vector, it inputs the code rate vector and each frame of the image to be processed in the video stream into the image processing model for processing.
[0059] Step S103, generate adjustment parameters based on the code rate vector through the image processing model, and perform image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted by the image processing model during the image quality enhancement process.
[0060] In this embodiment, after the bitrate vector and the image to be processed are input into the image processing model, the image processing model generates corresponding adjustment parameters according to the input bitrate vector, and performs image quality enhancement processing on the input image to be processed based on the adjustment parameters, adjusts the extracted feature map through the adjustment parameters, and finally outputs a high-quality image corresponding to the image to be processed.
[0061] It can be seen that the image quality enhancement method provided by the embodiment of the present invention obtains a bitrate vector according to the encoding bitrate of the video stream, inputs the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model, generates adjustment parameters based on the bitrate vector through the image processing model, and performs image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature map extracted during the image quality enhancement processing by the image processing model. Since the image processing model can adjust the extracted feature map based on the bitrate vector corresponding to the encoding bitrate, the image enhancement effect is highly correlated with the encoding bitrate. When the encoding bitrate changes, it can adapt to the changed encoding bitrate and ensure the image quality enhancement effect at this encoding bitrate without retraining the model, thus greatly shortening the R & D cycle of the model.
[0062] In this embodiment, in order to facilitate the image processing model to adapt to different encoding bitrates, during the training stage of the image processing model, the highest bitrate and the lowest bitrate of the video are usually determined with reference to the service scenario, and multiple bitrate levels are divided between the highest bitrate and the lowest bitrate, so that the image processing model is fully trained at each bitrate level, thereby ensuring the image quality enhancement effect at this bitrate level. Therefore, during the application stage of the image processing model, the bitrate vector input into the image processing model also needs to ensure that it is the bitrate vector corresponding to the set bitrate level.
[0063] Based on this, in step S101 above, it may specifically include: when there is a preset bitrate in the multiple preset bitrates that is the same as the encoding bitrate of the video stream, quantize the encoding bitrate of the video stream to obtain a bitrate vector; when there is no preset bitrate in the multiple preset bitrates that is the same as the encoding bitrate of the video stream, determine the target bitrate in the multiple preset bitrates that is closest to the encoding bitrate of the video stream, and quantize the target bitrate to obtain a bitrate vector.
[0064] In this embodiment, assume that the highest bitrate is b max , the lowest bitrate is b min , and S preset bitrates are divided, then the preset bitrate can be expressed as b s = b max - (b max - b min ) × s / (S - 1), s ∈ [0, S - 1]. When s = 0, b s is the highest bitrate, and when s = S - 1, bs is the lowest bit rate.
[0065] After obtaining the encoding bit rate of the video stream, it is possible to check whether there is a b that is the same as the encoding bit rate of the video stream among the S preset bit rates. s If there is a matching b s then directly quantize the encoding bit rate of the video stream to obtain a bit rate vector and input it into the image processing model; if there is no matching b s then it is possible to find the b that is closest to the encoding bit rate of the video stream s and quantize this closest b s to obtain a bit rate vector and input it into the image processing model.
[0066] For example, the highest bit rate is set to 5000k and the lowest bit rate is set to 300k. Since the highest and lowest bit rates are determined based on the reference service scenario, the encoding bit rate in the service must be between the highest and lowest bit rates. When the encoding bit rate of the video stream is reduced from 2000k to 1500k to save bandwidth in the service, then based on b s = b max -(b max -b min )×s / (S - 1), the preset bit rate b s closest to 1500k can be calculated, and b s is quantized to obtain a bit rate vector.
[0067] In one implementation, please refer to Figure 2 , the image processing model may include an image quality enhancement network and a bit rate modulation network. The above step S103 specifically includes: generating adjustment parameters based on the bit rate vector through the bit rate modulation network and outputting the adjustment parameters to the image quality enhancement network; performing image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed.
[0068] That is to say, the bit rate vector quantized based on the encoding bit rate of the video stream is input into the bit rate modulation network of the image processing model, and the bit rate modulation network outputs adjustment parameters to the image quality enhancement network based on the bit rate vector; the input of the image quality enhancement network is each frame of the image to be processed in the video stream. During the process of performing image quality enhancement processing on each frame of the image to be processed, the image quality enhancement network also uses the adjustment parameters output by the bit rate modulation network to adjust the data distribution of the extracted feature maps, so as to ensure a good image quality enhancement effect.
[0069] In one implementation, please refer to Figure 3, the bitrate modulation network specifically includes a feature extraction module, a first adjustment parameter generation module, and a second adjustment parameter generation module, and the image quality enhancement network specifically includes an encoder, a decoder, and a residual connection module. The output end of the feature extraction module is respectively connected to the input ends of the first adjustment parameter generation module and the second adjustment parameter generation module. The output end of the first adjustment parameter generation module is connected to the encoder, and the output end of the second adjustment parameter generation module is connected to the decoder. The input end of the encoder and the output end of the decoder are both connected to the residual connection module, and the output end of the encoder is connected to the input end of the decoder.
[0070] The above steps generate adjustment parameters based on the bitrate vector through the bitrate modulation network and output the adjustment parameters to the image quality enhancement network. Specifically, it can include: extracting features of the bitrate vector through the feature extraction module, and respectively outputting the obtained bitrate feature vectors to the first adjustment parameter generation module and the second adjustment parameter generation module; extracting features of the bitrate feature vector through the first adjustment parameter generation module and generating the first adjustment parameter to output to the image quality enhancement network; extracting features of the bitrate feature vector through the second adjustment parameter generation module and generating the second adjustment parameter to output to the image quality enhancement network.
[0071] Among them, the bitrate modulation network is mainly composed of fully connected layers, which can extract features from the input bitrate vector and generate different adjustment parameters for the image quality enhancement network to adjust the data distribution of the feature maps. That is to say, the feature extraction module, the first adjustment parameter generation module, and the second adjustment parameter generation module can all be implemented through fully connected layers.
[0072] In this embodiment, the feature extraction module is the backbone network. After the quantized bitrate vector is input into the backbone network to extract features, the first adjustment parameter and the second adjustment parameter are respectively generated through two branches of the first adjustment parameter generation module and the second adjustment parameter generation module. The first adjustment parameter generated by the first adjustment parameter generation module will be output to the encoder in the image quality enhancement network to adjust the feature map extracted in the encoder; the second adjustment parameter generated by the second adjustment parameter generation module will be output to the decoder in the image quality enhancement network to adjust the feature map extracted in the decoder. Among them, the first adjustment parameter and the second adjustment parameter are both vectors.
[0073] The above steps perform image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network, and obtain a high-quality image corresponding to the image to be processed. Specifically, it may include: extracting features of the image to be processed through an encoder, adjusting the data distribution of the depth feature map extracted in the encoder according to the first adjustment parameter, and finally outputting the encoder feature map to the decoder; extracting features of the encoder feature map through the decoder, adjusting the data distribution of the depth feature map extracted in the decoder according to the second adjustment parameter, and finally outputting the decoder feature map to the residual connection module; adding the decoder feature map and the image to be processed through the residual connection module to obtain a high-quality image corresponding to the image to be processed.
[0074] In this embodiment, the encoder and the decoder can be implemented by a fully convolutional network and are used to output feature maps of different sizes. The use of the residual connection module is to facilitate the training convergence of the model, and the model only needs to learn the difference between the input image and the reconstructed image.
[0075] In this embodiment, the encoder and the decoder can only adjust the generated depth feature maps and do not need to adjust the shallow feature maps. The depth feature map can be understood as a small-scale feature map, which reflects the deep features of the image. For example, it can be set to adjust the feature maps output by the last two convolutional layers in the encoder and the feature maps output by the first two convolutional layers in the decoder, so as to realize the adjustment of the data distribution of the entire feature map.
[0076] In one implementation, the adjusted depth feature map in the encoder is where F C represents the depth feature map before adjustment in the encoder, F′ C represents the depth feature map after adjustment in the encoder, μ(F C ) and σ(F C ) respectively represent the mean and standard deviation corresponding to F C , and σ C and μ C are obtained according to the first adjustment parameter; the adjusted depth feature map in the decoder is where F D represents the depth feature map before adjustment in the decoder, F′ D represents the depth feature map after adjustment in the decoder, μ(F D ) and σ(F D ) respectively represent the mean and standard deviation corresponding to F D , and σ D and μ D are obtained according to the second adjustment parameter.
[0077] For example, assume that the number of channels of the depth feature maps output by a certain layer in the encoder and decoder is both c, the rate vector input to the rate modulation network is an 8-dimensional vector, and the first adjustment parameter and the second adjustment parameter output are both 2c-dimensional vectors. Then, the first c-dimensional vector in the first adjustment parameter can be used as the mean μ C , and the last c-dimensional vector in the first adjustment parameter can be used as the standard deviation σ C . Similarly, the first c-dimensional vector in the second adjustment parameter is used as the mean μ D , and the last c-dimensional vector in the second adjustment parameter is used as the standard deviation σ D .
[0078] In one implementation, the above image processing model is trained through the following steps: Select a sample image and a high-quality sample image with the same content as the sample image from a pre-constructed training dataset; the training dataset includes image datasets corresponding to multiple preset rates, and the sample images in each image dataset are extracted from the sample videos corresponding to the preset rates. Different sample videos are obtained by encoding the same video using different preset rates; Quantize the preset rate corresponding to the sample image to obtain the corresponding preset rate vector; Use the high-quality sample image as the label of the sample image, input the labeled sample image and the preset rate vector into the pre-constructed image processing model to obtain the output result of the image processing model; Calculate the loss value according to the output result of the image processing model and the label of the sample image; Iteratively update the parameters of the image processing model according to the loss value, and finally obtain the trained image processing model.
[0079] In this embodiment, before training the image processing model, a training dataset will be constructed. First, it is necessary to determine the highest rate b max and the lowest rate b min of the video, and then determine the granularity between the rates, that is, how many rate levels (preset rates) are divided between the highest rate and the lowest rate. Assume that S preset rates are divided, then the preset rate can be expressed as b s = b max - (b max - b min ) × s / (S - 1), s ∈ [0, S - 1]. When s = 0, b s is the highest rate, and when s = S - 1, b s is the lowest rate. Encode the same video using different preset rates to obtain sample videos corresponding to different preset rates, then extract image frames from the sample videos corresponding to different preset rates to obtain sample images corresponding to different preset rates, and generate corresponding image datasets according to the sample images corresponding to each preset rate.
[0080] For example, encoding the same video using the highest bitrate and the lowest bitrate to obtain the highest bitrate video and the lowest bitrate video, with other encoding parameters adapted to the corresponding bitrates. Extracting image frames from the highest bitrate video and the lowest bitrate video to obtain the highest quality image and the lowest quality image. The video bitrate and the corresponding quality image form a paired bitrate-image dataset. The same bitrate can correspond to quality images of different contents, which is a one-to-many relationship. Thus, for multiple preset bitrates (from the lowest bitrate to the highest bitrate), corresponding image datasets (from the lowest quality image dataset to the highest quality image dataset) will be obtained, as Figure 4 shown.
[0081] Assume that the sample image I s is from the image dataset corresponding to the preset bitrate b s . Take the high-quality sample image I max with the same content as the sample image I s in the image dataset corresponding to the preset bitrate b hq as the label for supervised training. Quantize the preset bitrate b s corresponding to the sample image I s into a preset bitrate vector. For example, when S 8 = 256, the preset bitrate b s can be quantized into an 8-dimensional vector. For example, b0 is quantized into 00000000, and b 255 is quantized into 11111111. Input the labeled sample image I s and the preset bitrate vector into a pre-constructed image processing model (as Figure 3 shown). After the preset bitrate vector is processed by the feature extraction module, the first adjustment parameter generation module and the second adjustment parameter generation module respectively generate the first adjustment parameter (σ C and μ C ) and the second adjustment parameter (σ D and μ D ). In the image quality enhancement network, the first adjustment parameter and the second adjustment parameter are used to adjust the depth feature maps extracted from the encoder and the decoder respectively. Suppose the depth feature map output by a certain layer in the encoder is F C ∈R c×h×w , then the adjusted depth feature map is Then input the adjusted depth feature map into the next layer of the network for processing; similarly, the adjusted depth feature map in the decoder is F′ D = Thus, the adjustment of the data distribution of the feature map is realized. After obtaining the reconstructed high-quality image I sq output by the image quality enhancement network (the output result of the image processing model), use the difference between I sq and I hq to optimize the image processing model.
[0082] In this embodiment, it is possible to calculate the loss value according to I sq and I hq to measure the image quality difference between I sq and I hq through the loss value. Among them, the loss value includes pixel loss L mse , perceptual loss L mse and GAN loss L mse . The pixel loss can be expressed as L mse = ‖I hq - I sq ‖, the perceptual loss can be expressed as L perp = ||VGG(I hq ) - VGG(I sq )‖, and the GAN loss can be expressed as L gan = ||GAN(I hq ) - GAN(I sq )||. Among them, the perceptual loss is calculated using the VGG (Visual Geometry Group) network; the GAN loss is calculated using GAN (Generative Adversarial Network). The specific calculation formula adopts the existing technology and will not be elaborated here. The final loss value can be expressed as L = α1L mse + α2L perp + α3L gan , where α1, α2, and α3 respectively represent the weights of the pixel loss, perceptual loss, and GAN loss.
[0083] In this embodiment, since the depth feature map generated in the image processing model will be adjusted based on the preset bit rate b s corresponding to I s , it will also be affected by gradient descent. Therefore, when optimizing the loss L, the image quality enhancement network and the bit rate modulation network are optimized simultaneously. Finally, the effect of the image quality enhancement network obtained is highly correlated with the input bit rate.
[0084] During actual training, s can be set to 0 at the beginning of training, that is, b s = b max . Then, the input sample image at this time is the highest quality image corresponding to the highest bit rate b max (high-quality sample image I hq) Due to the existence of the global residual connection, the model only needs to learn an identity mapping at this time. As the number of training iterations increases, s gradually increases, corresponding to a decrease in the video bitrate (preset bitrate), and the image quality of the input sample images gradually deteriorates. The corresponding supervision label is still the highest-quality image corresponding to the highest bitrate. The learning difficulty of the model gradually increases until s is S - 1, that is, b s = b min , then the input sample image I s is the lowest-quality image corresponding to the lowest bitrate b min . Maintaining s at S - 1 for a period of time makes the entire model training more sufficient. That is to say, as s gradually increases, the learning difficulty of the model will gradually increase. Correspondingly, the number of iterative training times of the model at the corresponding bitrate will also increase. By adopting this training method with gradually increasing difficulty, the model can be guided to optimize in a fixed direction, accelerating the convergence of the model. At each preset bitrate, the model will be iteratively trained several times to ensure that the high-quality image effect can be fully fitted at the preset bitrate.
[0085] It can be seen that the embodiments of the present invention pre-construct a paired bitrate-image data set, sample from the image data set according to certain rules, and based on different bitrate parameters, the model adaptively adjusts the image quality, dynamically controls the image quality enhancement effect. For example, for images with higher image quality, the model needs to perform a small degree of image quality enhancement to avoid distortion problems caused by over-enhancement. For images with lower image quality, the model needs to perform a large degree of image quality enhancement to improve the clarity of the entire picture.
[0086] To execute the corresponding steps in the above embodiments and each possible manner, the following provides an implementation manner of an image quality enhancement device. Please refer to Figure 5 , which is a functional module diagram of an image quality enhancement device 600 provided by the embodiments of the present invention. It should be noted that the basic principle and the technical effects generated by the image quality enhancement device 600 provided in this embodiment are the same as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The image quality enhancement device 600 includes: a quantization module 610, an input module 620, and a processing module 630.
[0087] The quantization module 610 is used to obtain a bitrate vector according to the encoding bitrate of the video stream.
[0088] It can be understood that the quantization module 610 can execute the above step S101.
[0089] The input module 620 is used to input the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model.
[0090] It can be understood that the input module 620 can execute the above step S102.
[0091] A processing module 630, configured to generate an adjustment parameter based on a bitrate vector through an image processing model, and perform image quality enhancement processing on the image to be processed according to the adjustment parameter to obtain a high-quality image corresponding to the image to be processed; the adjustment parameter is used to adjust the feature map extracted by the image processing model during the image quality enhancement processing.
[0092] It can be understood that the processing module 630 can execute the above step S103.
[0093] Optionally, the quantization module 610 is specifically configured to quantize the encoding bitrate of the video stream to obtain a bitrate vector when there is a preset bitrate among multiple preset bitrates that is the same as the encoding bitrate of the video stream; when there is no preset bitrate among multiple preset bitrates that is the same as the encoding bitrate of the video stream, determine a target bitrate among the multiple preset bitrates that is closest to the encoding bitrate of the video stream, and quantize the target bitrate to obtain a bitrate vector.
[0094] Optionally, the image processing model includes an image quality enhancement network and a bitrate modulation network. The processing module 630 is specifically configured to generate an adjustment parameter based on the bitrate vector through the bitrate modulation network and output the adjustment parameter to the image quality enhancement network; perform image quality enhancement processing on the image to be processed based on the adjustment parameter through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed.
[0095] Optionally, the bitrate modulation network includes a feature extraction module, a first adjustment parameter generation module, and a second adjustment parameter generation module. The processing module 630 is specifically configured to perform feature extraction on the bitrate vector through the feature extraction module, and output the obtained bitrate feature vector to the first adjustment parameter generation module and the second adjustment parameter generation module respectively; perform feature extraction on the bitrate feature vector through the first adjustment parameter generation module and generate a first adjustment parameter to output to the image quality enhancement network; perform feature extraction on the bitrate feature vector through the second adjustment parameter generation module and generate a second adjustment parameter to output to the image quality enhancement network.
[0096] Optionally, the image quality enhancement network includes an encoder, a decoder, and a residual connection module. The input end of the encoder and the output end of the decoder are both connected to the residual connection module, and the output end of the encoder is connected to the input end of the decoder. The processing module 630 is specifically configured to extract features from the image to be processed through the encoder, adjust the data distribution of the depth feature map extracted in the encoder according to the first adjustment parameter, and finally output the encoder feature map to the decoder; extract features from the encoder feature map through the decoder, adjust the data distribution of the depth feature map extracted in the decoder according to the second adjustment parameter, and finally output the decoder feature map to the residual connection module; add the decoder feature map and the image to be processed through the residual connection module to obtain the high-quality image corresponding to the image to be processed.
[0097] Optionally, the adjusted depth feature map in the encoder is where F C represents the depth feature map before adjustment in the encoder, and F' C represents the depth feature map after adjustment in the encoder. μ(F C ) and σ(F C ) respectively represent the mean and standard deviation corresponding to F C , and σ C and μ C are obtained according to the first adjustment parameter. The adjusted depth feature map in the decoder is where F D represents the depth feature map before adjustment in the decoder, and F' D represents the depth feature map after adjustment in the decoder. μ(F D ) and σ(F D ) respectively represent the mean and standard deviation corresponding to F D , and σ D and μ D are obtained according to the second adjustment parameter.
[0098] Optionally, the above image processing model can be trained in the following way: select a sample image and a high-quality sample image with the same content as the sample image from a pre-constructed training dataset; the training dataset includes image datasets corresponding to multiple preset bitrates, and the sample images in each image dataset are extracted from the sample videos corresponding to the preset bitrates. Different sample videos are obtained by encoding the same video using different preset bitrates; quantize the preset bitrate corresponding to the sample image to obtain the corresponding preset bitrate vector; use the high-quality sample image as the label of the sample image, input the labeled sample image and the preset bitrate vector into the pre-constructed image processing model to obtain the output result of the image processing model; calculate the loss value according to the output result of the image processing model and the label of the sample image; iteratively update the parameters of the image processing model according to the loss value, and finally obtain the trained image processing model.
[0099] It can be seen that the image quality enhancement device provided by the embodiment of the present invention includes a quantization module, an input module, and a processing module. The quantization module is used to obtain a bitrate vector according to the encoding bitrate of the video stream; the input module is used to input the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model; the processing module is used to generate an adjustment parameter based on the bitrate vector through the image processing model, and perform image quality enhancement processing on the image to be processed according to the adjustment parameter to obtain a high-quality image corresponding to the image to be processed; the adjustment parameter is used to adjust the feature map extracted by the image processing model during the image quality enhancement processing. Since the image processing model can adjust the extracted feature map based on the bitrate vector corresponding to the encoding bitrate, the image enhancement effect is highly correlated with the encoding bitrate. When the encoding bitrate changes, it can adapt to the changed encoding bitrate and ensure the image quality enhancement effect at this encoding bitrate without retraining the model, thus greatly shortening the R & D cycle of the model.
[0100] Please refer to Figure 6 , which is a block diagram of an electronic device 100 provided by an embodiment of the present invention. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The elements of the memory 110, the processor 120, and the communication module 130 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.
[0101] Among them, the memory 110 is used to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), etc.
[0102] The processor 120 is used to read / write the data or programs stored in the memory 110 and perform corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the image quality enhancement method disclosed in the above embodiments can be implemented.
[0103] The communication module 130 is used to establish a communication connection between the electronic device 100 and other devices through a network, and is used to send and receive data through the network.
[0104] It should be understood that Figure 6 the structure shown is only a schematic diagram of the structure of the electronic device 100, and the electronic device 100 may also include more or fewer components than Figure 6 those shown in, or have a different configuration from Figure 6 those shown. Figure 6 Each component shown in can be implemented by hardware, software, or a combination thereof.
[0105] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 120, the image quality enhancement method disclosed in the above embodiments is implemented.
[0106] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0107] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0108] When the above-described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0109] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image quality enhancement method, characterized in that, The method includes: Obtaining a bitrate vector according to the encoding bitrate of the video stream; Inputting the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model; the image processing model includes an image quality enhancement network and a bitrate modulation network; Generating adjustment parameters based on the bitrate vector through the image processing model, and performing image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed, including: generating adjustment parameters based on the bitrate vector through the bitrate modulation network, and outputting the adjustment parameters to the image quality enhancement network; performing image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted during the image quality enhancement processing of the image processing model.
2. The image quality enhancement method according to claim 1, wherein The bitrate modulation network includes a feature extraction module, a first adjustment parameter generation module, and a second adjustment parameter generation module; The generating adjustment parameters based on the bitrate vector through the bitrate modulation network and outputting the adjustment parameters to the image quality enhancement network includes: Performing feature extraction on the bitrate vector through the feature extraction module, and respectively outputting the obtained bitrate feature vector to the first adjustment parameter generation module and the second adjustment parameter generation module; Performing feature extraction on the bitrate feature vector through the first adjustment parameter generation module, and generating a first adjustment parameter and outputting it to the image quality enhancement network; Performing feature extraction on the bitrate feature vector through the second adjustment parameter generation module, and generating a second adjustment parameter and outputting it to the image quality enhancement network.
3. The image quality enhancement method according to claim 2, wherein The image quality enhancement network includes an encoder, a decoder, and a residual connection module. The input end of the encoder and the output end of the decoder are both connected to the residual connection module, and the output end of the encoder is connected to the input end of the decoder; The performing image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed includes: Performing feature extraction on the image to be processed through the encoder, and adjusting the data distribution of the depth feature map extracted in the encoder according to the first adjustment parameter, and finally outputting the encoder feature map to the decoder; Performing feature extraction on the encoder feature map through the decoder, and adjusting the data distribution of the depth feature map extracted in the decoder according to the second adjustment parameter, and finally outputting the decoder feature map to the residual connection module; Adding the decoder feature map and the image to be processed through the residual connection module to obtain a high-quality image corresponding to the image to be processed.
4. The image quality enhancement method according to claim 3, wherein The adjusted depth feature map in the encoder is where F C represents the depth feature map before adjustment in the encoder, and F ′ C represents the adjusted depth feature map in the encoder. μ(F C ) and σ(F C ) respectively represent the mean and standard deviation corresponding to F C , and σ C and μ C are obtained according to the first adjustment parameter; The adjusted depth feature map in the decoder is where F ′ represents the depth feature map before adjustment in the decoder, and F ′ D represents the adjusted depth feature map in the decoder. μ(F D ) and σ(F D ) represent the mean and standard deviation corresponding to F D respectively, and σ D and μ D are obtained according to the second adjustment parameter.
5. The image quality enhancement method according to claim 1, wherein The obtaining a bitrate vector according to the encoding bitrate of the video stream includes: When there is a preset bitrate in multiple preset bitrates that is the same as the encoding bitrate of the video stream, quantizing the encoding bitrate of the video stream to obtain a bitrate vector; When there is no preset bitrate in multiple preset bitrates that is the same as the encoding bitrate of the video stream, determine the target bitrate in the multiple preset bitrates that is closest to the encoding bitrate of the video stream, and quantize the target bitrate to obtain a bitrate vector.
6. The image quality enhancement method according to any one of claims 1-5, characterized in that, The image processing model is trained through the following steps: Select a sample image and a high-quality sample image with the same content as the sample image from a pre-constructed training dataset; the training dataset includes image datasets corresponding to multiple preset bitrates, and the sample images in each image dataset are extracted from the sample videos corresponding to the preset bitrates. Different sample videos are obtained by encoding the same video using different preset bitrates; Quantize the preset bitrate corresponding to the sample image to obtain the corresponding preset bitrate vector; Use the high-quality sample image as the label of the sample image, and input the sample image with the label and the preset bitrate vector into a pre-constructed image processing model to obtain the output result of the image processing model; Calculate the loss value according to the output result of the image processing model and the label of the sample image; Iteratively update the parameters of the image processing model according to the loss value, and finally obtain a trained image processing model.
7. An image quality enhancement device, characterized in that, The device includes: A quantization module for obtaining a bitrate vector according to the encoding bitrate of the video stream; An input module for inputting the bitrate vector and each frame of the image to be processed in the video stream into a pre-trained image processing model; A processing module for generating adjustment parameters based on the bitrate vector through the image processing model, and performing image quality enhancement processing on the image to be processed according to the adjustment parameters to obtain a high-quality image corresponding to the image to be processed; the adjustment parameters are used to adjust the feature maps extracted by the image processing model during the image quality enhancement process; the image processing model includes an image quality enhancement network and a bitrate modulation network; the processing module is used to generate adjustment parameters based on the bitrate vector through the bitrate modulation network, and output the adjustment parameters to the image quality enhancement network; perform image quality enhancement processing on the image to be processed based on the adjustment parameters through the image quality enhancement network to obtain a high-quality image corresponding to the image to be processed.
8. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the image quality enhancement method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the image quality enhancement method according to any one of claims 1-6.
Citation Information
Patent Citations
Model training and image processing method and device, medium and electronic equipment
CN110163237A
Image quality enhancement method and device, terminal and readable storage medium
CN114390307A
Image processing method, device and equipment and readable storage medium
CN116797466A
Video data processing method and device
CN117319708A