Enhancement model training method, enhancement method, device, equipment and program product
By training video quality enhancement models, using noisy images and random inactivation branching technology, the problem of detail loss in video encoding is solved, and the clarity and bit rate reduction space of video texture details are improved.
Patent Information
- Application Number
- CN202411998182.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
During the video encoding process, the information compression effect leads to the loss of video picture details and the user's subjective feelings about video clarity decrease.
A training method for video quality enhancement models is adopted to train by obtaining high-quality and low-quality video images, and the model structure is adjusted using noise images and random inactivation branches, model losses are calculated and parameters are adjusted to enhance video texture details.
It effectively improves the texture detail clarity of the video picture, improves users' subjective feelings about video clarity, and brings greater code rate reduction space to video encoding without losing picture quality, and reduces video transmission bandwidth costs.
Smart Images

Figure CN119941534A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine learning, and in particular to a video quality enhancement model training method, a video quality enhancement method, a video quality enhancement model training device, a video quality enhancement device, an electronic device, and a computer program product. Background Art
[0002] With the rapid development of the short video industry and the continuous improvement of user demand, the video content on the platform is increasing, and users' requirements for video quality and fluency are also increasing.
[0003] However, during the video encoding process, the information compression effect brought about by the encoding often causes the problem of loss of details in the video picture, resulting in a decrease in the user's subjective perception of the clarity of the video being watched.
[0004] In view of this, the art urgently needs a video quality enhancement model training method and a video quality enhancement method that can enhance the texture details of the video picture and improve the user's subjective feeling in terms of clarity when watching the video.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The purpose of the present disclosure is to provide a video quality enhancement model training method, a video quality enhancement method, a video quality enhancement model training device, a video quality enhancement device, an electronic device and a computer program product, thereby being able to enhance the texture details of the video picture at least to a certain extent and improve the user's subjective feeling of clarity when watching the video.
[0007] According to a first aspect of the present disclosure, a method for training a video quality enhancement model is provided, comprising:
[0008] Acquire first training data of the video quality enhancement model, wherein the first training data includes a plurality of training data pairs, wherein the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality;
[0009] Obtaining a first input training image according to the first video image and the first noise, and inputting the first input training image into the video quality enhancement model to obtain a corresponding first output training image;
[0010] Calculating a first model loss of the video quality enhancement model according to a first model loss function, the first output training image and the second video image;
[0011] The model parameters of the video quality enhancement model are trained according to the first model loss to obtain the trained video quality enhancement model.
[0012] In an exemplary embodiment of the present disclosure, obtaining a first input training image according to the first video image and the first noise includes:
[0013] Adding first noise to the first video image to obtain a first noise image;
[0014] The first noise image is copied, and the copied first noise images are spliced to obtain the first input training image.
[0015] In an exemplary embodiment of the present disclosure, the method further includes:
[0016] Obtaining a second input training image according to the first video image and the second noise, and inputting the second input training image into the trained video quality enhancement model to obtain a corresponding second output training image;
[0017] Calculate a second model loss of the video quality enhancement model according to the first model loss function, the second model loss function and the third model loss function, as well as the first output training image and the second video image;
[0018] The model parameters of the video quality enhancement model are adjusted according to the second model loss to obtain the fine-tuned video quality enhancement model.
[0019] In an exemplary embodiment of the present disclosure, obtaining a second input training image according to the first video image and the second noise includes:
[0020] Adding second noise to the first video image to obtain a second noise image;
[0021] The first video image and the second noise image are spliced to obtain the second input training image.
[0022] In an exemplary embodiment of the present disclosure, the method further includes:
[0023] Determining a target convolutional layer from a plurality of convolutional layers of the video quality enhancement model;
[0024] A random dropout branch is connected in parallel to the target convolutional layer to adjust the structure of the video quality enhancement model.
[0025] According to a second aspect of the present disclosure, a video quality enhancement method is provided, comprising:
[0026] Acquire a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model, wherein the video quality enhancement model is obtained by the training method of the video quality enhancement model as described above;
[0027] An enhanced second quality video is obtained through the video quality enhancement model, and the video quality of the second quality video is higher than the video quality of the first quality video.
[0028] According to a third aspect of the present disclosure, a training device for a video quality enhancement model is provided, comprising:
[0029] a training data acquisition module, configured to acquire first training data of the video quality enhancement model, wherein the first training data includes a plurality of training data pairs, wherein the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality;
[0030] An input image determination module is configured to obtain a first input training image according to the first video image and the first noise, and input the first input training image into the video quality enhancement model to obtain a corresponding first output training image;
[0031] A model loss determination module is configured to calculate a first model loss of the video quality enhancement model according to a first model loss function, the first output training image and the second video image;
[0032] The enhancement model training module is configured to perform training on the model parameters of the video quality enhancement model according to the first model loss to obtain the trained video quality enhancement model.
[0033] In an exemplary embodiment of the present disclosure, the input image determination module includes:
[0034] A first noise image determining unit is configured to add first noise to the first video image to obtain a first noise image;
[0035] The first noise image splicing unit is configured to copy the first noise image and splice the copied first noise images to obtain the first input training image.
[0036] In an exemplary embodiment of the present disclosure, the video quality enhancement model training device further includes an enhancement model fine-tuning module, and the enhancement model fine-tuning module includes:
[0037] A second input training image determining unit is configured to obtain a second input training image according to the first video image and the second noise, and input the second input training image into the trained video quality enhancement model to obtain a corresponding second output training image;
[0038] A second model loss determination unit is configured to calculate a second model loss of the video quality enhancement model according to the first model loss function, the second model loss function and the third model loss function, as well as the first output training image and the second video image;
[0039] The model parameter fine-tuning unit is configured to adjust the model parameters of the video quality enhancement model according to the second model loss to obtain the fine-tuned video quality enhancement model.
[0040] In an exemplary embodiment of the present disclosure, the second input training image determining unit includes:
[0041] A second noise image determining unit is configured to add second noise to the first video image to obtain a second noise image;
[0042] The second noise image stitching unit is configured to stitch the first video image with the second noise image to obtain the second input training image.
[0043] In an exemplary embodiment of the present disclosure, the training device of the video quality enhancement model further includes a random deactivation module, and the random deactivation module includes:
[0044] A target convolutional layer determination unit, configured to determine a target convolutional layer from a plurality of convolutional layers of the video quality enhancement model;
[0045] The random deactivation branch parallel unit is configured to execute the target convolution layer in parallel with a random deactivation branch to adjust the structure of the video quality enhancement model.
[0046] According to a fourth aspect of the present disclosure, a video quality enhancement device is provided, comprising:
[0047] A video image input module is configured to obtain a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model, wherein the video quality enhancement model is obtained by the training method of the video quality enhancement model as described above;
[0048] The video image enhancement module is configured to execute a second quality video enhanced by the video quality enhancement model, wherein the video quality of the second quality video is higher than the video quality of the first quality video.
[0049] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the training method of the video quality enhancement model described in any one of the above.
[0050] According to a sixth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the training method of the video quality enhancement model described in any one of the above is implemented.
[0051] The exemplary embodiments of the present disclosure may have the following beneficial effects:
[0052] In the training method of the video quality enhancement model in the example implementation of the present disclosure, a first input training image is obtained by adding a first noise to a first video image of a first quality, and the first input training image is input into the video quality enhancement model to obtain a corresponding first output training image, and then the first model loss of the video quality enhancement model is calculated according to the first output training image and the second video image of a second quality, and then the model parameters of the video quality enhancement model are trained according to the first model loss to obtain the trained video quality enhancement model. The training method of the video quality enhancement model in the example implementation of the present disclosure, on the one hand, by introducing reasonable random perturbations in the training process of the model, and by using noise with a large degree of perturbation and randomness, the model's ability to generate texture details with a certain degree of randomness from the input is improved, thereby improving the model's ability to enhance texture and details, thereby effectively improving the model's ability to enhance the clarity of texture details, and alleviating the problem of model generalization difficulties; on the other hand, it can also bring greater bit rate reduction space for subsequent video encoding without losing image quality when the model scale is limited, and ultimately effectively reduce the bandwidth cost of the video during transmission, and improve the model's ability to process real online videos.
[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0055] Figure 1 A flowchart of a method for training a video quality enhancement model according to an exemplary embodiment of the present disclosure is shown;
[0056] Figure 2 A schematic diagram of a process of obtaining a first input training image according to a first video image and a first noise according to an exemplary embodiment of the present disclosure is shown;
[0057] Figure 3 A schematic flow chart showing a method for fine-tuning a video quality enhancement model according to an exemplary embodiment of the present disclosure;
[0058] Figure 4 A schematic diagram of a process of obtaining a second input training image according to a first video image and a second noise according to an exemplary embodiment of the present disclosure is shown;
[0059] Figure 5 A schematic diagram showing a flow chart of a video quality enhancement method according to an exemplary embodiment of the present disclosure;
[0060] Figure 6 A schematic flow chart of a method for training and fine-tuning a video quality enhancement model in a specific embodiment of the present disclosure is shown;
[0061] Figure 7 A block diagram of a video quality enhancement model training apparatus according to an exemplary embodiment of the present disclosure is shown;
[0062] Figure 8 A block diagram showing a video quality enhancement apparatus according to an exemplary embodiment of the present disclosure is shown;
[0063] Fig. 9 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0064] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0065] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein.
[0066] The following example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0067] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0068] In some relevant embodiments, the increasing number of short videos can be efficiently processed by developing corresponding video processing chips. The video processing chips can provide powerful hardware support for the processing and optimization of video content, and can be applied to short video platforms and live broadcast products to enhance the platform's real-time processing capabilities in video encoding, real-time rendering, image processing, video analysis, etc., and improve the user's video viewing experience. The video processing chip's video encoding processing can reduce the video transmission bandwidth cost and improve the video viewing fluency.
[0069] Since the information compression effect brought by encoding often causes the loss of details in the video during the video encoding process, a video quality enhancement algorithm can be used to pre-process the video before encoding. The goal is to enhance the texture details of the video so that the video has more details after encoding, thereby improving the user's subjective feeling of clarity when watching the video.
[0070] At the same time, the bandwidth cost of videos has gradually become an important expense that cannot be ignored in the operation of video platforms. Therefore, encoding and compressing videos has become an indispensable part of saving bandwidth costs. However, the loss of video texture caused by the video encoding process will affect the user's subjective clarity. Therefore, it is necessary to pre-process the input video before video encoding to enhance the detail texture and improve the subjective clarity experience of the video. For example, when designing a pre-processing algorithm on a lightweight video processing chip, it is necessary to consider improving the model's ability to generate texture and details under a limited model scale, thereby effectively improving the strength and generalization ability of the model to enhance texture details.
[0071] Deploying a pre-processing algorithm for video detail texture enhancement on a lightweight video processing chip requires considering the limited computing power of the video processing chip itself while achieving real-time processing of large quantities of video. Therefore, it is necessary to constrain the complexity of the pre-processing algorithm so that it can effectively improve the texture details of the video quality while occupying a small amount of computing power and computing space of the video processing chip.
[0072] The image quality enhancement algorithm for pre-processing of video processing chips mainly reduces the model complexity based on the re-parameterization mechanism and enhances the texture details of the video quality under the condition of limited computing power, such as RepSR (Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch Normalization). Specifically, the current algorithm uses a method based on RealESRGAN (real enhanced version of the blind image super-resolution model) to degrade video images in the training strategy, thereby constructing degraded image-original image pairs as training data. During the training process, the L1 loss function and the loss function based on GAN (generative adversarial network) are used for constraints, so that the network can improve the sharpness of the input image, so that the network outputs a sharper image. The current algorithm has the advantages of high efficiency and good real-time performance, and can also enhance the texture to a certain extent. However, due to the limited size of the model in the algorithm, this limits the texture enhancement capability of the current algorithm, which in turn hinders its application in higher bitrate video compression.
[0073] However, the above algorithm still has deficiencies in generating highly complex textures and details, which leads to the problem of severe detail loss when compressing video images at a higher bit rate. Specifically, the current algorithm only takes the input image as the only input and samples and generates textures and details from the input. Due to the determinism of a single input, the smaller model size of the current algorithm cannot sample texture details with a certain degree of randomness from a certain input, which limits the further improvement of the ability to generate texture details and ultimately leads to the current algorithm's shortcomings in texture enhancement. In addition, the current algorithm relies too much on the method of constructing data pairs based on RealESRGAN, which will also lead to the problem of difficulty in generalization of the trained model in real scenes, which will also lead to insufficient texture enhancement, which in turn leads to the problem of subjective clarity reduction in texture details in the encoded video.
[0074] Based on the above problems, this example implementation first provides a method for training a video quality enhancement model. Figure 1 As shown, the training method of the above video quality enhancement model may include the following steps:
[0075] Step S110. Acquire first training data of a video quality enhancement model, wherein the first training data includes a plurality of training data pairs, each of the training data pairs includes a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality.
[0076] Step S120: Obtain a first input training image according to the first video image and the first noise, and input the first input training image into a video quality enhancement model to obtain a corresponding first output training image.
[0077] Step S130. Calculate the first model loss of the video quality enhancement model according to the first model loss function, the first output training image and the second video image.
[0078] Step S140. Train the model parameters of the video quality enhancement model according to the first model loss to obtain a trained video quality enhancement model.
[0079] In the training method of the video quality enhancement model in the example implementation of the present disclosure, a first input training image is obtained by adding a first noise to a first video image of a first quality, and the first input training image is input into the video quality enhancement model to obtain a corresponding first output training image, and then the first model loss of the video quality enhancement model is calculated according to the first output training image and the second video image of a second quality, and then the model parameters of the video quality enhancement model are trained according to the first model loss to obtain the trained video quality enhancement model. The training method of the video quality enhancement model in the example implementation of the present disclosure, on the one hand, by introducing reasonable random perturbations in the training process of the model, and by using noise with a large degree of perturbation and randomness, the model's ability to generate texture details with a certain degree of randomness from the input is improved, thereby improving the model's ability to enhance texture and details, thereby effectively improving the model's ability to enhance the clarity of texture details, and alleviating the problem of model generalization difficulties; on the other hand, it can also bring greater bit rate reduction space for subsequent video encoding without losing image quality when the model scale is limited, and ultimately effectively reduce the bandwidth cost of the video during transmission, and improve the model's ability to process real online videos.
[0080] Next, combine Figures 2 to 4 The above steps of this exemplary embodiment are described in more detail.
[0081] In step S110, first training data of a video quality enhancement model is obtained, the first training data including a plurality of training data pairs, the training data pairs including a first video image of a first quality and a second video image of a second quality, the second quality being higher than the first quality.
[0082] In this example implementation, the initialized video quality enhancement model may be generatively pre-trained, and the training at this stage may be performed through first training data, such as 10,000-20,000 images, wherein a plurality of training data pairs are included, and the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality. The first video image of the first quality refers to a low-quality video image with a lower video definition and resolution, and the second video image of the second quality refers to a high-quality video image with a higher video definition and resolution.
[0083] In step S120, a first input training image is obtained according to the first video image and the first noise, and the first input training image is input into a video quality enhancement model to obtain a corresponding first output training image.
[0084] In this example implementation, Figure 2 As shown, obtaining a first input training image according to a first video image and a first noise may specifically include the following steps:
[0085] Step S210: Add first noise to the first video image to obtain a first noise image.
[0086] The first noise is Gaussian noise with a random variance within a first interval, for example, Gaussian noise with a random variance of 0.3-0.6. The first video image is added with Gaussian noise with a random variance within the first interval to obtain a first noise image.
[0087] Step S220: Copy the first noise image, and splice the copied first noise images to obtain a first input training image.
[0088] The first noise image is copied into two and spliced together to obtain a first input training image as input data of the model.
[0089] In this example implementation, the generative mechanism can determine the generated results through randomness, which not only enhances the diversity of the model generation results, but also improves the robustness and generalization ability of the algorithm. Proper introduction of randomness can improve the flexibility of the generation process.
[0090] In step S130, a first model loss of the video quality enhancement model is calculated according to the first model loss function, the first output training image and the second video image.
[0091] In this example implementation, during the training phase of the video quality enhancement model, only the first model loss function, i.e., the training loss function based on GAN, can be used for constraint calculation to calculate the first model loss. Generative pre-training is performed only under the guidance of GAN constraints to ensure that the model has a certain degree of generative capability.
[0092] In step S140, model parameters of the video quality enhancement model are trained according to the first model loss to obtain a trained video quality enhancement model.
[0093] Finally, the model parameters of the video quality enhancement model are updated based on the first model loss to obtain the trained video quality enhancement model. In this way, generative pre-training is training an image conversion task similar to denoising, aiming to convert the distribution of inputs that are severely affected by noise corruption to the distribution of clean images. Since only GAN is used as a constraint on the loss function, the entire model can sample textures and details that only conform to the distribution of clean images, without completely corresponding to the target clean image. In this way, the model can have a relatively strong randomness in generating textures and details, and the trained model has the ability to model a wide range of textures and details.
[0094] In this example implementation, by introducing appropriate random input and training mechanisms, the network's texture and detail generation capabilities are appropriately improved while maintaining a certain fidelity, thereby effectively improving the subjective clarity of video texture details.
[0095] After the video quality enhancement model is trained through the above steps, it can also be generatively fine-tuned to improve the model's ability to depict texture details.
[0096] In this example implementation, Figure 3 As shown, the fine-tuning method of the video quality enhancement model may specifically include the following steps:
[0097] Step S310: Obtain a second input training image according to the first video image and the second noise, and input the second input training image into the trained video quality enhancement model to obtain a corresponding second output training image.
[0098] For the pre-trained video quality enhancement model, generative fine-tuning can be performed. Less training data can be used for training at this stage. For example, fine-tuning training can be performed on 1000-2000 images.
[0099] In this example implementation, Figure 4 As shown, obtaining a second input training image according to the first video image and the second noise may specifically include the following steps:
[0100] Step S410: Add second noise to the first video image to obtain a second noise image.
[0101] The second noise is a Gaussian noise with a random variance size within the second interval, wherein the random variance size within the second interval is smaller than the random variance size within the first interval, for example, the second noise may be a Gaussian noise with a random variance size of 0-0.02. By using noise with a smaller perturbation degree and randomness, the process of enhancing texture details of the fitting model can be achieved while retaining a certain degree of randomness.
[0102] Step S420: splice the first video image and the second noise image to obtain a second input training image.
[0103] By splicing the first video image with the second noise image, a second input training image can be obtained as input data of the model.
[0104] Step S320. Calculate the second model loss of the video quality enhancement model according to the first model loss function, the second model loss function and the third model loss function, as well as the first output training image and the second video image.
[0105] In the process of generative fine-tuning, the model loss can be calculated based on a variety of loss functions. For example, the second model loss can be calculated using the L1 loss function with the target clean image, the loss function based on FFT (fast Fourier transform), and the GAN loss function.
[0106] Step S330: Adjust the model parameters of the video quality enhancement model according to the second model loss to obtain a fine-tuned video quality enhancement model.
[0107] Finally, the model parameters of the video quality enhancement model are adjusted based on the second model loss to obtain a fine-tuned video quality enhancement model. Since the input is strongly constrained by the target clean image during the generative fine-tuning process, the texture details can be enhanced with greater fidelity, thereby not destroying the semantic information of the image. At the same time, since the model is fine-tuned starting from the generative pre-trained model weights, and there are still GAN loss function constraints in the training process, and the input still has slight noise disturbances, the model still has a certain degree of randomness in generating textures and details, so that the model as a whole still has a strong ability to depict texture details.
[0108] In this example implementation, a Dropout (random deactivation) mechanism may also be introduced into the model during the fine-tuning process to reduce the sensitivity of the model to distribution changes and improve the generalization ability of the model to enhance texture details.
[0109] In this example implementation, a target convolution layer may be determined from multiple convolution layers of the video quality enhancement model, and a random deactivation branch may be connected in parallel to the target convolution layer to adjust the structure of the video quality enhancement model.
[0110] In order to avoid the model overfitting to the degraded space of RealESRGAN during fine-tuning, which would lead to a decrease in generalization ability, a Dropout-based connection strategy is introduced in the middle of the model. The specific implementation here is to bypass the Dropout branch with dynamic perception weights next to two convolutional layers of the model.
[0111] By introducing random perturbations based on Dropout, the model’s ability to learn the overall distribution of input to output can be enhanced, thereby avoiding overfitting to the design of the degradation operator and improving the generalization ability of texture detail enhancement.
[0112] On the other hand, this exemplary embodiment also provides a method for enhancing video quality. Figure 5 As shown, the above-mentioned video quality enhancement method may include the following steps:
[0113] Step S510: Obtain a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model.
[0114] The video quality enhancement model is obtained through the above video quality enhancement model training method.
[0115] Step S520: Obtain an enhanced second quality video through a video quality enhancement model, wherein the video quality of the second quality video is higher than the video quality of the first quality video.
[0116] The video quality enhancement method in this example implementation uses the video quality enhancement model obtained by the training method of the video quality enhancement model as described above to enhance the video image to be enhanced, which can effectively improve the texture details of the video and reduce the bandwidth cost of the video during transmission.
[0117] like Figure 6 The figure is a complete flow chart of a specific implementation of the present disclosure, which is an example of the above steps in this example implementation. The specific steps of the flow chart are as follows:
[0118] Step S610: Generative pre-training.
[0119] First, the initialized algorithm model is generatively pre-trained. This stage of training is performed on 10,000-20,000 images. Here, the input image is set to the input image itself plus Gaussian noise with a random variance of 0.3-0.6, and is copied into two copies and spliced together as input. The loss function is constrained only by the GAN-based training loss function. Generative pre-training is training an image conversion task similar to denoising, aiming to convert the distribution of inputs that are severely affected by noise damage to the distribution of clean images. Since only GAN is constrained as the loss function, the entire model can sample textures and details that only conform to the distribution of clean images, without completely corresponding to the target clean image. In this way, the model can have a relatively strong randomness in generating textures and details, and the trained model has the ability to model a wide range of textures and details.
[0120] Step S620: Generative fine-tuning.
[0121] Secondly, for the pre-trained algorithm model, generative fine-tuning is performed. This stage of training is performed on 1000-2000 images. Here, the input image is set to the input image itself, and the input image itself is spliced with Gaussian noise with a random variance of 0-0.02 as the model input. The loss function uses the L1 loss function with the target clean image, the loss function based on FFT transformation, and the GAN loss function. Since the input is strongly constrained by the target clean image during the generative fine-tuning process, the texture details can be enhanced with more fidelity, thereby not destroying the semantic information of the image. At the same time, since the model is fine-tuned from the model weights of generative pre-training, and the GAN loss function is still constrained during the training process, and the input is still slightly disturbed by noise, the model still has a certain degree of randomness in generating textures and details, so that the model still has a strong ability to depict texture details as a whole.
[0122] Step S630: Introduce Dropout.
[0123] In addition, in order to avoid overfitting to the degradation space of RealESRGAN during fine-tuning, which would lead to a decrease in generalization ability, a Dropout-based connection strategy can be introduced in the middle of the model. Specifically, a Dropout branch with dynamic perception weights can be bypassed next to two convolutional layers of the model. By introducing random perturbations based on Dropout, the model's ability to learn the overall distribution from input to output can be enhanced, thereby avoiding overfitting to the design of the degradation operator and improving the generalization ability for texture detail enhancement.
[0124] As for the improvement of clarity, the video quality enhancement model training method and the video quality enhancement method in this example implementation can reduce the bit rate by 7 percent while still achieving no decrease in clarity in subjective blind testing.
[0125] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0126] It should be noted that the user information involved in this disclosure (including but not limited to user device information, user personal information, etc.) is all information authorized by the user or fully authorized by all parties.
[0127] Furthermore, the present disclosure also provides a training device for a video quality enhancement model. Figure 7As shown, the video quality enhancement model training device may include a training data acquisition module 710, an input image determination module 720, a model loss determination module 730, and an enhancement model training module 740. Among them:
[0128] The training data acquisition module 710 is configured to acquire first training data of the video quality enhancement model, wherein the first training data includes a plurality of training data pairs, wherein the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality;
[0129] An input image determination module 720 is configured to obtain a first input training image according to the first video image and the first noise, and input the first input training image into a video quality enhancement model to obtain a corresponding first output training image;
[0130] The model loss determination module 730 is configured to calculate a first model loss of the video quality enhancement model according to the first model loss function, the first output training image and the second video image;
[0131] The enhancement model training module 740 is configured to perform training on the model parameters of the video quality enhancement model according to the first model loss to obtain a trained video quality enhancement model.
[0132] In some exemplary embodiments of the present disclosure, the input image determination module 720 may include a first noise image determination unit and a first noise image splicing unit.
[0133] A first noise image determining unit is configured to add first noise to the first video image to obtain a first noise image;
[0134] The first noise image splicing unit is configured to copy the first noise image and splice the copied first noise images to obtain a first input training image.
[0135] In some exemplary embodiments of the present disclosure, a video quality enhancement model training device provided by the present disclosure may further include an enhancement model fine-tuning module, which may include a second input training image determination unit, a second model loss determination unit, and a model parameter fine-tuning unit. Wherein:
[0136] A second input training image determining unit is configured to obtain a second input training image according to the first video image and the second noise, and input the second input training image into the trained video quality enhancement model to obtain a corresponding second output training image;
[0137] A second model loss determination unit is configured to calculate a second model loss of the video quality enhancement model according to the first model loss function, the second model loss function and the third model loss function, as well as the first output training image and the second video image;
[0138] The model parameter fine-tuning unit is configured to adjust the model parameters of the video quality enhancement model according to the second model loss to obtain a fine-tuned video quality enhancement model.
[0139] In some exemplary embodiments of the present disclosure, the second input training image determining unit may include a second noise image determining unit and a second noise image splicing unit.
[0140] A second noise image determining unit is configured to add second noise to the first video image to obtain a second noise image;
[0141] The second noise image stitching unit is configured to stitch the first video image with the second noise image to obtain a second input training image.
[0142] In some exemplary embodiments of the present disclosure, a video quality enhancement model training device provided by the present disclosure may further include a random deactivation module, which may include a target convolutional layer determination unit and a random deactivation branch parallel unit.
[0143] A target convolution layer determination unit is configured to determine a target convolution layer from a plurality of convolution layers of a video quality enhancement model;
[0144] The random deactivation branch parallel unit is configured to execute a random deactivation branch in parallel with the target convolution layer to adjust the structure of the video quality enhancement model.
[0145] Furthermore, the present disclosure also provides a video quality enhancement device. Figure 8 As shown, the video quality enhancement device may include a module 810 and a module 820. Among them:
[0146] The video image input module 810 is configured to obtain a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model, wherein the video quality enhancement model is obtained by the above video quality enhancement model training method;
[0147] The video image enhancement module 820 is configured to execute a second quality video enhanced by a video quality enhancement model, wherein the video quality of the second quality video is higher than the video quality of the first quality video.
[0148] The specific details of the training device of the above-mentioned video quality enhancement model and each module / unit in the video quality enhancement device have been described in detail in the corresponding method embodiment part, and will not be repeated here.
[0149] Fig. 9 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present disclosure is shown.
[0150] It should be noted that Fig. 9 The computer system 900 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0151] like Fig. 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for system operation are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0152] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed.
[0153] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 909, and / or installed from a removable medium 911. When the computer program is executed by a central processing unit (CPU) 901, various functions defined in the system of the present disclosure are executed.
[0154] The exemplary embodiments of the present disclosure also provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, the training method of the video quality enhancement model and the video quality enhancement method are implemented.
[0155] In one embodiment, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing a computer program. The readable storage medium may be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid-state drive (SSD), and the like. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing a computer program, such as a read-only memory, a NAND flash memory, and the like.
[0156] In one embodiment, the computer program product may be an intangible product including a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as a digital file storing an executable file, an installation package, etc. of the computer program.
[0157] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed entirely on the user computing device, or partially on the user computing device, or as a separate software package, or partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (e.g., an Internet connection provided by an operator).
[0158] The computer program may be carried or transmitted via electrical, magnetic, optical, electromagnetic, infrared, or other signals. The electronic device may convert the signal carrying the computer program into a digital signal, thereby running the computer program. When the computer program is run on an electronic device, its code is used to enable the electronic device to execute (more specifically, the processor of the electronic device may execute) the method steps of various exemplary embodiments of the present disclosure, such as the training method of the video quality enhancement model and the video quality enhancement method described above.
[0159] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0160] It should be noted that, although several modules of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules described above can be concretized in one module. Conversely, the features and functions of one module described above can be further divided into multiple modules to be concretized.
[0161] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.
[0162] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for training a video quality enhancement model, characterized in that: include: Acquire first training data of the video quality enhancement model, wherein the first training data includes a plurality of training data pairs, wherein the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality; Obtaining a first input training image according to the first video image and the first noise, and inputting the first input training image into the video quality enhancement model to obtain a corresponding first output training image; Calculating a first model loss of the video quality enhancement model according to a first model loss function, the first output training image and the second video image; The model parameters of the video quality enhancement model are trained according to the first model loss to obtain the trained video quality enhancement model.
2. The video quality enhancement model training method according to claim 1, characterized in that: The step of obtaining a first input training image according to the first video image and the first noise includes: Adding first noise to the first video image to obtain a first noise image; The first noise image is copied, and the copied first noise images are spliced to obtain the first input training image.
3. The video quality enhancement model training method according to claim 1, characterized in that: The method further comprises: Obtaining a second input training image according to the first video image and the second noise, and inputting the second input training image into the trained video quality enhancement model to obtain a corresponding second output training image; Calculate a second model loss of the video quality enhancement model according to the first model loss function, the second model loss function and the third model loss function, as well as the first output training image and the second video image; The model parameters of the video quality enhancement model are adjusted according to the second model loss to obtain the fine-tuned video quality enhancement model.
4. The video quality enhancement model training method according to claim 3, characterized in that: The step of obtaining a second input training image according to the first video image and the second noise includes: Adding second noise to the first video image to obtain a second noise image; The first video image and the second noise image are spliced to obtain the second input training image.
5. The video quality enhancement model training method according to claim 1, characterized in that: The method further comprises: Determining a target convolutional layer from a plurality of convolutional layers of the video quality enhancement model; A random dropout branch is connected in parallel to the target convolutional layer to adjust the structure of the video quality enhancement model.
6. A video quality enhancement method, characterized in that: include: Acquire a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model, wherein the video quality enhancement model is obtained by the training method of the video quality enhancement model according to any one of claims 1 to 5; An enhanced second quality video is obtained through the video quality enhancement model, and the video quality of the second quality video is higher than the video quality of the first quality video.
7. A training device for a video quality enhancement model, characterized in that: include: a training data acquisition module, configured to acquire first training data of the video quality enhancement model, wherein the first training data includes a plurality of training data pairs, wherein the training data pairs include a first video image of a first quality and a second video image of a second quality, wherein the second quality is higher than the first quality; An input image determination module is configured to obtain a first input training image according to the first video image and the first noise, and input the first input training image into the video quality enhancement model to obtain a corresponding first output training image; A model loss determination module is configured to calculate a first model loss of the video quality enhancement model according to a first model loss function, the first output training image and the second video image; The enhancement model training module is configured to perform training on the model parameters of the video quality enhancement model according to the first model loss to obtain the trained video quality enhancement model.
8. A video quality enhancement device, characterized in that: include: a video image input module, configured to obtain a first quality video to be enhanced, and input a video image in the first quality video into a video quality enhancement model, wherein the video quality enhancement model is obtained by the training method of the video quality enhancement model according to any one of claims 1 to 5; The video image enhancement module is configured to execute a second quality video enhanced by the video quality enhancement model, wherein the video quality of the second quality video is higher than the video quality of the first quality video.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the video quality enhancement model training method according to any one of claims 1 to 5 or the video quality enhancement method according to claim 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of the video quality enhancement model according to any one of claims 1 to 5 or the video quality enhancement method according to claim 6 is implemented.