Reference frame selection method and apparatus, device, and storage medium
By acquiring multiple adjacent frames of the video frame to be processed, using an AI model to select the most suitable reference frame and perform image enhancement, the problem of insufficient reconstructed frame quality in existing technologies is solved, and higher quality video reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2026-03-24
AI Technical Summary
Existing video compression algorithms cannot effectively improve the image quality of reconstructed frames when selecting reference frames, resulting in visually unpleasant compression artifacts.
By acquiring multiple adjacent frames of the video frame to be processed, an AI model is used to select the reference frame that best matches the image content of the video frame to be processed, and the image quality is improved through an image enhancement model.
It improves the image quality of reconstructed frames, enhances the visual effects of video frames, and reduces compression artifacts.
Smart Images

Figure CN115243044B_ABST
Abstract
Description
Technical Field
[0001] This application relates to video image technology, including but not limited to reference frame selection methods, apparatus, devices, and storage media. Background Technology
[0002] Video has become the most popular form of content consumption. According to reports, video viewing accounted for 82% of all internet traffic by 2022. To reduce transmission bandwidth and storage costs, video service providers typically compress videos. However, some video compression algorithms, due to their block-transform-based coding methods, are prone to producing visually unpleasant compression artifacts. Therefore, developing video enhancement algorithms is essential.
[0003] In related technologies, a reference frame is first selected from the adjacent frames of the current frame, and then the current frame is enhanced based on the reference frame to obtain a reconstructed frame with better image quality than the current frame; however, the image quality of the reconstructed frame obtained based on related technologies cannot meet the quality requirements. Summary of the Invention
[0004] In view of this, the reference frame selection method, apparatus, device, and storage medium provided in this application are intended to select a better reference frame, thereby helping to better enhance the image quality of the video frame to be processed and improve the image quality of the reconstructed frame.
[0005] According to one aspect of the embodiments of this application, a reference frame selection method is provided, comprising: acquiring E first adjacent frames of a video frame to be processed; wherein, E is greater than 1; selecting a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; wherein, the first reference frame is used to enhance the image quality of the video frame to be processed.
[0006] Thus, since the image content of the video frame to be processed and its E first adjacent frames are considered when selecting the reference frame, rather than simply selecting the reference frame based on the positional relationship between the adjacent frames and the video frame to be processed (e.g., using the previous and next frames of the video frame to be processed as its reference frames), the selected reference frame is more compatible with the image content of the video frame to be processed, thereby better enhancing the image quality of the video frame to be processed and improving the image enhancement quality of the video frame to be processed.
[0007] According to one aspect of the embodiments of this application, a video enhancement method is provided, comprising: acquiring E first adjacent frames of a video frame to be processed; wherein, E is greater than 1; selecting a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; and enhancing the image quality of the video frame to be processed based on the first reference frame.
[0008] According to one aspect of the embodiments of this application, a reference frame selection apparatus is provided, comprising: an acquisition module configured to acquire E first adjacent frames of a video frame to be processed; wherein, E is greater than 1; and a selection module configured to select a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; wherein, the first reference frame is used to enhance the image quality of the video frame to be processed.
[0009] According to one aspect of the embodiments of this application, a video enhancement apparatus is provided, comprising: an acquisition module configured to acquire E first adjacent frames of a video frame to be processed; wherein E is greater than 1; a selection module configured to select a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; and an enhancement module configured to enhance the image quality of the video frame to be processed based on the first reference frame.
[0010] According to one aspect of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor executes the program to implement the method described in the embodiments of the present application.
[0011] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the methods provided in the embodiments of this application.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0014] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0015] Figure 1A schematic diagram illustrating the implementation flow of the reference frame selection method provided in the embodiments of this application;
[0016] Figure 2 A schematic diagram illustrating the implementation flow of another reference frame selection method provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram illustrating the relationship between the first adjacent frame and the video frame to be processed, provided in an embodiment of this application.
[0018] Figure 4 A schematic diagram illustrating the process of determining the trained image enhancement model provided in the embodiments of this application;
[0019] Figure 5 A schematic diagram illustrating the process of determining the trained reference frame selection model provided in an embodiment of this application;
[0020] Figure 6 A schematic diagram illustrating the process of determining the trained target image enhancement model provided in an embodiment of this application;
[0021] Figure 7 This is a schematic diagram of the video decompression and distortion process provided in an embodiment of this application;
[0022] Figure 8 A schematic diagram illustrating the workflow of the adaptive reference frame selection module provided in an embodiment of this application;
[0023] Figure 9 This is a schematic diagram of the training process provided in an embodiment of this application;
[0024] Figure 10 A schematic diagram comparing the subjective performance of decompression distortion of embodiments of this application and the heuristic reference frame selection method provided for the embodiments of this application;
[0025] Figure 11 This is a schematic diagram of the structure of the reference frame selection device provided in the embodiments of this application;
[0026] Figure 12 This is a schematic diagram of the hardware entity of the electronic device according to an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0029] In the following description, references to "some embodiments," "this embodiment," "this application embodiment," and examples, etc., describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0030] It should be noted that the terms "first, second, third, fourth, fifth," etc., used in the embodiments of this application are for distinguishing similar or different objects and do not represent a specific order of objects. It is understood that "first, second, third, fourth, fifth" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0031] This application provides a reference frame selection method applied to an electronic device. This electronic device can be of various types with information processing capabilities, such as mobile phones, tablets, desktop computers, televisions, and projection devices. The function implemented by this method can be achieved by a processor in the electronic device calling program code. The program code can be stored in a computer storage medium. Therefore, the electronic device includes at least a processor and a storage medium.
[0032] Figure 1 This is a schematic diagram illustrating the implementation flow of the reference frame selection method provided in the embodiments of this application, as follows: Figure 1 As shown, the method may include the following steps 101 to 102:
[0033] Step 101: Obtain the E first adjacent frames of the video frame to be processed; where E is greater than 1;
[0034] Step 102: Based on the image content of the video frame to be processed and the E first adjacent frames, select a first reference frame of the video frame to be processed from the E first adjacent frames; wherein, the first reference frame is used to enhance the image quality of the video frame to be processed.
[0035] In this embodiment, since the image content of the video frame to be processed and its E first adjacent frames are considered when selecting the reference frame, rather than simply selecting the reference frame based on the positional relationship between the adjacent frames and the video frame to be processed (e.g., taking the previous and next frames of the video frame to be processed as its reference frames), the selected reference frame is more compatible with the image content of the video frame to be processed, thereby better enhancing the image quality of the video frame to be processed and improving the image enhancement quality of the video frame to be processed.
[0036] In the embodiments of this application, the processor that performs steps 101 and 102 and the processor used to enhance the image quality of the video frame to be processed can be the same processor or different processors, and there is no limitation on this.
[0037] This application embodiment further provides a reference frame selection method. Figure 2 This is a schematic diagram illustrating the implementation flow of another reference frame selection method provided in an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps 201 to 203:
[0038] Step 201: Obtain the E first adjacent frames of the video frame to be processed;
[0039] Step 202: Using a pre-trained reference frame selection model, process the image content of the video frame to be processed and the E first adjacent frames to obtain the first reference frame of the video frame to be processed; wherein, the reference frame selection model is an AI model.
[0040] Step 203: Input the video frame to be processed and the corresponding first reference frame into the pre-trained target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
[0041] The following sections will describe further optional implementation methods for each of the above steps, as well as related terms.
[0042] In step 201, the E first adjacent frames of the video frame to be processed are obtained.
[0043] Whether it is the first adjacent frame here, or the second and third adjacent frames mentioned below, the term "adjacent frame" is a broad concept, referring to video frames within a certain range from the video frame or sample frame to be processed, including video frames at least one moment before the video frame or sample frame to be processed and / or video frames at least one moment after the video frame or sample frame to be processed.
[0044] For example, taking the video frame to be processed as an example, such as Figure 3 As shown, assume X t For the video frame to be processed, video frame X t-1 Up to video frame X t-N (i.e., video frame X) t The first N frames are all video frames X. t The first adjacent frame, video frame X t+1 Up to video frame X t+N (i.e., video frame X) t The last N frames are all video frames X. t The first adjacent frame; where N is greater than 0.
[0045] In step 202, the image content of the video frame to be processed and the E first adjacent frames is processed by a pre-trained reference frame selection model to obtain the first reference frame of the video frame to be processed; wherein, the reference frame selection model is an AI model.
[0046] In the embodiments of this application, the structure of the reference frame selection model is not limited and can be a variety of AI models. For example, it can be linear regression, logistic regression, linear discriminant analysis, decision tree, Bayesian, K-nearest neighbor, learned vector quantization, support vector machine, bagging and random forest or deep neural network, etc.
[0047] Furthermore, in some embodiments, the reference frame selection model is a convolutional neural network, comprising at least: a convolutional layer, a pooling layer, a fully connected layer, and an output layer; wherein, the convolutional layer is used to perform convolution operations on the image content of the video frame to be processed and its E first neighboring frames to obtain a feature map; the pooling layer is used to perform pooling operations on the feature map to obtain a pooled feature map; the fully connected layer is used to determine the probability that the first neighboring frame is selected as the first reference frame based on the pooled feature map; and the output layer is used to select the first reference frame based on the probability that the E first neighboring frames are selected as the first reference frame.
[0048] In some embodiments, the video frame to be processed and the E first adjacent frames can be input into the convolutional layer one by one, and the convolutional layer performs convolution operation on the input image content one by one; in other embodiments, these video frames can be merged before being input into the convolutional layer, that is, the video frame to be processed and the E first adjacent frames are merged along the channel dimension to obtain a merged video frame, and then the merged video frame is input into the convolutional layer to extract the features of the merged video frame using the convolution operation.
[0049] The merging operation refers to the simple splicing or combination of these video frames to obtain a larger image. For example, suppose the images to be merged include image A and image B; where image A is represented as... Image B is represented as After merging these two images, the resulting merged video frame is:
[0050] In some embodiments, the fully connected layer is further configured to determine the probability of each of the F different frame numbers being used as the number of first reference frames based on the pooled feature map; wherein F is greater than 1; correspondingly, the output layer is configured to determine the target number based on the probability of each of the F different frame numbers being used as the number of first reference frames; and to select the first reference frames for the target number based on the probability of each of the E first adjacent frames being selected as the first reference frames.
[0051] Furthermore, in some embodiments, the number of first reference frames with the highest probability or a probability greater than a number threshold can be used as the target number.
[0052] In step 203, the video frame to be processed and the corresponding first reference frame are input into the pre-trained target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
[0053] The process of determining the target image enhancement model is divided into three stages, namely, the first stage, the second stage, and the third stage in sequence. The first and second stages can obtain a trained reference frame selection model, and then the third stage is to retrain the image enhancement model based on the reference frame selection model obtained in the second stage.
[0054] The first stage involves pre-training an initial model of the image enhancement model to obtain a robust and reliable image enhancement model that can achieve good enhancement results for various reference frames. The second stage involves fixing the pre-trained image enhancement model and training the initial model of the reference frame selection model to achieve rapid convergence and obtain a well-trained reference frame selection model. The third stage involves retraining the image enhancement model based on the reference frame selection model obtained in the second stage to obtain a target image enhancement model with better performance.
[0055] The detailed descriptions of the above three stages are as follows:
[0056] Phase 1: Obtaining a trained image augmentation model, including:
[0057] like Figure 4 As shown, based on multiple fourth neighboring frames of the third sample frame and the third sample frame, the model parameters of the second initial model of the image enhancement model 401 are subjected to a second adjustment process to obtain the adjusted second initial model; wherein, the second adjustment process includes: sampling at least one third reference frame from the multiple fourth neighboring frames; inputting the third sample frame and the at least one third reference frame into the second initial model to obtain the third reconstructed frame of the third sample frame; determining the second loss of the third reconstructed frame based at least on the third reconstructed frame and the standard frame of the third sample frame; and adjusting the model parameters of the second initial model based on the second loss;
[0058] Based on multiple fifth neighboring frames of the fourth sample frame and the fourth sample frame, the model parameters of the adjusted second initial model are subjected to a second adjustment process until the corresponding second loss or the number of iterations meets the cutoff condition, thus obtaining image enhancement model 401. In other words, the model parameters of the second initial model are continuously trained through a large number of sample frames and corresponding neighboring frames, and finally, when the training result meets the cutoff condition, the image enhancement model, that is, the trained second initial model, is obtained.
[0059] Understandably, when returning to perform a second adjustment on the model parameters of the second initial model, the data used is new sample data. For example, in the step of "performing a second adjustment on the model parameters of the adjusted second initial model based on multiple fifth adjacent frames of the fourth sample frame and the fourth sample frame", the data used is the fourth sample frame and the multiple fifth adjacent frames, rather than the third sample frame and the multiple fourth adjacent frames.
[0060] The standard frame for the third sample frame mentioned in this article refers to an image frame whose image quality meets the required specifications; for example, the standard frame is a lossless frame. The standard frames for other sample frames mentioned below can be understood by referring to the explanation of the standard frame for the third sample frame here.
[0061] In this embodiment, the sampling method is not limited. The third reference frame can be randomly sampled from the plurality of fourth adjacent frames, or other predetermined sampling strategies can be used. In short, the reference frame is positioned differently relative to the sample frame in multiple iterations; thus, the trained image enhancement model can output a reconstructed frame with good image quality for any input reference frame.
[0062] In this embodiment, the second loss satisfying the cutoff condition includes the second loss being less than the first threshold. The number of iterations satisfying the cutoff condition includes the number of iterations reaching the second threshold.
[0063] For the step of "determining the second loss of the third reconstructed frame based at least on the standard frame of the third sample frame and the third reconstructed frame", in some embodiments, the second loss can be determined based on the difference between the third reconstructed frame and the standard frame of the third sample frame. For example, the second loss L2 can be calculated using the following formula (1):
[0064]
[0065] In equation (1), For the third reconstructed frame, Y t The standard frame is the third sample frame, and ε is set to 1e-6.
[0066] Understandably, after training the second initial model, i.e., obtaining the image enhancement model, this image enhancement model can be used as an evaluator to assess the quality of the reference frames selected by the first initial model of the reference frame selection model, thereby achieving the training of the first initial model. Specifically, refer to the following detailed description of the second stage.
[0067] The second stage: obtaining a trained reference frame selection model, including:
[0068] like Figure 5 As shown, based on multiple second adjacent frames of the first sample frame and the first sample frame, the model parameters of the first initial model of the reference frame selection model 501 are adjusted to obtain the adjusted first initial model.
[0069] The first adjustment process includes: inputting the plurality of second adjacent frames and the first sample frame into the first initial model to obtain a second reference frame of the first sample frame; obtaining a first reconstructed frame of the first sample frame, wherein the first reconstructed frame is obtained by a pre-trained image enhancement model 401 based on the first sample frame and the corresponding second reference frame; determining a first loss of the first reconstructed frame based at least on the first reconstructed frame and the standard frame of the first sample frame; and adjusting the model parameters of the first initial model based on the first loss.
[0070] Based on multiple third adjacent frames of the second sample frame and the second sample frame, the model parameters of the adjusted first initial model are subjected to the first adjustment process until the corresponding first loss or iteration number meets the cutoff condition, thus obtaining the reference frame selection model 501.
[0071] Thus, on the one hand, since the impact of reference frame selection on the enhancement effect of the image enhancement model is considered in the second stage, the trained reference frame selection model can obtain better reference frames, thereby facilitating better enhancement of the image quality of the video frames to be processed. On the other hand, using a pre-trained image enhancement model to train the reference frame selection model allows the training process to converge quickly, thereby saving computational power.
[0072] Understandably, the reference frame selection model has the same structure as the first initial model, the difference being the values of their model parameters. In some embodiments, the reference frame selection model is a convolutional neural network. The step of inputting the plurality of second neighboring frames and the first sample frame into the first initial model to obtain the second reference frame of the first sample frame includes: performing a convolution operation on the image content of the input first sample frame and its second neighboring frames through a convolutional layer to obtain a feature map; then performing a pooling operation on the feature map output by the convolutional layer, and outputting it to a fully connected layer; the fully connected layer determines the probability that each second neighboring frame is selected as the second reference frame based on the pooled feature map; and the output layer selects the second reference frame according to the probability that the plurality of second neighbors are selected as the second reference frame.
[0073] It should be noted that the first reconstructed frame is obtained by the image enhancement model 401 obtained in the first stage based on the input first sample frame and the corresponding second reference frame.
[0074] The "determine the first loss of the first reconstructed frame based at least on the standard frame of the first reconstructed frame and the first sample frame" in the first adjustment process can be implemented through the following embodiment 1 or embodiment 2. Of course, the first loss can also be determined by other methods.
[0075] In Example 1, the method for determining the second loss is the same as that for the first loss, namely, the first loss L1 is determined by the following formula (2):
[0076]
[0077] In equation (2), For the first reconstructed frame, Y t The first sample frame is the standard frame, and ε is set to 1e-6.
[0078] In Embodiment 2, the first loss can also be determined as follows: A first reward for the first reconstructed frame is determined based on the standard frame of the first reconstructed frame and the first sample frame; starting from the first sample frame, M1 consecutive frames before the first sample frame and M2 consecutive frames after the first sample frame are used as the fifth reference frame; the fifth reference frame and the first sample frame are input into the image enhancement model to obtain the second reconstructed frame of the first sample frame; wherein M1 and M2 are greater than 0 and less than or equal to half the number of the plurality of second adjacent frames; a second reward for the second reconstructed frame is determined based on the standard frame of the second reconstructed frame and the first sample frame; the first loss is determined based on the first reward, the second reward, and the probability that the second reference frame is selected as a reference frame; thus, in this embodiment, the calculation of the first loss is based not only on the loss of the first reconstructed frame (i.e., the reconstructed frame calculated by the image enhancement model based on the reference frame output by the first initial model of the reference frame selection model in this embodiment) but also on the loss of the second reconstructed frame (i.e., the reconstructed frame calculated by the image enhancement model based on the reference frame selected by the benchmark method), thus making the reference frame output by the finally trained reference frame selection model superior to the reference frame selected by the benchmark method.
[0079] Furthermore, in some embodiments, the first reward can be determined according to the following formula (3).
[0080]
[0081] In equation (3), the function f(·) is used to calculate the PSNR of the first reconstructed frame, Y t The standard frame is the first sample frame. This is the first reconstructed frame;
[0082] Based on this, in order to maximize the expected reward, the first loss L1 is calculated using the loss function shown in the following formula (4):
[0083]
[0084] In equation (4), K represents the total number of second reference frames. This indicates the probability that the second reference frame is selected as the reference frame. This refers to the first reward. This refers to the second reward.
[0085] The third stage: To obtain a better-performing image enhancement model, it is necessary to retrain the image enhancement model based on the reference frame selection model trained in the second stage. In some embodiments, this includes:
[0086] like Figure 6As shown, based on multiple sixth adjacent frames of the fifth sample frame and the fifth sample frame, the model parameters of the image enhancement model 401 are adjusted in the third way to obtain the adjusted image enhancement model.
[0087] The third adjustment process includes: inputting the plurality of sixth adjacent frames and the fifth sample frame into the reference frame selection model 501 to obtain the fourth reference frame of the fifth sample frame; obtaining the fourth reconstructed frame of the fifth sample frame; the fourth reconstructed frame is obtained by the image enhancement model 401 based on the fifth sample frame and the corresponding fourth reference frame; determining the third loss of the fourth reconstructed frame based at least on the standard frames of the fourth reconstructed frame and the fifth sample frame; and adjusting the model parameters of the image enhancement model 401 based on the third loss.
[0088] Based on the sixth sample frame and its multiple seventh neighboring frames, the model parameters of the adjusted image enhancement model are adjusted until the corresponding third loss or the number of iterations meets the cutoff condition, thus obtaining the target image enhancement model.
[0089] In this way, by using a fixed, well-trained reference frame selection model to retrain the image enhancement model, the performance of the image enhancement model can be further improved, thereby obtaining reconstructed frames with better image quality through the target image enhancement model during online use.
[0090] It should be noted that the method for determining the third loss can be understood by referring to the methods for determining the first or second loss, and will not be described again here.
[0091] In this embodiment, the network structure of the image enhancement model is not limited and can be any video enhancement network. For example, the network structure of the image enhancement model is EDVR.
[0092] This application provides another video enhancement method, including: acquiring E first adjacent frames of a video frame to be processed; wherein E is greater than 1; selecting a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; and enhancing the image quality of the video frame to be processed based on the first reference frame.
[0093] In some embodiments, enhancing the image quality of the video frame to be processed based on the first reference frame includes: inputting the video frame to be processed and the corresponding first reference frame into the target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
[0094] It should be noted that the description of the above video enhancement method embodiments is similar to the description of the above reference frame selection method embodiments, and has similar beneficial effects. For technical details not disclosed in the video enhancement method embodiments of this application, please refer to the description of the reference frame selection method embodiments of this application for understanding.
[0095] The reference frame selection method and video enhancement method provided in this application are applicable to scenarios such as video decompression and video deblurring. In this application, the application scenarios of the method are not limited. In short, they are applicable to any scenario that requires enhancing the image quality of the video frame to be processed based on the reference frame.
[0096] Some video compression algorithms, due to their block-transform-based coding methods, are prone to producing visually unpleasant compression artifacts. Therefore, developing video artifact removal algorithms is essential. Considering the temporal redundancy in videos, video artifact removal algorithms extract spatiotemporal information from reference frames to remove compression distortion from the current frame. For example... Figure 7 As shown, the video decompression distortion algorithm first uses a reference frame selection module to select a reference frame from adjacent frames, and then inputs the reference frame and the current frame into the decompression distortion module to obtain the reconstructed frame. To improve the quality of the reconstructed frame, researchers focus on designing a better decompression distortion module, paying less attention to the design of the reference frame selection module. However, in this embodiment, the focus is on optimizing the reference frame selection module.
[0097] The relevant reference frame selection modules are all designed based on heuristic rules. For example, in Example 1, the two most recent peak quality frames are used as reference frames; in Example 2, the previous frame is used as the reference frame; in Example 3, adjacent preceding and following frames are used as reference frames; and in Example 4, adjacent video frames or I / P frames with lower quantization parameters are used as reference frames.
[0098] However, the reference frame selection module designed based on heuristic rules described above cannot adaptively select reference frames according to the video content and is prone to finding suboptimal reference frames. Specifically, the method described in Example 1 ignores high-quality detail information in low-quality frames. The method described in Example 2 ignores information from subsequent frames. The method described in Example 3 is not necessarily optimal for each frame because the quality fluctuations of adjacent frames are usually different. The method described in Example 4 also ignores high-quality detail information in low-quality frames and is highly dependent on information from the decoding end.
[0099] Based on this, the following will describe an exemplary application of the embodiments of this application in a practical application scenario.
[0100] To address the issues of inability to adaptively select reference frames based on video content and susceptibility to suboptimal solutions in reference frame selection modules, this application provides a video decompression distortion method based on adaptive reference frame selection. This method comprises two modules: an adaptive reference frame selection module (an example of a reference frame selection model) and a decompression distortion module (an example of an image enhancement model). First, the adaptive reference frame selection module selects reference frames based on information from the current frame and its neighboring frames. Then, the decompression distortion module performs decompression distortion operations on the current frame based on the selected reference frames.
[0101] The workflow of the adaptive reference frame selection module is as follows: Figure 8 As shown, this module selects a reference frame based on information from the current frame and its adjacent frames. The specific steps include steps 801 to 805 as follows:
[0102] Step 801: First, given the current video frame, its preceding N video frames, and its following N video frames, for a total of 2N+1 video frames, merge these video frames along the channel dimension.
[0103] Step 802: Use convolution operations to extract features from the merged video frames and output a feature map with smaller spatial resolution and more channels.
[0104] Step 803: Use average pooling to transform the feature map into a 1-dimensional vector;
[0105] Step 804: Use a fully connected layer to transform the 1-dimensional vector into a probability distribution, namely X. t-N To X t-1 and X t+1 To X t+N The probability of each being selected as a reference frame;
[0106] Step 805, based on the probability distribution, select from 2N adjacent frames (i.e., X... t-N To X t-1 and X t+1 To X t+N K frames are selected as reference frames in the )
[0107] The input to the decompression distortion module is the current frame and the selected reference frame, and its output is the current frame with compression distortion removed. Its network structure can be any video enhancement network. In this example, the video enhancement network EDVR is used.
[0108] The training method involved in the embodiments of this application is described below. The training is divided into three stages: in stage one, the decompression distortion module is trained based on a random sampling reference frame selection strategy; in stage two, the decompression distortion module is fixed and the adaptive reference frame selection module is trained; in stage three, the adaptive reference frame selection module is fixed and the decompression distortion module is retrained. The detailed description of the three stages is as follows.
[0109] Phase one involves training a decompression distortion module based on a reference frame selection strategy using random sampling. For example... Figure 9 As shown, firstly, K frames are uniformly sampled from 2N adjacent frames as reference frames. Then, the reference frames and the current frame are input together into the decompression distortion module. Finally, as shown in formula (5), the decompression distortion module is optimized using the Charbonnier loss function:
[0110]
[0111] In equation (5), It is the output of the decompression distortion module, Y t It is a lossless image, with ε set to 1e-6. C, H, and W correspond to the number of channels, height, and width of the output image, respectively.
[0112] Phase two involves fixing the decompression distortion module and training the adaptive reference frame selection module. For example... Figure 9 As shown, a reinforcement learning-based training method is employed, which optimizes the adaptive reference frame selection module based on the quality of the reconstructed image. The definitions of state, action, and reward in this training method are as follows.
[0113] state: It is defined as 2N+1 consecutive frames as input.
[0114] action: From the probability distribution p∈R 2N The sampled value is P, which is the output of the adaptive reference frame selection module and satisfies... P 0:N and P N:2N Corresponding to the previous frame X [t-N:t-1] and subsequent frames X [t-N:t-1] The probability of being selected.
[0115] award: Reflects the state Take action The value of the reconstructed image is shown in formula (6) below, where the quality of the reconstructed image is used as a reward:
[0116]
[0117] In equation (6), f is used to calculate the PSNR of the reconstructed image.
[0118] To maximize the expected reward, the loss function shown in (7) is used:
[0119]
[0120] In equation (7), For adjacent preceding and following frames {X t-K / 2 ...,X t-1 ,X t+1 ,...,X t+K / 2 The reconstruction quality of the reference frame.
[0121] Phase 3: Fix the adaptive reference frame selection module and retrain the decompression distortion module. For example... Figure 9 As shown, after the adaptive reference frame selection module has been trained, the decompression distortion module is retrained based on the reference frame selection strategy it has learned.
[0122] To verify the effectiveness of the adaptive reference frame selection module in the above method, it was compared with three common heuristic reference frame selection methods: Adjacent, MQF, and PQF. The Adjacent method selects adjacent preceding and following frames {X... t-K / 2 ...,X t-1 ,X t+1 ,...,X t+K / 2 As reference frames, the MQF method uses the K-frame with the highest quality among adjacent frames as the reference frame, while the PQF method uses the two most recent peak quality frames as reference frames. The reference frame search radius N in the adaptive reference frame selection module is set to 10.
[0123] The test data consists of 18 test sequences from a publicly available dataset. These test sequences have resolutions ranging from 352×240 to 2560×1680 and were compressed using an HEVC encoder from a certain mobile phone.
[0124] Table 1 below quantitatively compares the advantages and disadvantages of the method provided in this application and the heuristic reference frame selection method. ΔPSNR and ΔSSIM represent the average PSNR and SSIM improvements of the reference frame selection method across 18 test sequences, respectively. As can be seen from Table 1, under different numbers of reference frames, the method provided in this application achieves higher ΔPSNR and ΔSSIM than the heuristic reference frame selection method. This verifies the effectiveness of the method provided in this application.
[0125] Table 1
[0126]
[0127] like Figure 10As shown, it qualitatively compares the above-mentioned method and the heuristic reference frame selection method provided in this application. Figure 10 As shown, compared to heuristic reference frame selection methods, the method provided in this application recovers more detailed information. This is mainly due to the fact that the method provided in this application can provide high-quality reference information, i.e., superior reference frame selection.
[0128] This application provides a video decompression distortion method based on adaptive reference frame selection. Compared to other video decompression distortion methods, the method provided in this application has two major advantages in reference frame selection. First, this method can adaptively select reference frames based on the video content and does not depend on decoding end information. Second, this method uses a data-driven approach to learn reference frame selection, which can find a better solution than heuristic reference frame selection methods. In other words, this application considers the impact of reference frame selection on the enhancement effect, thus finding a better solution than heuristic reference frame selection methods, thereby facilitating better enhancement of the image quality of the video frames to be processed.
[0129] In this embodiment, the EDVR network in the decompression distortion module can also be replaced with other video enhancement networks, and the network structure of the decompression distortion module is not limited.
[0130] The algorithm of the adaptive reference frame selection module is also applicable to tasks such as video compression and video deblurring.
[0131] Based on the aforementioned adaptive reference frame selection module, a reference frame number selection branch is added to adaptively determine the number and position of reference frames according to the spatiotemporal information of the current frame.
[0132] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps; or steps from different embodiments may be combined into a new technical solution.
[0133] Based on the foregoing embodiments, this application provides a reference frame selection device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be an AI acceleration engine (such as NPU), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.
[0134] Figure 11 This is a schematic diagram of the structure of the reference frame selection device provided in the embodiments of this application, as shown below. Figure 11 As shown, the reference frame selection device 110 includes:
[0135] The acquisition module 1101 is configured to acquire the E first adjacent frames of the video frame to be processed;
[0136] Selection module 1102 is configured to select a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; wherein the first reference frame is used to enhance the image quality of the video frame to be processed.
[0137] In some embodiments, the selection module 1102 is configured to: process the image content of the video frame to be processed and the E first adjacent frames using a pre-trained reference frame selection model to obtain a first reference frame of the video frame to be processed; wherein the reference frame selection model is an AI model.
[0138] In some embodiments, the reference frame selection model is a convolutional neural network, comprising at least: a convolutional layer, a pooling layer, a fully connected layer, and an output layer; wherein, the convolutional layer is used to perform a convolution operation on the image content of the input video frame to be processed and the E first adjacent frames to obtain a feature map; the pooling layer is used to perform a pooling operation on the feature map to obtain a pooled feature map; the fully connected layer is used to determine the probability that the first adjacent frame is selected as the first reference frame based on the pooled feature map; and the output layer is used to select the first reference frame based on the probability that the E first adjacent frames are selected as the first reference frame.
[0139] In some embodiments, the fully connected layer is further configured to determine the probability of each of the F different frame numbers being used as the number of first reference frames based on the pooled feature map; correspondingly, the output layer is configured to determine the target number based on the probability of each of the F different frame numbers being used as the number of first reference frames; and to select the first reference frames for the target number based on the probability of each of the E first adjacent frames being selected as the first reference frames.
[0140] In some embodiments, the process of determining the reference frame selection model includes: performing a first adjustment process on the model parameters of a first initial model of the reference frame selection model based on a plurality of second adjacent frames of a first sample frame and the first sample frame, to obtain an adjusted first initial model; wherein, the first adjustment process includes: inputting the plurality of second adjacent frames and the first sample frame into the first initial model to obtain a second reference frame of the first sample frame; obtaining a first reconstructed frame of the first sample frame, wherein the first reconstructed frame is obtained by a pre-trained image enhancement model based on the first sample frame and the corresponding second reference frame; determining a first loss of the first reconstructed frame based at least on the first reconstructed frame and a standard frame of the first sample frame; and adjusting the model parameters of the first initial model based on the first loss;
[0141] Based on multiple third adjacent frames of the second sample frame and the second sample frame, the model parameters of the adjusted first initial model are subjected to the first adjustment process until the corresponding first loss or iteration number meets the cutoff condition, thereby obtaining the reference frame selection model.
[0142] It should be noted that the process of determining the reference frame selection model can be performed by the reference frame selection device 110 or by other devices, and there is no limitation on this.
[0143] In some embodiments, determining the first loss of the first reconstructed frame based at least on the first reconstructed frame and the standard frame of the first sample frame includes: determining a first reward of the first reconstructed frame based on the first reconstructed frame and the standard frame of the first sample frame; taking the first sample frame as a starting point, taking M1 consecutive frames before the first sample frame and M2 consecutive frames after the first sample frame as fifth reference frames, and inputting the fifth reference frames and the first sample frame into the image enhancement model to obtain a second reconstructed frame of the first sample frame; wherein M1 and M2 are greater than 0 and less than or equal to half the number of the plurality of second adjacent frames; determining a second reward of the second reconstructed frame based on the second reconstructed frame and the standard frame of the first sample frame; and determining the first loss based on the first reward, the second reward, and the probability that the second reference frame is selected as a reference frame.
[0144] In some embodiments, the process of determining the image enhancement model includes: performing a second adjustment process on the model parameters of the second initial model of the image enhancement model based on a plurality of fourth neighboring frames of the third sample frame and the third sample frame, to obtain an adjusted second initial model; wherein the second adjustment process includes: sampling at least one third reference frame from the plurality of fourth neighboring frames; inputting the third sample frame and the at least one third reference frame into the second initial model to obtain a third reconstructed frame of the third sample frame; determining a second loss of the third reconstructed frame based at least on the third reconstructed frame and a standard frame of the third sample frame; adjusting the model parameters of the second initial model based on the second loss; and performing the second adjustment process on the model parameters of the adjusted second initial model based on a plurality of fifth neighboring frames of the fourth sample frame and the fourth sample frame, until the corresponding second loss or the number of iterations meets the cutoff condition, to obtain the image enhancement model.
[0145] It should be noted that the process of determining the image enhancement model can be performed by the reference frame selection device 110 or by other devices, and there is no limitation on this.
[0146] In some embodiments, the reference frame selection device 110 further includes an input module, which is configured to input the video frame to be processed and the corresponding first reference frame into a pre-trained target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
[0147] In some embodiments, the process of determining the target image enhancement model includes: performing a third adjustment process on the model parameters of the image enhancement model based on a plurality of sixth neighboring frames of the fifth sample frame and the fifth sample frame, to obtain an adjusted image enhancement model; wherein, the third adjustment process includes: inputting the plurality of sixth neighboring frames and the fifth sample frame into the reference frame selection model to obtain a fourth reference frame of the fifth sample frame; obtaining a fourth reconstructed frame of the fifth sample frame; the fourth reconstructed frame is obtained by the image enhancement model based on the fifth sample frame and the corresponding fourth reference frame; determining a third loss of the fourth reconstructed frame based at least on the fourth reconstructed frame and the standard frame of the fifth sample frame; and adjusting the model parameters of the image enhancement model based on the third loss;
[0148] Based on the sixth sample frame and its multiple seventh neighboring frames, the model parameters of the adjusted image enhancement model are adjusted until the corresponding third loss or the number of iterations meets the cutoff condition, thus obtaining the target image enhancement model.
[0149] It should be noted that the process of determining the target image enhancement model can be performed by the reference frame selection device 110 or by other devices, and there is no limitation on this.
[0150] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0151] This application embodiment further provides a video enhancement device, including: an acquisition module configured to acquire E first adjacent frames of a video frame to be processed; wherein E is greater than 1; a selection module configured to select a first reference frame of the video frame to be processed from the E first adjacent frames based on the image content of the video frame to be processed and the E first adjacent frames; and an enhancement module configured to enhance the image quality of the video frame to be processed based on the first reference frame.
[0152] In some embodiments, the enhancement module is configured to input the video frame to be processed and the corresponding first reference frame into the target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
[0153] It should be noted that the description of the above video enhancement device embodiments is similar to the description of the above reference frame selection method embodiments, and has similar beneficial effects. For technical details not disclosed in the video enhancement device embodiments of this application, please refer to the description of the reference frame selection method embodiments of this application for understanding.
[0154] It should be noted that the module division of the device described in the above embodiments is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0155] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0156] This application provides an electronic device. Figure 12 This is a schematic diagram of the hardware entity of the electronic device according to an embodiment of this application, such as... Figure 12 As shown, the electronic device 120 includes a memory 1201 and a processor 1202. The memory 1201 stores a computer program that can run on the processor 1202. When the processor 1202 executes the program, it implements the steps in the method provided in the above embodiments.
[0157] It should be noted that the memory 1201 is configured to store instructions and applications executable by the processor 1202, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 1202 and the various modules in the electronic device 120. It can be implemented by flash memory or random access memory (RAM).
[0158] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.
[0159] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0160] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0161] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0162] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0163] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0164] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0165] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0166] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0167] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0168] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0169] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0170] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0171] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0172] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A reference frame selection method, characterized in that, The method includes: Obtain the E first adjacent frames of the video frame to be processed; where E is greater than 1; A pre-trained reference frame selection model is used to process the image content of the video frame to be processed and the E first adjacent frames to obtain a first reference frame for the video frame to be processed. The reference frame selection model is an AI model, specifically a convolutional neural network, comprising at least a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The convolutional layer performs convolution operations on the image content of the video frame to be processed and the E first adjacent frames to obtain a feature map. The pooling layer performs pooling operations on the feature map to obtain a pooled feature map. The fully connected layer determines the probability that the first adjacent frame is selected as the first reference frame based on the pooled feature map. The output layer selects the first reference frame based on the probability that the E first adjacent frames are selected as the first reference frame. The first reference frame is used to enhance the image quality of the video frame to be processed.
2. The method according to claim 1, characterized in that, The fully connected layer is also used to determine the probability of using F different frame numbers as the number of first reference frames based on the pooled feature map; where F is greater than 1. Accordingly, the output layer is used to determine the target number based on the probability that the F different number of frames are respectively used as the number of first reference frames; and to select the first reference frames of the target number based on the probability that the E first adjacent frames are selected as the first reference frames.
3. The method according to claim 1 or 2, characterized in that, The process of determining the reference frame selection model includes: Based on multiple second adjacent frames of the first sample frame and the first sample frame, the model parameters of the first initial model of the reference frame selection model are adjusted to obtain the adjusted first initial model. The first adjustment process includes: inputting the plurality of second adjacent frames and the first sample frame into the first initial model to obtain a second reference frame of the first sample frame; obtaining a first reconstructed frame of the first sample frame, wherein the first reconstructed frame is obtained by a pre-trained image enhancement model based on the first sample frame and the corresponding second reference frame; determining a first loss of the first reconstructed frame based at least on the first reconstructed frame and the standard frame of the first sample frame; and adjusting the model parameters of the first initial model based on the first loss. Based on multiple third adjacent frames of the second sample frame and the second sample frame, the model parameters of the adjusted first initial model are subjected to the first adjustment process until the corresponding first loss or iteration number meets the cutoff condition, thereby obtaining the reference frame selection model.
4. The method according to claim 3, characterized in that, The step of determining the first loss of the first reconstructed frame based at least on the first reconstructed frame and the standard frame of the first sample frame includes: Based on the first reconstructed frame and the standard frame of the first sample frame, determine the first reward of the first reconstructed frame; Starting from the first sample frame, the consecutive M1 frames before the first sample frame and the consecutive M2 frames after the first sample frame are used as the fifth reference frame. The fifth reference frame and the first sample frame are input into the image enhancement model to obtain the second reconstructed frame of the first sample frame; wherein, M1 and M2 are greater than 0 and less than or equal to half the number of the plurality of second adjacent frames. The second reward of the second reconstructed frame is determined based on the standard frame of the second reconstructed frame and the first sample frame; The first loss is determined based on the first reward, the second reward, and the probability that the second reference frame is selected as the reference frame.
5. The method according to claim 3, characterized in that, The process of determining the image enhancement model includes: Based on multiple fourth adjacent frames of the third sample frame and the third sample frame, the model parameters of the second initial model of the image enhancement model are adjusted in a second way to obtain the adjusted second initial model. The second adjustment process includes: sampling at least one third reference frame from the plurality of fourth adjacent frames; inputting the third sample frame and the at least one third reference frame into the second initial model to obtain a third reconstructed frame of the third sample frame; determining a second loss of the third reconstructed frame based at least on the standard frame of the third reconstructed frame and the third sample frame; and adjusting the model parameters of the second initial model based on the second loss. Based on the multiple fifth adjacent frames of the fourth sample frame and the fourth sample frame, the model parameters of the adjusted second initial model are subjected to the second adjustment process until the corresponding second loss or the number of iterations meets the cutoff condition, thereby obtaining the image enhancement model.
6. The method according to claim 3, characterized in that, The method further includes: Based on multiple sixth adjacent frames of the fifth sample frame and the fifth sample frame, the model parameters of the image enhancement model are adjusted in a third way to obtain the adjusted image enhancement model. The third adjustment process includes: inputting the plurality of sixth adjacent frames and the fifth sample frame into the reference frame selection model to obtain the fourth reference frame of the fifth sample frame; obtaining the fourth reconstructed frame of the fifth sample frame; the fourth reconstructed frame is obtained by the image enhancement model based on the fifth sample frame and the corresponding fourth reference frame; determining the third loss of the fourth reconstructed frame based at least on the standard frames of the fourth reconstructed frame and the fifth sample frame; and adjusting the model parameters of the image enhancement model based on the third loss. Based on the sixth sample frame and its multiple seventh neighboring frames, the model parameters of the adjusted image enhancement model are adjusted until the corresponding third loss or the number of iterations meets the cutoff condition, thus obtaining the target image enhancement model.
7. The method according to claim 6, characterized in that, The method further includes: The video frame to be processed and the corresponding first reference frame are input into the target image enhancement model to obtain the fifth reconstructed frame of the video frame to be processed.
8. A video enhancement method, characterized in that, The method includes: Obtain the E first adjacent frames of the video frame to be processed; where E is greater than 1; A pre-trained reference frame selection model is used to process the image content of the video frame to be processed and the E first adjacent frames to obtain a first reference frame for the video frame to be processed. The reference frame selection model is an AI model, specifically a convolutional neural network, comprising at least a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The convolutional layer performs convolution operations on the image content of the video frame to be processed and the E first adjacent frames to obtain a feature map. The pooling layer performs pooling operations on the feature map to obtain a pooled feature map. The fully connected layer determines the probability that the first adjacent frame is selected as the first reference frame based on the pooled feature map. The output layer selects the first reference frame based on the probability that the E first adjacent frames are selected as the first reference frame. Based on the first reference frame, the image quality of the video frame to be processed is enhanced.
9. A reference frame selection device, characterized in that, The device includes: The acquisition module is configured to acquire E first adjacent frames of the video frame to be processed; where E is greater than 1. The selection module is configured to process the image content of the video frame to be processed and the E first adjacent frames using a pre-trained reference frame selection model to obtain a first reference frame for the video frame to be processed. The reference frame selection model is an AI model, specifically a convolutional neural network, comprising at least a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The convolutional layer performs convolution operations on the image content of the video frame to be processed and the E first adjacent frames to obtain a feature map. The pooling layer performs pooling operations on the feature map to obtain a pooled feature map. The fully connected layer determines the probability that the first adjacent frame is selected as the first reference frame based on the pooled feature map. The output layer selects the first reference frame based on the probability that the E first adjacent frames are selected as the first reference frame. The first reference frame is used to enhance the image quality of the video frame to be processed.
10. A video enhancement device, characterized in that, The device includes: The acquisition module is configured to acquire E first adjacent frames of the video frame to be processed; where E is greater than 1. The selection module is configured to process the image content of the video frame to be processed and the E first adjacent frames using a pre-trained reference frame selection model to obtain a first reference frame for the video frame to be processed. The reference frame selection model is an AI model, specifically a convolutional neural network, comprising at least a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The convolutional layer performs convolution operations on the image content of the video frame to be processed and the E first adjacent frames to obtain a feature map. The pooling layer performs pooling operations on the feature map to obtain a pooled feature map. The fully connected layer determines the probability that the first adjacent frame is selected as the first reference frame based on the pooled feature map. The output layer selects the first reference frame based on the probability that the E first adjacent frames are selected as the first reference frame. The enhancement module is configured to enhance the image quality of the video frame to be processed based on the first reference frame.
11. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 7, or when the processor executes the program, it implements the method according to claim 8.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7, or when the computer program is executed by a processor, it implements the method as described in claim 8.
Citation Information
Patent Citations
Video denoising method and device, electronic equipment and computer readable storage medium
CN113556442A