Method and apparatus for implementing super-resolution of image frames
By utilizing trained super-resolution networks and reference information, the method addresses frame quality variations in video streaming, enhancing resolution and efficiency in video frame processing.
Patent Information
- Application Number
- CN202010664759.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-07-10
AI Technical Summary
In the video streaming media service, the frame sequences undergo different degrees of compression, resulting in large differences in the quality of adjacent frames, affecting the super-resolution effect, and high terminal processing resource consumption.
The transmitting end obtains the super-score reference information, including quantization parameters and image quality scores, selects M high-quality image frames to input the video super-resolution network for processing, and the terminal selects the corresponding network for super-resolution processing based on the quantization parameters, saving terminal resources and improving processing efficiency.
It improves the super-resolution processing effect of image frames, saves terminal processing resources, and improves processing efficiency.
Smart Images

Figure CN113920010B_ABST
Abstract
Description
Technical Field
[0001] This application relates to video super-resolution technology, and in particular to a method and apparatus for implementing super-resolution of image frames. Background Art
[0002] With the rise of video application software, people's demand for high-definition resolution videos, and even ultra-high definition (UHD) videos, is increasing day by day. In multimedia services such as online video playback and live streaming based on information transmission between a sending end and a mobile terminal, how to improve the encoding efficiency, reduce the target bit rate, and enhance the video playback effect has become a challenge.
[0003] In related technologies, one frame before and one frame after or two frames before and two frames after the current frame are selected, and three or five consecutive frames are grouped into an image set and input into a deep learning network, and finally a high-resolution image of the current frame is output.
[0004] However, when the above method is applied to an actual streaming media service, the frame sequence may be compressed to varying degrees, resulting in a large quality difference between adjacent frames, thereby affecting the super-resolution effect of the current frame. Summary of the Invention
[0005] This application provides a method and apparatus for implementing super-resolution of image frames to save the processing resources of a terminal, improve the processing efficiency of the terminal, and enhance the effect of super-resolution processing of image frames.
[0006] In a first aspect, this application provides a method for implementing super-resolution of image frames, including: obtaining super-resolution reference information, where the super-resolution reference information includes a quantization parameter and a set of image quality scores, and the set of image quality scores includes the image quality scores of multiple image frames; selecting M image frames from the multiple image frames according to the set of image quality scores, where M is greater than or equal to 1; obtaining a video super-resolution network corresponding to the quantization parameter, where the video super-resolution network has a super-resolution function; inputting the M image frames and a first image frame into the video super-resolution network, and the video super-resolution network is configured to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
[0007] The sending end has the encoding function of the video stream, and encodes each image frame in the video stream respectively. Usually, this encoding process may include image block division, mode selection, residual calculation, transformation, quantization, inverse quantization, inverse transformation, reconstruction, and filtering. The processes related to this application include mode selection, quantization, and reconstruction.
[0008] Mode selection is to select a segmentation and prediction mode for an image block. Segmentation can divide (or partition) an image frame into smaller parts, such as image blocks in the shape of a square or rectangle. The prediction mode can provide the best match or the minimum residual (the minimum residual means better compression in transmission or storage), or provide the minimum signaling overhead (the minimum signaling overhead means better compression in transmission or storage), or consider or balance both of the above at the same time. The prediction mode can include, for example, an intra prediction mode and / or an inter prediction mode. Among them, the intra prediction mode is used to generate an intra prediction block by using the reconstructed pixel points of adjacent blocks within the same current image frame and output intra prediction parameters. The inter prediction mode is used to select the reconstructed pixel points of a reference block from multiple reference blocks in multiple other image frames to generate an inter prediction block and output inter prediction parameters. The transmitting end obtains the reference frame indication information of the image frame from the mode selection, and the reference frame indication information is used to indicate the above-mentioned intra adjacent blocks or reference blocks.
[0009] Quantization quantizes the transform residual coefficients through, for example, scalar quantization or vector quantization to obtain quantized residual coefficients. The quantization process can reduce the bit depth related to some or all of the transform residual coefficients. For example, during quantization, the n-bit transform residual coefficients can be rounded down to m-bit transform residual coefficients, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, the appropriate quantization step size can be indicated by the QP. A smaller quantization parameter can correspond to fine quantization (a smaller quantization step size), and a larger quantization parameter can correspond to coarse quantization (a larger quantization step size), and vice versa. Quantization is a lossy operation, and the larger the quantization step size, the greater the loss. In a possible implementation, the video encoder can be used to output the QP so that the video decoder can receive and use the QP for decoding. The transmitting end obtains the quantization parameter of the image frame from the quantization.
[0010] Reconstruction is a process of inverse quantizing and inverse transforming the quantized residual coefficients (obtained after the residual block is transformed and quantized) to obtain a reconstructed residual block, and then adding the reconstructed residual block to the prediction block to obtain a reconstructed block in the pixel domain. The image frame obtained by the transmitting end during the reconstruction process is the reconstructed frame, and then the image quality score of the image frame can be obtained based on the difference between the reconstructed frame and the original image frame. For example, the transmitting end can obtain the image quality score of the first image frame according to the peak signal to noise ratio (PSNR), the structural similarity (SSIM), or the video multimethod assessment fusion (VMAF) of multi-feature fusion.
[0011] PSNR can be calculated using the following formula:
[0012]
[0013] Among them, x(i, j) represents the pixel value of the i-th row and j-th column in the first image frame, x′(i, j) represents the pixel value of the i-th row and j-th column in the reconstructed frame of the first image frame, and the resolution of the first image is M×N. In this application, the calculated PSNR can be used as the image quality score of the first image frame.
[0014] The calculation formula of SSIM is based on three comparison metrics between the first image frame x and the reconstructed frame y of the first image frame: luminance l(x, y), contrast c(x, y), and structure s(x, y):
[0015]
[0016]
[0017]
[0018] Among them, c3 = c2 / 2, μ x represents the mean difference of x, μ y represents the mean difference of y, represents the variance of x, represents the variance of y, σ xy represents the covariance of x and y, c1 = (k1L) 2 , c2 = (k2L) 2 , L represents the range of pixel values, k1 = 0.01, k2 = 0.03.
[0019] SSIM(x, y) = [l(x, y) α ·c(x, y) β ·s(x, y) γ
[0020] Assuming that α, β, and γ are all 1, we can get:
[0021]
[0022] For each calculation, an N×N window can be selected from the image frame, then the window is slid for calculation, and finally the average value is taken as the SSIM of the entire image frame. In this application, the calculated SSIM can be used as the image quality score of the first image frame.
[0023] VMAF is a model trained by machine learning, which can score the input image frame, that is, input the image frame to be scored, and the model directly outputs the image quality score of this image frame.
[0024] In this application, the super-resolution reference information is obtained by the sender and transmitted to the terminal. The terminal can obtain M image frames based on this super-resolution reference information, and input the M image frames and the first image frame to be processed into the video super-resolution network together to achieve super-resolution processing of the first image frame. Based on this, the sender can include the above quantization parameters, image quality scores, and reference frame indication information in the super-resolution reference information, and all three pieces of information are used by the terminal to select the above M image frames.
[0025] In addition to the function of encoding the video stream, the sender also has the ability to train neural networks, that is, the video super-resolution network used in this application is trained by the sender. Before the terminal executes the super-resolution implementation method of the image frames provided in this application, it has downloaded the trained video super-resolution network from the sender. The sender can train different video super-resolution networks for one or more quantization parameters, so that there is a corresponding relationship between the quantization parameters and the video super-resolution networks. This corresponding relationship can be one-to-one (one quantization parameter corresponds to one video super-resolution network), or many-to-one (multiple quantization parameters correspond to one video super-resolution network). After obtaining a quantization parameter in the super-resolution reference information, the terminal can obtain the corresponding video super-resolution network according to this quantization parameter.
[0026] Super-resolution is a process of obtaining a high-resolution image frame from multiple low-resolution image frames, which can improve the resolution of the original image. Resolution refers to the amount of information stored in an image frame, which is used to represent how many pixels are included per inch of the image. The unit of resolution is pixels per inch (PPI). The higher the resolution, the more pixels are included per inch. Super-resolution processing is to increase the number of pixels included per inch, making the details of the image frame become rich and improving its clarity. This application can use a trained neural network to implement the super-resolution function.
[0027] Optionally, this application can also use image processing algorithms to implement the super-resolution function, and other methods can also be used to implement the super-resolution function, which is not specifically limited herein.
[0028] The terminal of the present application selects a video super-resolution network to be used and reference image frames participating in super-resolution processing by means of super-resolution reference information from the sending end, so as to perform super-resolution processing on the first image frame (i.e., the image frame to be super-resolved) and improve the resolution of the first image frame. On the one hand, the terminal selects a corresponding video super-resolution network according to the quantization parameters of the first image frame, which can better improve the resolution. On the other hand, the reference image frames are selected by using the scores given by the sending end to each image frame to achieve the purpose of maximizing the quality after multi-frame fusion super-resolution, and improve the effect of super-resolution processing of the image frame. On the third hand, the resources of the sending end are used to score the image frames in the video stream, which can save the processing resources of the terminal, reduce the computing amount of the terminal, and improve the efficiency of its super-resolution processing.
[0029] In a possible implementation manner, the selecting M image frames from the multiple image frames according to the set of image quality scores specifically includes: when the multiple image frames include a first set of image frames, and the number of image frames in the first set of image frames is greater than or equal to M, selecting the M image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame.
[0030] M is the number of reference image frames required for super-resolution processing by the video super-resolution network, and it is related to the training data used by the sending end to train the video super-resolution network. The present application does not make specific limitations on this. Selecting M image frames from multiple image frames whose image quality scores are known, and the image quality scores of the M image frames are all higher than the image quality score of the first image frame, which meets the requirements of super-resolution processing, that is, using high-score image frames to super-resolve the image frame to be super-resolved.
[0031] In a possible implementation manner, the M image frames include the first M image frames in the first set of image frames arranged in descending order of image quality score.
[0032] In addition to the above conditions, the higher the scores of the selected M image frames, the more beneficial it is to improve the super-resolution effect. Therefore, further, the M image frames with the highest scores can be selected from multiple image frames with scores higher than the first image frame.
[0033] In a possible implementation, the super-resolution reference information further includes reference frame indication information of the first image frame; the step of selecting M image frames from the multiple image frames according to the set of image quality scores further includes: when the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is less than M, selecting all image frames from the first set of image frames, where the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; when the number of all image frames is less than M, selecting the image frames corresponding to the reference frame indication information from the multiple image frames; when the number of all image frames and the image frames corresponding to the reference frame indication information is less than M, selecting at least one image frame with an increasing time interval from the first image frame from the other image frames in the multiple image frames except the all image frames and the image frames corresponding to the reference frame indication information until M image frames are selected.
[0034] If the number of all image frames with an image quality score higher than that of the first image frame is less than M, but the video super-resolution network must input M image frames to complete the super-resolution processing of the first image frame. At this time, in addition to selecting all image frames with an image quality score higher than that of the first image frame, other methods can be used to obtain image frames to make up M image frames. The reference frame of the first image frame is an optional object, which is the reference frame (intra-frame reference frame or inter-frame reference frame) selected by the sender when selecting the mode of the first image frame. Since this reference frame can be used to predict the first image frame during encoding, it can also be used as a reference image frame when performing super-resolution processing on the first image frame. If the number of all image frames with an image quality score higher than that of the first image frame and the reference frame of the first image frame still cannot reach M image frames, then image frames can be selected from the multiple image frames in the order of increasing time interval from the first image frame, that is, starting from the first image frame, selecting image frames frame by frame forward and backward. For example, if the serial number of the first image frame is n and it is assumed that 4 more image frames need to be selected, then the four image frames can be n-2, n-1, n+1, n+2. If any of these four image frames has already been selected, then the next image frame will be selected. For example, if n+1 and n-1 have already been selected, then n-3 and n+3 can be selected; or, if n-1 and n-2 have already been selected, then n-3 and n+3 can be selected.
[0035] It should be noted that there is a possibility that when selecting all the above-mentioned image frames with an image quality score higher than that of the first image frame, the reference frame of the first image frame is already included. In this case, image frames can be directly selected from the multiple image frames in the order of increasing time interval from the first image frame.
[0036] In a possible implementation, the super-resolution reference information further includes reference frame indication information of the first image frame; the step of selecting M image frames from the multiple image frames according to the set of image quality scores further includes: when the multiple image frames do not include the first image frame set, determining whether the first image frame is an I frame, where the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame; if the first image frame is an I frame, the M image frames include M copy samples of the first image frame; if the first image frame is not an I frame, selecting the image frame corresponding to the reference frame indication information from the multiple image frames; when the number of image frames corresponding to the reference frame indication information is less than M, selecting at least one image frame with an increasing time interval from the first image frame from the other image frames in the multiple image frames except the image frames corresponding to the reference frame indication information until M image frames are selected.
[0037] During video coding, according to the prediction type, image frames can be divided into I frames, P frames, and B frames. Among them, an I frame is an image frame encoded as an independent static image, providing a random access point in the video stream; a P frame is an image frame predicted from its adjacent previous I frame or P frame as a reference frame, and it can be used as a reference frame for subsequent P frames or B frames; a B frame is an image frame obtained by bidirectional prediction using the two nearest adjacent frames (which can be I frames or P frames) as reference frames.
[0038] If the first image frame is an I frame, then two situations may occur: one is that a scene switch has occurred, and the first image frame is the start of another scene, so the content of the previous scene is different from the current scene, and the information of the image frames in the previous scene cannot be utilized; the other is that the first image frame is an I frame inserted at a fixed interval, and usually, the quantization parameter of the first image frame is lower than that of the surrounding image frames, and its image quality may be much higher than that of the surrounding image frames. Therefore, the information of the surrounding image frames cannot be utilized either. Based on these two situations, if the first image frame is an I frame, it is very likely that no reference image frame that can be used for super-resolution processing can be found from the surrounding image frames. In order to meet the input requirements of the video super-resolution network in this application, the first image frame can be copied multiple times to obtain M copy samples. The information and data of each copy sample are the same as those of the first image frame. Then, the M copy samples are input into the video super-resolution network, which is equivalent to inputting the information and data of M image frames, and the information and data of these M image frames are the same.
[0039] If the first image frame is not an I-frame and there is no image frame with a higher image quality score among its surrounding image frames, other methods can be used to obtain image frames to make up the M image frames. As described above, the reference frame of the first image frame can be selected. If there are still not enough M image frames at this time, image frames can also be selected from multiple image frames in the order of increasing time interval from the first image frame, that is, starting from the first image frame, select image frames frame by frame forward and backward. For example, if the serial number of the first image frame is n and it is assumed that 4 more image frames need to be selected, then these four image frames can be n-2, n-1, n+1, n+2. If any of these four image frames has already been selected, then continue to select the next image frame. For example, if n+1 and n-1 have already been selected, then n-3 and n+3 can be selected; or, if n-1 and n-2 have already been selected, then n-3 and n+3 can be selected.
[0040] In a possible implementation manner, the obtaining of the super-resolution reference information specifically includes: receiving the super-resolution reference information sent by the sending end.
[0041] After the sending end encodes the video stream, a bitstream of the video stream is obtained. In one case, when encoding, the sending end encodes the data of the image frames in the video stream and the super-resolution reference information, such as quantization parameters, image quality scores, and reference frame indication information, together to obtain a bitstream containing the data of the image frames and the super-resolution reference information, and transmits this bitstream to the terminal; in another case, the sending end encodes the data of the image frames in the video stream to obtain a bitstream of the image frames, then encodes the super-resolution reference information to obtain a bitstream of the super-resolution reference information, and transmits the bitstream of the image frames and the bitstream of the super-resolution reference information after splicing to the terminal. The present application does not make specific limitations on the implementation manner of the bitstream.
[0042] In a possible implementation manner, the multiple image frames include consecutive multiple image frames in the video stream, and the multiple image frames include the first image frame.
[0043] The fact that the alternative multiple image frames include the first image frame means that these multiple image frames are before and / or after the first image frame in the video stream. The shorter the time interval from the first image frame, the greater the correlation between the image frames. In the case of slow motion or even a static image, it is possible that the two consecutive image frames are almost exactly the same. Therefore, when selecting M image frames, the range of the multiple image frames can be limited first. For example, the first image frame is located in the middle position (exactly in the center or within a certain set range in the middle) of the multiple image frames. Then, no matter which method is used to select M image frames, it will not exceed this range. This also avoids the situation where the selected image frame may have a very high image quality score but is far apart from the first image frame in terms of time, losing its reference value and resulting in poor super-resolution effects.
[0044] In a possible implementation, the video super-resolution network includes a convolutional neural network (CNN), a deep neural network (DNN), or a recurrent neural network (RNN).
[0045] In a possible implementation, the video super-resolution network includes a convolutional layer and an activation layer.
[0046] In a possible implementation, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; and the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0047] In a second aspect, the present application provides a method for implementing super-resolution of an image frame, including: obtaining quantization parameters of a first image frame; obtaining an image quality score of the first image frame according to the first image frame and a reconstructed frame of the first image frame; and sending a video stream and super-resolution reference information to a terminal, where the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameters and the image quality score.
[0048] For the obtaining of the quantization parameters and the image quality score, reference may be made to the first aspect above, which will not be elaborated herein.
[0049] During the encoding process of the video stream by the transmitting end of the present application, the quantization parameters and the image quality score of the first image frame in the video stream are obtained and added to the super-resolution reference information and sent to the terminal, so as to implement super-resolution processing of the first image frame (i.e., the image frame to be super-resolved) and improve the resolution of the first image frame. On the one hand, the video super-resolution network corresponds to the quantization parameters, which can better improve the resolution. On the other hand, the transmitting end scores each image frame and uses this as the basis for selecting the reference image frame to achieve the purpose of maximizing the quality after multi-frame fusion super-resolution and improving the effect of super-resolution processing of the image frame. In a third aspect, by using the resources of the transmitting end to score the image frames in the video stream, the processing resources of the terminal can be saved, the computing amount of the terminal can be reduced, and the efficiency of its super-resolution processing can be improved.
[0050] In a possible implementation, before sending the video stream and the super-resolution reference information to the terminal, it further includes: obtaining reference frame indication information of the first image frame; correspondingly, the super-resolution reference information further includes the reference frame indication information.
[0051] In a possible implementation, the obtaining of the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame specifically includes: obtaining the image quality score of the first image frame according to the peak signal-to-noise ratio (PSNR), the structural similarity (SSIM), or the video multi-feature fusion evaluation criterion (VMAF).
[0052] In a possible implementation, the video super-resolution network includes a convolutional neural network (CNN), a deep neural network (DNN), or a recurrent neural network (RNN).
[0053] In a possible implementation, the video super-resolution network includes a convolutional layer and an activation layer.
[0054] In a possible implementation, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; and the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0055] In a possible implementation, the method further includes: obtaining a training data set, where the training data set includes first-resolution images and second-resolution images of multiple image frames respectively, and multiple quantization parameters, and the resolution of the first-resolution images is higher than that of the second-resolution images; training multiple video super-resolution networks according to the training data set, and the multiple video super-resolution networks correspond to the multiple quantization parameters.
[0056] When the transmitting end trains the video super-resolution network, it collects a training data set, which includes high-resolution images and low-resolution images (which may include multiple low-resolution images) of multiple image frames, and quantization parameters of multiple image frames. The high-resolution images of the multiple image frames are all obtained by the same downsampling method and compression method to obtain the corresponding low-resolution images. According to the high-resolution images and low-resolution images of the multiple image frames, based on the corresponding relationship between the high-resolution images and the low-resolution images, the training engine can learn the rules of how to process the high-resolution images into low-resolution images, and the corresponding relationship between the rules and the quantization parameters, and then form multiple video super-resolution networks respectively corresponding to different quantization parameters.
[0057] In a third aspect, the present application provides a method for implementing super-resolution of an image frame, including: the sending end obtains quantization parameters of a first image frame, and obtains an image quality score of the first image frame according to the first image frame and a reconstructed frame of the first image frame; the sending end sends a video stream and super-resolution reference information to a terminal, the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameters and the image quality score; the terminal obtains the first image frame according to the video stream, and obtains the quantization parameters and an image quality score set according to the super-resolution reference information, the image quality score set includes image quality scores of multiple image frames respectively; the terminal selects M image frames from the multiple image frames according to the image quality score set, M is greater than or equal to 1; the terminal obtains a video super-resolution network corresponding to the quantization parameters, the video super-resolution network has a super-resolution function; the terminal inputs the M image frames and the first image frame into the video super-resolution network, and the video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
[0058] The terminal of the present application uses the super-resolution reference information from the sending end to select the video super-resolution network to be used and the reference image frames participating in the super-resolution processing, so as to implement the super-resolution processing of the first image frame (i.e., the image frame to be super-resolved) and improve the resolution of the first image frame. On the one hand, the terminal selects the corresponding video super-resolution network according to the quantization parameters of the first image frame, which can better improve the resolution. On the other hand, the reference image frames are selected by using the scores given by the sending end to each image frame to achieve the purpose of maximizing the quality after multi-frame fusion super-resolution and improve the effect of the super-resolution processing of the image frame. On the third hand, using the resources of the sending end to score the image frames in the video stream can save the processing resources of the terminal, reduce the computing amount of the terminal, and improve the efficiency of its super-resolution processing.
[0059] In a fourth aspect, the present application provides a terminal device, including: an acquisition module, configured to acquire super-resolution reference information, the super-resolution reference information includes quantization parameters and an image quality score set, the image quality score set includes image quality scores of multiple image frames respectively; select M image frames from the multiple image frames according to the image quality score set, M is greater than or equal to 1; acquire a video super-resolution network corresponding to the quantization parameters, the video super-resolution network has a super-resolution function; a processing module, configured to input the M image frames and a first image frame into the video super-resolution network, and the video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
[0060] In a possible implementation, the obtaining module is specifically configured to, when the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is greater than or equal to M, select the M image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame.
[0061] In a possible implementation, the M image frames include the first M image frames in the first set of image frames arranged in descending order of image quality score.
[0062] In a possible implementation, the super-resolution reference information further includes reference frame indication information of the first image frame; the obtaining module is further configured to, when the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is less than M, select all the image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; when the number of all the image frames is less than M, select the image frames corresponding to the reference frame indication information from the multiple image frames; when the number of all the image frames and the image frames corresponding to the reference frame indication information is less than M, select at least one image frame with an increasing interval time from the first image frame from the other image frames in the multiple image frames except the all the image frames and the image frames corresponding to the reference frame indication information until M image frames are selected.
[0063] In a possible implementation, the super-resolution reference information further includes reference frame indication information of the first image frame; the obtaining module is further configured to, when the multiple image frames do not include a first set of image frames, determine whether the first image frame is an I-frame, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; if the first image frame is an I-frame, the M image frames include M replicated samples of the first image frame; if the first image frame is not an I-frame, select the image frames corresponding to the reference frame indication information from the multiple image frames; when the number of the image frames corresponding to the reference frame indication information is less than M, select at least one image frame with an increasing interval time from the first image frame from the other image frames in the multiple image frames except the image frames corresponding to the reference frame indication information until M image frames are selected.
[0064] In a possible implementation, the obtaining module is specifically configured to receive the super-resolution reference information sent by the sending end.
[0065] In a possible implementation, the quantization parameter includes the quantization parameter used by the transmitting end during the quantization process of the first image frame.
[0066] In a possible implementation, the multiple image frames include a continuous plurality of image frames in the video stream, and the multiple image frames include the first image frame.
[0067] In a possible implementation, the video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
[0068] In a possible implementation, the video super-resolution network includes a convolutional layer and an activation layer.
[0069] In a possible implementation, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0070] In a fifth aspect, the present application provides a transmitting end device, including: an acquisition module, configured to acquire the quantization parameter of the first image frame; obtain the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame; a sending module, configured to send a video stream and super-resolution reference information to a terminal device, where the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameter and the image quality score.
[0071] In a possible implementation, the acquisition module is further configured to acquire the reference frame indication information of the first image frame; correspondingly, the super-resolution reference information further includes the reference frame indication information.
[0072] In a possible implementation, the acquisition module is specifically configured to obtain the image quality score of the first image frame according to the peak signal-to-noise ratio PSNR, the structural similarity SSIM, or the video evaluation criterion VMAF of multi-feature fusion.
[0073] In a possible implementation, the video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
[0074] In a possible implementation, the video super-resolution network includes a convolutional layer and an activation layer.
[0075] In a possible implementation, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0076] In a possible implementation, it further includes: a training module; the obtaining module is further configured to obtain a training data set, where the training data set includes the first-resolution images and the second-resolution images of multiple image frames respectively, as well as multiple quantization parameters, and the resolution of the first-resolution images is higher than that of the second-resolution images; the training module is configured to train multiple video super-resolution networks according to the training data set, and the multiple video super-resolution networks correspond to the multiple quantization parameters.
[0077] In a sixth aspect, the present application provides an image processing system, including: a sending-end device and a terminal device, where the sending-end device adopts the device described in any one of the above fifth aspects, and the terminal device adopts the device described in any one of the above fourth aspects.
[0078] In a seventh aspect, the present application provides a terminal, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the above first aspects.
[0079] In an eighth aspect, the present application provides a sending end, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the above second aspects.
[0080] In a ninth aspect, the present application provides a computer-readable storage medium, including a computer program, when the computer program is executed on a computer, the computer executes the method described in any one of the above first to third aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 FIG. shows an exemplary schematic diagram of an architecture applicable to the method for implementing super-resolution of image frames in the present application;
[0082] Figures 2a - 2e FIG. shows an exemplary architecture of a video super-resolution network;
[0083] Figure 3 FIG. shows a schematic structural diagram of the terminal 300;
[0084] Figure 4 FIG. is a flowchart of an embodiment of the method for implementing super-resolution of image frames in the present application;
[0085] Figures 5a - 5c FIG. shows an exemplary schematic diagram of the process of playing a video online;
[0086] Figures 6a - 6c An exemplary schematic diagram showing the switching process of video resolution;
[0087] Figure 7 A schematic structural diagram of an embodiment of the terminal device of the present application;
[0088] Figure 8 A schematic structural diagram of an embodiment of the sending - end device of the present application. Detailed implementation manners
[0089] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0090] The terms "first", "second", etc. in the description, claims and drawings of the present application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, including a series of steps or units, a method, a system, a product or a device does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0091] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a, b and c", where a, b, c can be single or multiple.
[0092] Figure 1 An exemplary schematic diagram showing an architecture applicable to the super - resolution implementation method of the image frame of the present application, as Figure 1As shown in the figure, the application framework includes a sending end and a terminal. The sending end can be, for example, the cloud, or other devices or servers with image encoding functions. The sending end and the terminal can be connected through a wireless communication network. The sending end includes a downsampling module and an encoding module. The downsampling module is used to perform downsampling processing on each image frame in the input source video stream. Methods such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation can be used. The downsampling scale is one-half or one-fourth, etc. This application does not make specific limitations on this. The encoding module is used to encode the data of the downsampled image frames to obtain a video bitstream. The terminal includes a decoding module and a super-resolution module. The decoding module is used to decode the received video bitstream to obtain the data of the image frames, and then reconstruct the video stream composed of the reconstructed image frame sequences according to the data of the image frames. The types of encoding / decoding modules used in this application include but are not limited to H.264 / 265. The specific encoding / decoding methods and parameters are not specifically limited.
[0093] The above-mentioned sending end further includes a training engine, which is used to train a video super-resolution network to perform super-resolution processing on the reconstructed image frames.
[0094] The training data in this application includes: a training data set, which includes high-resolution images and low-resolution images (which can include multiple low-resolution images) of multiple image frames, as well as quantization parameters of multiple image frames. The high-resolution images of the multiple image frames all use the same downsampling method and compression method to obtain the corresponding low-resolution images. Based on the corresponding relationship between the high-resolution images and the low-resolution images of the multiple image frames, the training engine can learn the rules on how to process high-resolution images into low-resolution images, and the corresponding relationship between these rules and the quantization parameters, and then form multiple video super-resolution networks corresponding to different quantization parameters respectively.
[0095] Since the resolution of the video will go through two degradation processes of downsampling and encoding during the process of being processed into a bitstream, where the degradation degree of downsampling is the same for all image frames in the video stream, but the degradation degree of encoding distortion will be different according to different quantization parameters. Therefore, the training engine at the sending end will train different video super-resolution networks according to different quantization parameters when training the video super-resolution network, so that the terminal can better perform video super-resolution recovery. Thus, when the terminal performs super-resolution processing on the image frames according to the video super-resolution network, it needs to determine the corresponding video super-resolution network according to the QP used for quantization processing of the image frames.
[0096] The above training data can be stored in a database (not shown), and the training engine trains a neural network based on the training data, such as a video super-resolution network. It should be noted that the embodiments of the present application do not limit the source of the training data. For example, the training data can be obtained from the sending end or other places for training.
[0097] The video super-resolution network can be used to implement the super-resolution implementation method of the image frame provided by the embodiments of the present application, that is, the terminal inputs the reconstructed image frame into the video super-resolution network based on the relevant information of the sending end, and a high-resolution image frame can be obtained. The following will be combined with Figures 2a - 2e to elaborate on the video super-resolution network in detail.
[0098] The video super-resolution network trained by the training engine can be applied to Figure 1 the application framework shown, especially the sending end. The training engine can train the video super-resolution network at the sending end, and then the terminal downloads and uses the video super-resolution network from the sending end. For example, the training engine trains the video super-resolution network, the terminal downloads the video super-resolution network from the sending end, and then can perform super-resolution processing on the input reconstructed image frame according to the video super-resolution network to obtain a high-resolution image frame.
[0099] The above sending end can be a server, such as a streaming media server, a video website server, etc.
[0100] The above terminal can be, for example, a mobile phone, a tablet, etc., or a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, etc. The present application does not limit this.
[0101] It should be noted that the sending end and the terminal can be independent devices, each implementing corresponding functions. Optionally, the sending end and the terminal can also be regarded as a whole and interact with each other to implement corresponding functions. Among them, the sending end can adopt Figure 8 the device shown, and the terminal can adopt Figure 7 the device shown.
[0102] In a possible implementation manner, Figure 1 the framework shown can be applied to the following application scenarios:
[0103] 1. The sending end provides a video library. When a user opens a video playback software on a terminal and selects a video to play, the terminal sends a playback request to the sending end, and the request carries the identification information of the video. The sending end device obtains the corresponding video according to the identification information, encodes and compresses it, and then sends it to the terminal. To increase the bit rate, the sending end may perform downsampling on the video stream. When the user plays such a video stream and feels that the picture is not clear, the user can select a higher resolution through the control provided by the video playback software. After receiving the instruction to change the resolution, the terminal uses the method provided in this application to perform super-resolution processing on each image frame in the video stream to increase the resolution of the video.
[0104] 2. When the terminal plays an online video and finds that the video can be played at a higher resolution, the terminal actively uses the method provided in this application to perform super-resolution processing on each image frame in the video stream to increase the resolution of the video.
[0105] 3. The video comes from other devices. For example, the video is stored on other computers or terminals, and the video playback software provides a function of playing while receiving, that is, when the terminal receives a sufficient amount of video stream data, the video playback software can start playing the video. At this time, the resolution of the video can be increased by the user's selection or the terminal's active trigger, and the terminal uses the method provided in this application to perform super-resolution processing on each image frame in the video stream to increase the resolution of the video.
[0106] It should be noted that the above examples describe several scenarios in which the super-resolution implementation method of the image frames provided in this application can be applied, but this does not limit the application scenarios of this application. As long as there are scenarios with requirements for video encoding and decoding, video transmission, and video playback, the method provided in this application can be used, and no specific limitation is made here.
[0107] Since the embodiments of this application involve the application of neural networks, for the sake of easy understanding, some nouns or terms used in the embodiments of this application will be explained below, and these nouns or terms are also part of the invention content.
[0108] (1) Neural network
[0109] A neural network (neural network, NN) is a machine learning model. A neural network can be composed of neural units, and a neural unit can refer to an operation unit with x s and intercept 1 as inputs, and the output of this operation unit can be:
[0110]
[0111] where s = 1, 2,... n, n is a natural number greater than 1, and W s is xs where \(w\) is the weight of the neuron, \(b\) is the bias of the neuron, and \(f\) is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0112] (2) Deep Neural Network
[0113] A deep neural network (DNN), also known as a multi - layer neural network, can be understood as a neural network with many hidden layers, where "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the \(i\) - th layer must be connected to any neuron in the \((i + 1)\) - th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression:
[0114]
[0115] where, \(\mathbf{x}\) is the input vector, \(\mathbf{y}\) is the output vector, \(\mathbf{b}\) is the offset vector, \(W\) is the weight matrix (also known as the coefficient), and \(\alpha(\cdot)\) is the activation function. Each layer simply performs the following simple operation on the input vector \(\mathbf{x}\) to obtain the output vector Since the DNN has many layers, the number of the coefficient \(W\) and the offset vector \(\mathbf{b}\) is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient \(W\) as an example: Suppose in a three - layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient \(W\) is located, and the subscripts correspond to the index 2 of the output in the third layer and the index 4 of the input in the second layer. In summary: The coefficient from the \(k\) - th neuron in the \((L - 1)\) - th layer to the \(j\) - th neuron in the \(L\) - th layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0116] (3) Convolutional Neural Network
[0117] A convolutional neural network (CNN) is a deep neural network with a convolutional structure and a deep learning architecture. A deep learning architecture refers to performing multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, and each neuron in this feed-forward artificial neural network can respond to the input image. A convolutional neural network contains a feature extractor composed of convolutional layers and pooling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane (feature map).
[0118] A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. The convolutional layer can include many convolutional operators, which are also called kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolutional operation on the image, the weight matrix usually processes one pixel after another (or two pixels after two pixels... depending on the value of the stride) along the horizontal direction of the input image, thus completing the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image. During the convolutional operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices with the same size (rows × columns), that is, multiple matrices of the same type, are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, where the dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. The sizes (rows × columns) of these multiple weight matrices are the same, and the sizes of the feature maps extracted by these multiple weight matrices of the same size are also the same. Then, the multiple feature maps of the same size that are extracted are combined to form the output of the convolutional operation. The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can be used to extract information from the input image, so that the convolutional neural network can make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network increases, the features extracted by the subsequent convolutional layers become more and more complex, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0119] Since it is often necessary to reduce the number of training parameters, pooling layers are often introduced periodically after convolutional layers. It can be a pooling layer following a single convolutional layer, or one or more pooling layers following multiple convolutional layers. In the process of image processing, the sole purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of pixel values in the image within a specific range as the result of average pooling. The max pooling operator can select the pixel with the maximum value within the specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after passing through the pooling layer can be smaller than that of the image input to the pooling layer. Each pixel in the image output by the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0120] (4) Recurrent Neural Networks
[0121] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, it is from the input layer to the hidden layer and then to the output layer, with full connections between layers, while there are no connections between individual nodes within each layer. Although this ordinary neural network has solved many problems, it is still powerless for many issues. For example, when predicting the next word in a sentence, usually the previous words are needed because the words in a sentence are not independent. The reason why RNNs are called recurrent neural networks is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes within the hidden layer itself are no longer unconnected but connected, and the input to the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNNs can process sequential data of any length. The training of RNNs is the same as that of traditional CNNs or DNNs. The error backpropagation algorithm is also used, but there is one difference: that is, if the RNN is unfolded, the parameters such as W are shared; while in the traditional neural network mentioned above, this is not the case. And in using the gradient descent algorithm, the output at each step depends not only on the network at the current step but also on the states of the networks in several previous steps. This learning algorithm is called back propagation through time (BPTT).
[0122] Since we already have convolutional neural networks, why do we still need recurrent neural networks? The reason is simple. In convolutional neural networks, there is a premise assumption that elements are independent of each other, and the input and output are also independent, such as cats and dogs. However, in the real world, many elements are interconnected. For example, the change of stocks over time, or a person says: "I like traveling, and my favorite place is Yunnan. I must go there if I have the chance in the future." Here, fill in the blank, and humans should all know that the answer is "Yunnan". Because humans can make inferences based on the context. But how can we make machines do this? This is where RNN comes into being. RNN aims to enable machines to have the ability to remember like humans. Therefore, the output of RNN needs to depend on the current input information and historical memory information.
[0123] (5) Loss function
[0124] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, we need to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0125] (6) Backpropagation algorithm
[0126] Convolutional neural networks can use the error backpropagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial super-resolution model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0127] (7) Generative adversarial network
[0128] Generative adversarial networks (GAN) is a deep learning model. There are at least two modules in this model: one module is the Generative Model, and the other module is the Discriminative Model. Through the mutual game learning of these two modules, better outputs can be generated. Both the generative model and the discriminative model can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Taking the GAN for generating pictures as an example, assume there are two networks, G (Generator) and D (Discriminator). Among them, G is a network for generating pictures. It receives a random noise z and generates a picture through this noise, denoted as G(z); D is a discriminative network used to determine whether a picture is "real". Its input parameter is x, where x represents a picture, and the output D(x) represents the probability that x is a real picture. If it is 1, it means the picture is 100% real. If it is 0, it means the picture cannot be real. During the training process of this generative adversarial network, the goal of the generative network G is to generate as real pictures as possible to deceive the discriminative network D, while the goal of the discriminative network D is to distinguish the pictures generated by G from the real pictures as much as possible. In this way, G and D constitute a dynamic "game" process, that is, the "adversarial" in the "generative adversarial network". Finally, in the ideal state, the result of the game is that G can generate pictures G(z) that are "indistinguishable from the real ones", and D is difficult to determine whether the pictures generated by G are real, that is, D(G(z)) = 0.5. In this way, an excellent generative model G is obtained, which can be used to generate pictures.
[0129] The following will be combined with Figures 2a - 2e describe the video super-resolution network (also known as the neural network) in detail.
[0130] As Figure 2a shown, Input 1 is processed through a 3×3 convolutional layer (3×3Conv) and an activation layer (Relu). Input 2 is processed through another 3×3 convolutional layer and another activation layer. Then the results obtained after the above processing are merged (concat), and then passed through a block processing layer (Res-Block), …, block processing layer, 3×3 convolutional layer, activation layer, 3×3 convolutional layer to obtain a residual value. After adding Input 1 and the residual value, the output is obtained.
[0131] As Figure 2b shown, the above block processing layer may include a 3×3 convolutional layer, an activation layer, and a 3×3 convolutional layer. After processing the input through these three layers, the result obtained after the processing is added to the initial input to obtain the output.
[0132] As Figure 2c shown, the above block processing layer may include a 3×3 convolutional layer, an activation layer, a 3×3 convolutional layer, and an activation layer. After processing the input through the 3×3 convolutional layer, the activation layer, and the 3×3 convolutional layer, the processed result is added to the initial input, and finally, an output is obtained through an activation layer.
[0133] As Figure 2d shown, the input is processed through a 3×3 convolutional layer, an activation layer, a block processing layer,..., a block processing layer, a 3×3 convolutional layer, an activation layer, and a 3×3 convolutional layer to obtain an output.
[0134] As Figure 2e shown, input 1 is processed through a 3×3 convolutional layer and an activation layer. After multiplying input 1 and input 2, the result is processed through another 3×3 convolutional layer and another activation layer. Then, the processed results are merged (concat), and further processed through block processing layers,..., block processing layers, a 3×3 convolutional layer, an activation layer, and a 3×3 convolutional layer to obtain a residual value. The output is obtained by adding input 1 and the residual value.
[0135] It should be noted that the neural network shown as Figures 2a - 2e above is only several examples of neural networks. In specific applications, the neural network may also exist in the form of other network models, which are not specifically limited in this application. In addition, the input and output of the video super-resolution network depend on its training process, and the above input and output, as well as their quantities, do not constitute limitations. Input 1 and input 2 may respectively correspond to any image frame input to the video super-resolution network, and the output may correspond to the high-resolution image frame output by the video super-resolution network.
[0136] Figure 3 shows a schematic structural diagram of the terminal 300.
[0137] The terminal 300 may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 432, an antenna 1, an antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, a headphone jack 370D, a sensor module 380, a button 390, a motor 391, an indicator 392, a camera 393, a display screen 394, and a subscriber identification module (SIM) card interface 395, etc. The sensor module 380 may include a pressure sensor 380A, a gyroscope sensor 380B, a barometric pressure sensor 380C, a magnetic sensor 380D, an acceleration sensor 380E, a distance sensor 380F, a proximity light sensor 380G, a fingerprint sensor 380H, a temperature sensor 380J, a touch sensor 380K, an ambient light sensor 380L, a bone conduction sensor 380M, etc.
[0138] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the terminal 300. In some other embodiments of the present application, the terminal 300 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0139] The processor 310 may include one or more processing units. For example, the processor 310 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0140] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0141] A memory can also be provided in the processor 310 for storing instructions and data. In some embodiments, the memory in the processor 310 is a cache memory. This memory can hold the instructions or data that the processor 310 has just used or recycled. If the processor 310 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses and reduces the waiting time of the processor 310, thus improving the efficiency of the system.
[0142] In some embodiments, the processor 310 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0143] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 310 may include multiple groups of I2C buses. The processor 310 can be respectively coupled to the touch sensor 380K, the charger, the flashlight, the camera 393, etc. through different I2C bus interfaces. For example: The processor 310 can be coupled to the touch sensor 380K through the I2C interface, enabling the processor 310 and the touch sensor 380K to communicate through the I2C bus interface to implement the touch function of the terminal 300.
[0144] The I2S interface can be used for audio communication. In some embodiments, the processor 310 may include multiple groups of I2S buses. The processor 310 can be coupled to the audio module 370 through the I2S bus to achieve communication between the processor 310 and the audio module 370. In some embodiments, the audio module 370 can transmit an audio signal to the wireless communication module 360 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0145] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 370 and the wireless communication module 360 can be coupled through the PCM bus interface. In some embodiments, the audio module 370 can also transmit audio signals to the wireless communication module 360 through the PCM interface to implement the function of answering calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0146] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 310 and the wireless communication module 360. For example, the processor 310 communicates with the Bluetooth module in the wireless communication module 360 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 370 can transmit audio signals to the wireless communication module 360 through the UART interface to implement the function of playing music through a Bluetooth headset.
[0147] The MIPI interface can be used to connect the processor 310 to peripheral devices such as the display screen 394 and the camera 393. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 310 and the camera 393 communicate through the CSI interface to implement the shooting function of the terminal 300. The processor 310 and the display screen 394 communicate through the DSI interface to implement the display function of the terminal 300.
[0148] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 310 to the camera 393, the display screen 394, the wireless communication module 360, the audio module 370, the sensor module 380, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0149] The USB interface 330 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 330 can be used to connect a charger to charge the terminal 300, and can also be used to transfer data between the terminal 300 and peripheral devices. It can also be used to connect a headset to play audio through the headset. This interface can also be used to connect other terminals, such as AR devices, etc.
[0150] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the terminal 300. In other embodiments of the present application, the terminal 300 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0151] The charging management module 340 is configured to receive a charging input from a charger. Herein, the charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 340 may receive the charging input of the wired charger through the USB interface 330. In some embodiments of wireless charging, the charging management module 340 may receive the wireless charging input through the wireless charging coil of the terminal 300. While charging the battery 432, the charging management module 340 may also supply power to the terminal through the power management module 341.
[0152] The power management module 341 is used to connect the battery 432, the charging management module 340, and the processor 310. The power management module 341 receives the inputs from the battery 432 and / or the charging management module 340 and supplies power to the processor 310, the internal memory 321, the display screen 394, the camera 393, the wireless communication module 360, etc. The power management module 341 may also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 341 may also be disposed in the processor 310. In other embodiments, the power management module 341 and the charging management module 340 may also be disposed in the same device.
[0153] The wireless communication function of the terminal 300 may be implemented through the antenna 1, the antenna 2, the mobile communication module 350, the wireless communication module 360, the modulation and demodulation processor, and the baseband processor, etc.
[0154] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal 300 may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 may be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.
[0155] The mobile communication module 350 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal 300. The mobile communication module 350 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 350 may receive electromagnetic waves through the antenna 1, filter, amplify, and perform other processing on the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 350 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 350 may be provided in the processor 310. In some embodiments, at least some functional modules of the mobile communication module 350 and at least some modules of the processor 310 may be provided in the same device.
[0156] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 370A, receiver 370B, etc.), or displays an image or video through the display screen 394. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 310 and be provided in the same device as the mobile communication module 350 or other functional modules.
[0157] The wireless communication module 360 can provide wireless communication solutions applied to the terminal 300, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 360 can be one or more devices integrating at least one communication processing module. The wireless communication module 360 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 310. The wireless communication module 360 can also receive the signals to be sent from the processor 310, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0158] In some embodiments, antenna 1 of terminal 300 is coupled to mobile communication module 350, and antenna 2 is coupled to wireless communication module 360, enabling terminal 300 to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0159] Terminal 300 implements the display function through the GPU, display screen 394, and application processor, etc. The GPU is a microprocessor for image processing, connected to display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 310 may include one or more GPUs, which execute program instructions to generate or change display information.
[0160] The display screen 394 is used to display images, videos, etc. The display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal 300 may include one or N display screens 394, where N is a positive integer greater than 1.
[0161] The terminal 300 can implement the shooting function through an ISP, a camera 393, a video codec, a GPU, a display screen 394, an application processor, etc.
[0162] The ISP is used to process the data fed back by the camera 393. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 393.
[0163] The camera 393 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the terminal 300 may include one or N cameras 393, where N is a positive integer greater than 1.
[0164] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 300 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc. In this application, the terminal can obtain M image frames based on the super-resolution reference information, and input the M image frames and the first image frame to be processed into the video super-resolution network together to achieve the super-resolution processing of the first image frame. After obtaining a quantization parameter in the super-resolution reference information, the terminal can obtain the corresponding video super-resolution network according to the quantization parameter. Super-resolution is a process of obtaining a high-resolution image frame from multiple low-resolution image frames. The higher the resolution, the more pixel points are included per inch. Super-resolution processing is to increase the number of pixel points included per inch, making the details of the image frame rich and improving its clarity. With the help of the super-resolution reference information from the sending end, the terminal selects the video super-resolution network to be used and the reference image frames participating in the super-resolution processing to achieve the super-resolution processing of the first image frame (i.e., the image frame to be super-resolved), and improve the resolution of the first image frame. On the one hand, the terminal selects the corresponding video super-resolution network according to the quantization parameter of the first image frame, which can better improve the resolution. On the other hand, the reference image frames are selected by using the scores given by the sending end to each image frame to achieve the purpose of maximizing the quality after multi-frame fusion super-resolution, and improve the effect of the super-resolution processing of the image frame. On the third hand, using the resources of the sending end to score the image frames in the video stream can save the processing resources of the terminal, reduce the computing amount of the terminal, and improve the efficiency of its super-resolution processing.
[0165] The video codec is used to compress or decompress digital videos. The terminal 300 can support one or more video codecs. In this way, the terminal 300 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0166] The NPU is a neural-network (NN) computing processor. By referring to the biological neural network structure, such as referring to the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as the intelligent cognition of the terminal 300 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.
[0167] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 300. The external memory card communicates with the processor 310 through the external memory interface 320 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0168] The internal memory 321 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 321 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the terminal 300 (such as audio data, phone book, etc.). In addition, the internal memory 321 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 310 executes various functional applications and data processing of the terminal 300 by running the instructions stored in the internal memory 321 and / or the instructions stored in the memory provided in the processor.
[0169] The terminal 300 can implement audio functions through the audio module 370, the speaker 370A, the receiver 370B, the microphone 370C, the headphone jack 370D, and the application processor, etc. For example, music playback, recording, etc.
[0170] The audio module 370 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 370 can also be used to encode and decode audio signals. In some embodiments, the audio module 370 can be provided in the processor 310, or some functional modules of the audio module 370 can be provided in the processor 310.
[0171] The speaker 370A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal 300 can listen to music or hands-free calls through the speaker 370A.
[0172] The receiver 370B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal 300 answers a call or a voice message, the voice can be listened to by bringing the receiver 370B close to the human ear.
[0173] The microphone 370C, also known as a "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone 370C with their mouth to input the sound signal into the microphone 370C. The terminal 300 can be provided with at least one microphone 370C. In some other embodiments, the terminal 300 can be provided with two microphones 370C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the terminal 300 can also be provided with three, four or more microphones 370C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0174] The headphone jack 370D is used to connect a wired headphone. The headphone jack 370D can be a USB interface 330, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0175] The pressure sensor 380A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 380A can be disposed on the display screen 394. There are many types of pressure sensors 380A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 380A, the capacitance between the electrodes changes. The terminal 300 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 394, the terminal 300 detects the intensity of the touch operation according to the pressure sensor 380A. The terminal 300 can also calculate the position of the touch according to the detection signal of the pressure sensor 380A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0176] The gyroscope sensor 380B can be used to determine the motion posture of the terminal 300. In some embodiments, the angular velocity of the terminal 300 about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 380B. The gyroscope sensor 380B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 380B detects the angle of jitter of the terminal 300, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to offset the jitter of the terminal 300 through reverse movement to achieve anti-shake. The gyroscope sensor 380B can also be used for navigation and somatosensory game scenarios.
[0177] The barometric pressure sensor 380C is used to measure barometric pressure. In some embodiments, the terminal 300 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 380C to assist in positioning and navigation.
[0178] The magnetic sensor 380D includes a Hall sensor. The terminal 300 can use the magnetic sensor 380D to detect the opening and closing of the flip leather case. In some embodiments, when the terminal 300 is a flip phone, the terminal 300 can detect the opening and closing of the flip according to the magnetic sensor 380D. Furthermore, according to the detected opening and closing state of the leather case or the flip, features such as automatic flip unlocking can be set.
[0179] The acceleration sensor 380E can detect the magnitude of the acceleration of the terminal 300 in various directions (generally three axes). When the terminal 300 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the terminal and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0180] The distance sensor 380F is used to measure distance. The terminal 300 can measure distance through infrared or laser. In some embodiments, in a shooting scenario, the terminal 300 can use the distance sensor 380F to measure distance to achieve rapid focusing.
[0181] The proximity light sensor 380G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The terminal 300 emits infrared light outward through the light-emitting diode. The terminal 300 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 300. When insufficient reflected light is detected, the terminal 300 can determine that there is no object near the terminal 300. The terminal 300 can use the proximity light sensor 380G to detect when the user holds the terminal 300 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 380G can also be used for automatic unlocking and locking in the leather case mode and pocket mode.
[0182] The ambient light sensor 380L is used to sense the ambient light brightness. The terminal 300 can adaptively adjust the brightness of the display screen 394 according to the sensed ambient light brightness. The ambient light sensor 380L can also be used to automatically adjust the white balance during photography. The ambient light sensor 380L can also cooperate with the proximity light sensor 380G to detect whether the terminal 300 is in a pocket to prevent accidental touch.
[0183] The fingerprint sensor 380H is used to collect fingerprints. The terminal 300 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access to application locks, fingerprint photography, fingerprint answering of incoming calls, etc.
[0184] The temperature sensor 380J is used to detect temperature. In some embodiments, the terminal 300 uses the temperature detected by the temperature sensor 380J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 380J exceeds a threshold, the terminal 300 reduces the performance of the processor located near the temperature sensor 380J in order to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the terminal 300 heats the battery 432 to prevent abnormal shutdown of the terminal 300 caused by low temperature. In still other embodiments, when the temperature is lower than yet another threshold, the terminal 300 boosts the output voltage of the battery 432 to prevent abnormal shutdown caused by low temperature.
[0185] The touch sensor 380K, also known as a "touch control device". The touch sensor 380K can be disposed on the display screen 394, and the touch sensor 380K and the display screen 394 form a touch screen, also known as a "touch screen". The touch sensor 380K is used to detect touch operations acting on or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 394. In other embodiments, the touch sensor 380K can also be disposed on the surface of the terminal 300, at a different position from the display screen 394.
[0186] The bone conduction sensor 380M can acquire vibration signals. In some embodiments, the bone conduction sensor 380M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 380M can also contact the human pulse and receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 380M can also be disposed in the earphone to form a bone conduction earphone. The audio module 370 can parse out voice signals based on the vibration signals of the vibrating bone mass of the vocal part acquired by the bone conduction sensor 380M to implement voice functions. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 380M to implement heart rate detection functions.
[0187] The button 390 includes a power-on button, volume buttons, etc. The button 390 can be a mechanical button or a touch button. The terminal 300 can receive button inputs and generate key signal inputs related to the user settings and function controls of the terminal 300.
[0188] The motor 391 can generate vibration prompts. The motor 391 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations for different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations on different areas of the display screen 394, the motor 391 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminders, receiving messages, alarms, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0189] The indicator 392 can be an indicator light and can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0190] The SIM card interface 395 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 395 to achieve contact and separation from the terminal 300. The terminal 300 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 395 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 395 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 395 can also be compatible with different types of SIM cards. The SIM card interface 395 can also be compatible with external memory cards. The terminal 300 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the terminal 300 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal 300 and cannot be separated from the terminal 300.
[0191] Figure 4 This is a flowchart of an embodiment of the method for implementing super-resolution of an image frame in this application. As Figure 4 shown, the method of this embodiment can be applied to Figure 1 the architecture shown, and its execution subject can include a sending end and the terminal shown in Figure 2. The method for implementing super-resolution of the image frame can include:
[0192] Step 401, the terminal sends a video playback request to the sending end.
[0193] Based on the application scenario of the method provided in this application, when a user plays a video online on a terminal, after the user selects the video to be played, for example, clicks on the thumbnail or video name of the video, the terminal obtains the instruction triggered by this operation, and then the terminal sends a video play request to the sender of the video according to this instruction.
[0194] For example, the user uses a video application (APP) installed on the terminal to play a video. After opening the video APP, the user browses the video list of the video APP and clicks on the corresponding video to trigger its playback on the terminal. The user's operation triggers the video APP to generate a video play request, and the identification information of the video to be played, such as a uniform resource locator (URL), video index, etc., is carried in the video play request. The terminal sends the video play request to the sender that stores the video to be played through the communication network with the sender.
[0195] Step 402: The sender performs downsampling processing on the first image frame, where the first image frame is any one of the images in the video indicated by the video play request.
[0196] After receiving the video play request from the terminal, the sender extracts the corresponding video from the video library according to the identification information carried in the video play request. To reduce the bit rate of video transmission, the sender can perform downsampling processing on each image frame included in the video to reduce the image data. The downsampling methods can include nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, etc., and the downsampling scale can be one-half or one-fourth, etc. It should be noted that this application does not make specific limitations on the specific downsampling methods or parameters used by the sender.
[0197] Step 403: The sender encodes the first image frame after downsampling processing, and at the same time obtains the quantization parameter, image quality score, and reference frame indication information of the first image frame.
[0198] The encoding module of the sender encodes the first image frame after downsampling processing to obtain the bitstream of the first image frame. This encoding process involves mode selection, quantization processing, and reconstruction. During the mode selection process, the sender can determine the reference frame of the first image frame and thus obtain the reference frame indication information; during the quantization processing process, the sender can obtain the quantization parameter; during the reconstruction process, the sender can obtain the image quality score of the first image frame according to the original first image frame and the reconstructed frame of the first image frame.
[0199] The sender can perform quality scoring based on the first image frame and the reconstructed frame of the first image frame, so as to obtain the image quality score of the first image frame. The sender can use scoring tools such as PSNR, SSIM, VMAF, etc. to obtain the above-mentioned image quality score.
[0200] Step 404: The sender sends a video stream and super-resolution reference information to the terminal. The video stream includes the first image frame, and the super-resolution reference information includes quantization parameters, image quality scores, and reference frame indication information.
[0201] After the sender encodes the video stream, it obtains the bitstream of the video stream: In one case, when the sender encodes, it encodes the data of the image frames in the video stream and the super-resolution reference information, such as quantization parameters, image quality scores, and reference frame indication information, etc., together to obtain a bitstream containing the data of the image frames and the super-resolution reference information, and transmits this bitstream to the terminal; In another case, the sender encodes the data of the image frames in the video stream to obtain the bitstream of the image frames, then encodes the super-resolution reference information to obtain the bitstream of the super-resolution reference information, and transmits the bitstream of the image frames and the bitstream of the super-resolution reference information after splicing to the terminal. The present application does not make specific limitations on the implementation manner of the bitstream.
[0202] Step 405: The terminal plays the video.
[0203] After the terminal receives the bitstream from the sender, the decoding device in it decodes the bitstream to obtain the image frames and super-resolution reference information in the video stream, and plays them frame by frame on the video APP in the chronological order of the image frames. The bitstream of the image frames carries the timestamp information of the image frames, such as the time point of the image frame in the video stream, or the serial number of the image frame in the video stream, or the offset between the image frame and the first image frame in the video stream, etc.
[0204] Step 406: The terminal obtains the super-resolution reference information of the first image frame. The super-resolution reference information includes the quantization parameters and reference frame indication information of the first image frame, and the set of image quality scores corresponding to the first image frame.
[0205] During the video playback process, the user may be dissatisfied with the resolution of the currently played video and hope to see a clearer picture. At this time, the user can click on the control provided on the video APP (the implementation process can be seen in the following embodiments and will not be elaborated here), select a higher resolution or the highest resolution. This operation will trigger the generation of a resolution transformation instruction, which includes the resolution selected by the user, and the resolution selected by the user is higher than the current resolution. Based on the resolution transformation instruction, the terminal starts to perform super-resolution processing on the image frames in the video stream.
[0206] Optionally, during the process of playing a video by a video APP, it is detected that the received video stream can support online playback at a higher resolution. Therefore, the super-resolution processing of the image frames in the video stream is actively triggered.
[0207] The quantization parameter and reference frame indication information of the first image frame can be directly obtained from the bitstream of the first image frame, while the set of image quality scores corresponding to the first image frame can be obtained from the bitstreams of multiple image frames including the first image frame within a preset range.
[0208] In order to perform super-resolution processing on the first image frame, the video super-resolution network needs to input multiple frames of images as references. Therefore, the terminal can first obtain the set of image quality scores according to a preset range. The set of image quality scores can include the image quality scores of N consecutive image frames in the video stream in chronological order, and the N image frames include the first image frame. Assume that the serial number of the first image frame in the video sequence is n, and the serial number range of the N image frames is [1, N], where 1 ≤ n ≤ N. For example, when N = 3, the N image frames can include n - 1, n, n + 1; for another example, when N = 4, the N image frames can include n - 2, n - 1, n, n + 1; for another example, when N = 8, the N image frames can include n - 3, n - 2, n - 1, n, n + 1, n + 2, n + 3, n + 4. The value of N should not be too large. If the value of N is too large, the content span of the N image frames may be too large, resulting in a poor super-resolution processing effect on the first image frame. It should be noted that the first image frame can be located in the middle of the N image frames, or at the front or back end of the N image frames. How to select specifically can depend on the preset selection rule, and the present application does not make specific limitations on this.
[0209] Step 407: The terminal selects M image frames from the multiple image frames according to the set of image quality scores.
[0210] The value of M is related to the training model of the video super-resolution network. For example, when the training engine at the sending end trains the video super-resolution network, the input training data includes M image frames, then M image frames also need to be input when using this video super-resolution network, and M ≥ 1.
[0211] The terminal can select the M image frames with the highest image quality scores from N image frames. It can be seen that the M image frames selected by the terminal may or may not be the M consecutive image frames in chronological order. For example, when N = 3 and M = 2, the two selected image frames include n - 1 and n + 1, and these two image frames plus the first image frame are consecutive in sequence; for another example, when N = 4 and M = 3, the three selected image frames can include n - 2, n - 1, and n + 1, and these three image frames plus the first image frame are consecutive in sequence; for another example, when N = 8 and M = 4, the four selected image frames can include n - 3, n - 1, n + 2, and n + 4, and these four image frames are not consecutive even when the first image frame is added.
[0212] It should be noted that in order to ensure the effect of super-resolution processing of the first image frame and actually improve the resolution of the first image frame, in addition to the condition of selecting the M image frames with the highest image quality scores from N image frames as described above, it is also necessary to satisfy that the image quality scores of these M image frames are all higher than the image quality score of the first image frame. Therefore, there are the following three possibilities:
[0213] (1) When the N image frames include a first image frame set (the image quality scores of the image frames included in this first image frame set are all higher than the image quality score of the first image frame), and the number of image frames in the first image frame set is greater than or equal to M, select the M image frames with the highest image quality scores from the first image frame set.
[0214] This situation indicates that there are enough image frames in the N image frames whose image quality scores are higher than the image quality score of the first image frame. Therefore, M image frames can be directly selected from the first image frame set.
[0215] (2) When the N image frames include a first image frame set, but the number of image frames in the first image frame set is less than M, select all the image frames from the first image frame set, and the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame. When the number of all the image frames is less than M, select the image frames corresponding to the reference frame indication information from multiple image frames. When the number of all the image frames and the image frames corresponding to the reference frame indication information is less than M, select at least one image frame with the shortest to longest time interval from the first image frame from the other image frames except all the image frames and the image frames corresponding to the reference frame indication information in multiple image frames until M image frames are selected.
[0216] This situation indicates that among the N image frames, there are image frames but not enough whose image quality scores are higher than that of the first image frame. Therefore, in addition to all the image frames in the first image frame set, it is also necessary to select other image frames that meet other conditions until the number of selected image frames reaches M. The first preference for the other conditions is the reference frame of the first image frame. This reference frame is the reference frame used by the sender during video prediction and can be obtained through the reference frame indication information carried in the bitstream. It should be noted that the reference frame of the first image frame may already be included in the first image frame set, so the second preference condition among the other conditions needs to be considered. If the number of image frames still does not reach M after selecting the reference frame of the first image frame, the second preference condition among the other conditions can be considered until the number of selected image frames reaches M. The second preference condition can be at least one image frame with an interval time from short to long with the first image frame among the other image frames in the N image frames except all the image frames in the first image frame set and the image frames corresponding to the reference frame indication information.
[0217] (3) The N image frames do not include the first image frame set. Determine whether the first image frame is an I frame. If the first image frame is an I frame, then the M image frames refer to M copy samples of the first image frame; if the first image frame is not an I frame, select the image frame corresponding to the reference frame indication information from the multiple image frames. When the number of image frames corresponding to the reference frame indication information is less than M, select at least one image frame with an interval time from short to long with the first image frame from the other image frames in the multiple image frames except the image frames corresponding to the reference frame indication information until M image frames are selected.
[0218] This situation indicates that there are no image frames in the N image frames whose image quality scores are higher than that of the first image frame. The first image frame may have two situations. One is that the first image frame is an I frame, and a scene switch may occur starting from the first image frame. Then the content of the previous scene is different from that of the switched scene, and the information of the image frames before the first image frame cannot be utilized; or, the first image frame is an I frame inserted at a fixed interval, then the quality of the first image frame is usually much higher than that of the surrounding image frames, and the information of the surrounding image frames cannot be utilized either. Therefore, the selected M image frames are M copy samples of the first image frame. The other is that the first image frame is not an I frame (B frame or P frame). Based on the image quality scores, there are no selectable image frames among the N image frames. Therefore, first select the reference frame of the first image frame, and then select at least one image frame with an interval time from short to long with the first image frame from the other image frames in the N image frames except all the image frames in the first image frame set and the image frames corresponding to the reference frame indication information until the number of selected image frames reaches M.
[0219] Step 408: The terminal obtains a video super-resolution network corresponding to the quantization parameter, and this video super-resolution network has a super-resolution function.
[0220] The video super-resolution network is a neural network with a super-resolution function obtained through training. For the training process of the video super-resolution network, refer to the relevant description of the training engine in the sender, which will not be elaborated here. As described above, for the correspondence between the video super-resolution network and the QP, after the terminal obtains the QP of the first image frame from the code stream, it can obtain the corresponding video super-resolution network according to this QP and use it as the neural network for super-resolution processing of the first image frame.
[0221] Step 409: The terminal inputs M image frames and the first image frame into the video super-resolution network. The video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
[0222] The terminal inputs the M image frames obtained in Step 407 into the video super-resolution network. After being processed by the video super-resolution network, a second image frame is obtained, and the resolution of this second image frame is the resolution indicated by the resolution transformation instruction in Step 406.
[0223] The terminal in this application selects the video super-resolution network to be used and the reference image frames participating in the super-resolution processing with the help of the super-resolution reference information from the sender to achieve the super-resolution processing of the first image frame (i.e., the image frame to be super-resolved) and improve the resolution of the first image frame. On the one hand, the terminal selects the corresponding video super-resolution network according to the quantization parameter of the first image frame, which can better improve the resolution. On the other hand, the reference image frames are selected by using the scores given by the sender to each image frame to achieve the purpose of maximizing the quality after multi-frame fusion super-resolution and improve the effect of the super-resolution processing of the image frame. On the third hand, using the resources of the sender to score the image frames in the video stream can save the processing resources of the terminal, reduce the computational amount of the terminal, and improve the efficiency of its super-resolution processing.
[0224] Figures 5a - 5c Shows an exemplary schematic diagram of the process of playing a video online.
[0225] As Figure 5a shown, the user clicks on the icon of the video APP on the desktop of the terminal to open the video APP.
[0226] As Figure 5b shown, the user selects the video to be played in the video list. The video list includes the thumbnails and names of selected videos, as well as the thumbnails and names of popular videos, and the titles of video classifications, etc.
[0227] AsFigure 5c As shown, enter the playback interface of the video selected by the user, and the video starts playing full screen on the screen of the terminal.
[0228] Figures 6a - 6c An exemplary schematic diagram showing the process of switching video resolution is shown.
[0229] As Figure 6a shown, there is a control for selecting the resolution in the lower right corner of the video playback interface. At this time, the currently adopted resolution is displayed on the control, for example, 270P. Clicking on this control brings up a pull-down menu. The 270P has an underline indicating that this is the resolution currently used for playing the video. As long as this resolution is not the highest definition resolution, the user can select a higher definition resolution than the current one.
[0230] As Figure 6b shown, the user selects a resolution from those with a higher definition than the current resolution. At this time, the terminal receives a resolution transformation instruction generated based on this operation, and then adopts the method in the embodiment shown in Figure 4 to adjust the resolution of the video being played, specifically, to perform super-resolution processing on the image frames in the video.
[0231] As Figure 6c shown, after the super-resolution processing, the resolution of the video is switched to the resolution selected by the user, and the words "Resolution switched from 270P to 1080P" are displayed on the playback interface. At this time, the currently adopted resolution is displayed on the control, for example, 1080P. Moreover, the user can clearly see that the video being played is much clearer than before the resolution transformation.
[0232] It should be noted that Figures 5a - 5c and Figures 6a - 6c are both examples provided by this application and do not constitute any limitation. The video playback process and the process of switching video resolution, including the display interface, the implementation method of the control, the representation method of the resolution (for example, standard definition, high definition, ultra high definition, etc.), the resolution switching method, etc., can also adopt other implementation methods, and this application does not make specific limitations in this regard.
[0233] Figure 7 is a schematic structural diagram of an embodiment of the terminal device of this application. As Figure 7 shown, this device can be applied to the terminal shown in Figure 2. The terminal device of this embodiment may include: an acquisition module 701 and a processing module 702. Among them,
[0234] An acquisition module 701, configured to acquire super-resolution reference information, where the super-resolution reference information includes quantization parameters and a set of image quality scores, and the set of image quality scores includes the image quality scores of multiple image frames respectively; select M image frames from the multiple image frames according to the set of image quality scores, where M is greater than or equal to 1; acquire a video super-resolution network corresponding to the quantization parameters, and the video super-resolution network has a super-resolution function; a processing module 702, configured to input the M image frames and a first image frame into the video super-resolution network, and the video super-resolution network is configured to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
[0235] In a possible implementation manner, the acquisition module 701 is specifically configured to, when the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is greater than or equal to M, select the M image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame.
[0236] In a possible implementation manner, the M image frames include the first M image frames in the first set of image frames arranged in descending order of image quality scores.
[0237] In a possible implementation manner, the super-resolution reference information further includes reference frame indication information of the first image frame; the acquisition module 701 is further configured to, when the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is less than M, select all the image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; when the number of all the image frames is less than M, select the image frames corresponding to the reference frame indication information from the multiple image frames; when the number of all the image frames and the image frames corresponding to the reference frame indication information is less than M, select at least one image frame with an interval time from short to long with the first image frame from the other image frames in the multiple image frames except all the image frames and the image frames corresponding to the reference frame indication information until M image frames are selected.
[0238] In a possible implementation, the super-resolution reference information further includes reference frame indication information of the first image frame; the obtaining module 701 is further configured to, when the multiple image frames do not include the first image frame set, determine whether the first image frame is an I-frame, and the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame; if the first image frame is an I-frame, the M image frames include M copy samples of the first image frame; if the first image frame is not an I-frame, select the image frame corresponding to the reference frame indication information from the multiple image frames; when the number of image frames corresponding to the reference frame indication information is less than M, select at least one image frame with an increasing interval time from the first image frame from the other image frames except the image frames corresponding to the reference frame indication information in the multiple image frames until M image frames are selected.
[0239] In a possible implementation, the obtaining module 701 is specifically configured to receive the super-resolution reference information sent by the sending end.
[0240] In a possible implementation, the quantization parameter includes the quantization parameter used by the sending end during the quantization process of the first image frame.
[0241] In a possible implementation, the multiple image frames include consecutive multiple image frames in the video stream, and the multiple image frames include the first image frame.
[0242] In a possible implementation, the video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
[0243] In a possible implementation, the video super-resolution network includes a convolutional layer and an activation layer.
[0244] In a possible implementation, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0245] The device in this embodiment can be used to execute Figure 4 the technical solution of the method embodiment shown, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0246] Figure 8 This is a schematic structural diagram of the sending end device embodiment of the present application. As Figure 8 shown, this device can be applied to the sending end. The sending end device in this embodiment may include: an obtaining module 801, a sending module 802, and a training module 803. Among them,
[0247] An acquisition module 801, configured to acquire quantization parameters of a first image frame; obtain an image quality score of the first image frame according to the first image frame and a reconstructed frame of the first image frame; a sending module 802, configured to send a video stream and super-resolution reference information to a terminal device, where the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameters and the image quality score.
[0248] In a possible implementation manner, the acquisition module 801 is further configured to acquire reference frame indication information of the first image frame; correspondingly, the super-resolution reference information further includes the reference frame indication information.
[0249] In a possible implementation manner, the acquisition module 801 is specifically configured to obtain the image quality score of the first image frame according to a peak signal-to-noise ratio (PSNR), a structural similarity (SSIM), or a video multi-feature fusion evaluation criterion (VMAF).
[0250] In a possible implementation manner, the video super-resolution network includes a convolutional neural network (CNN), a deep neural network (DNN), or a recurrent neural network (RNN).
[0251] In a possible implementation manner, the video super-resolution network includes a convolutional layer and an activation layer.
[0252] In a possible implementation manner, the depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolution kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
[0253] In a possible implementation manner, it further includes: a training module 803; the acquisition module 801 is further configured to acquire a training data set, where the training data set includes a first-resolution image and a second-resolution image of each of multiple image frames, and multiple quantization parameters, and the resolution of the first-resolution image is higher than the resolution of the second-resolution image; the training module 803 is configured to train multiple video super-resolution networks according to the training data set, and the multiple video super-resolution networks correspond to the multiple quantization parameters.
[0254] The device in this embodiment can be used to execute Figure 4 the technical solution of the method embodiment shown, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0255] In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in this application can be directly implemented by the hardware encoding processor, or implemented by the combination of the hardware and software modules in the encoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0256] The memories mentioned in the above embodiments may be volatile memories or non-volatile memories, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include but are not limited to these and any other suitable types of memories.
[0257] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0258] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0259] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0260] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0261] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0262] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0263] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A method for implementing super-resolution of image frames, characterized in that, Including: Obtain super-resolution reference information of a first image frame, where the super-resolution reference information includes quantization parameters and a set of image quality scores, and the set of image quality scores includes the image quality scores of multiple image frames respectively; Select M image frames from the multiple image frames according to the set of image quality scores, where M is greater than or equal to 1; Obtain a video super-resolution network corresponding to the quantization parameters, and the video super-resolution network has a super-resolution function; Input the M image frames and the first image frame into the video super-resolution network, and the video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
2. The method according to claim 1, wherein The step of selecting M image frames from the multiple image frames according to the set of image quality scores specifically includes: When the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is greater than or equal to M, select the M image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame.
3. The method according to claim 2, wherein The M image frames include the first M image frames in the first set of image frames arranged in descending order of image quality score.
4. The method according to any one of claims 1-3, characterized in that, The super-resolution reference information further includes reference frame indication information of the first image frame; The step of selecting M image frames from the multiple image frames according to the set of image quality scores further includes: When the multiple image frames include a first set of image frames and the number of image frames in the first set of image frames is less than M, select all the image frames from the first set of image frames, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; When the number of all the image frames is less than M, select the image frames corresponding to the reference frame indication information from the multiple image frames; When the number of all the image frames and the image frames corresponding to the reference frame indication information is less than M, select at least one image frame with the shortest to longest time interval from the first image frame from the other image frames in the multiple image frames except the all the image frames and the image frames corresponding to the reference frame indication information until M image frames are selected.
5. The method according to any one of claims 1-3, characterized in that The super-resolution reference information further includes reference frame indication information of the first image frame; The step of selecting M image frames from the multiple image frames according to the set of image quality scores further includes: When the multiple image frames do not include a first set of image frames, determine whether the first image frame is an I frame, and the image quality scores of the image frames included in the first set of image frames are all higher than the image quality score of the first image frame; If the first image frame is an I frame, the M image frames include M replicated samples of the first image frame; If the first image frame is not an I frame, select the image frames corresponding to the reference frame indication information from the multiple image frames; When the number of image frames corresponding to the reference frame indication information is less than M, at least one image frame with an increasing time interval from the first image frame is selected from the other image frames in the multiple image frames except the image frames corresponding to the reference frame indication information until M image frames are selected.
6. The method according to any one of claims 1 to 3, characterized in that The obtaining of the super-resolution reference information specifically includes: Receiving the super-resolution reference information sent by the sending end.
7. The method according to claim 6, wherein The quantization parameter includes the quantization parameter used by the sending end during the quantization process of the first image frame.
8. The method according to any one of claims 1-3 and 7, characterized in that, The multiple image frames include consecutive multiple image frames in the video stream, and the multiple image frames include the first image frame.
9. The method according to any one of claims 1-3, 7, characterized in that The video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
10. The method according to claim 9, wherein The video super-resolution network includes a convolutional layer and an activation layer.
11. The method according to claim 10, wherein The depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
12. A method for implementing super-resolution of an image frame, characterized in that, including: Obtaining the quantization parameter of the first image frame, and for the quantization parameter, obtaining a video super-resolution network corresponding to the quantization parameter, where the video super-resolution network has a super-resolution function; Obtaining the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame; Sending a video stream and super-resolution reference information to the terminal, where the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameter and the image quality score.
13. The method according to claim 12, wherein Before sending the video stream and super-resolution reference information to the terminal, it further includes: Obtaining the reference frame indication information of the first image frame; Correspondingly, the super-resolution reference information further includes the reference frame indication information.
14. The method according to claim 12 or 13, characterized in that, The obtaining of the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame specifically includes: Obtaining the image quality score of the first image frame according to the peak signal-to-noise ratio PSNR, the structural similarity SSIM, or the video evaluation criterion VMAF of multi-feature fusion.
15. The method according to claim 12 or 13, characterized in that, The video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
16. The method according to claim 15, wherein The video super-resolution network includes a convolutional layer and an activation layer.
17. The method according to claim 16, wherein The depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
18. The method according to any one of claims 16-17, characterized in that The method further includes: Obtaining a training data set, where the training data set includes the first-resolution images and second-resolution images of multiple image frames respectively, and multiple quantization parameters, and the resolution of the first-resolution images is higher than the resolution of the second-resolution images; Training multiple video super-resolution networks according to the training data set, and the multiple video super-resolution networks correspond to the multiple quantization parameters.
19. A method for implementing super-resolution of an image frame, characterized in that, including: The sender obtains the quantization parameters of the first image frame, and obtains the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame; The sender sends a video stream and super-resolution reference information to the terminal, the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameters and the image quality score; The terminal obtains the first image frame according to the video stream, and obtains the quantization parameters and the set of image quality scores according to the super-resolution reference information, the set of image quality scores includes the image quality scores of multiple image frames respectively; The terminal selects M image frames from the multiple image frames according to the set of image quality scores, where M is greater than or equal to 1; The terminal obtains a video super-resolution network corresponding to the quantization parameters, and the video super-resolution network has a super-resolution function; The terminal inputs the M image frames and the first image frame into the video super-resolution network, and the video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
20. A terminal device, characterized in that, Including: An obtaining module, configured to obtain super-resolution reference information of a first image frame, the super-resolution reference information includes quantization parameters and a set of image quality scores, the set of image quality scores includes the image quality scores of multiple image frames respectively; select M image frames from the multiple image frames according to the set of image quality scores, where M is greater than or equal to 1; obtain a video super-resolution network corresponding to the quantization parameters, and the video super-resolution network has a super-resolution function; A processing module, configured to input the M image frames and the first image frame into the video super-resolution network, and the video super-resolution network is used to perform super-resolution processing on the first image frame according to the M image frames to obtain a second image frame, and the resolution of the second image frame is higher than that of the first image frame.
21. The device according to claim 20, wherein, The obtaining module is specifically configured to, when the multiple image frames include a first image frame set and the number of image frames in the first image frame set is greater than or equal to M, select the M image frames from the first image frame set, and the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame.
22. The device according to claim 21, characterized in that, The M image frames include the first M image frames arranged in descending order of image quality score in the first image frame set.
23. The device according to any one of claims 20-22, characterized in that, The super-resolution reference information further includes reference frame indication information of the first image frame; the obtaining module is further configured to, when the multiple image frames include a first image frame set and the number of image frames in the first image frame set is less than M, select all the image frames from the first image frame set, and the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame; when the number of all the image frames is less than M, select the image frames corresponding to the reference frame indication information from the multiple image frames; When the number of all the image frames and the image frames corresponding to the reference frame indication information is less than M, at least one image frame with an interval time from short to long with respect to the first image frame is selected from the other image frames in the multiple image frames except the all image frames and the image frames corresponding to the reference frame indication information until M image frames are selected.
24. The device according to any one of claims 20-22, characterized in that, The super-resolution reference information further includes the reference frame indication information of the first image frame; the obtaining module is further configured to, when the multiple image frames do not include the first image frame set, determine whether the first image frame is an I frame, and the image quality scores of the image frames included in the first image frame set are all higher than the image quality score of the first image frame; if the first image frame is an I frame, the M image frames include M replication samples of the first image frame; if the first image frame is not an I frame, the image frame corresponding to the reference frame indication information is selected from the multiple image frames; When the number of the image frames corresponding to the reference frame indication information is less than M, at least one image frame with an interval time from short to long with respect to the first image frame is selected from the other image frames in the multiple image frames except the image frames corresponding to the reference frame indication information until M image frames are selected.
25. The device according to any one of claims 20 - 22, characterized in that, The obtaining module is specifically configured to receive the super-resolution reference information sent by the sending end.
26. The device according to claim 25, characterized in that, The quantization parameter includes the quantization parameter used by the sending end during the quantization process of the first image frame.
27. The device according to any one of claims 20-22, 26, characterized in that, The multiple image frames include a continuous plurality of image frames in the video stream, and the multiple image frames include the first image frame.
28. The device according to any one of claims 20 - 22, 26, characterized in that The video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
29. The device according to claim 28, wherein, The video super-resolution network includes a convolutional layer and an activation layer.
30. The device according to claim 29, characterized in that, The depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolution kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
31. A transmitting end device, characterized in that, Comprising: An obtaining module, configured to obtain the quantization parameter of the first image frame, and for the quantization parameter, obtain a video super-resolution network corresponding to the quantization parameter, the video super-resolution network having a super-resolution function; obtain the image quality score of the first image frame according to the first image frame and the reconstructed frame of the first image frame; A sending module, configured to send a video stream and super-resolution reference information to a terminal device, the video stream includes the first image frame, and the super-resolution reference information includes the quantization parameter and the image quality score.
32. The device according to claim 31, characterized in that, The obtaining module is further configured to obtain the reference frame indication information of the first image frame; correspondingly, the super-resolution reference information further includes the reference frame indication information.
33. The device according to claim 31 or 32, characterized in that, The obtaining module is specifically configured to obtain the image quality score of the first image frame according to the peak signal-to-noise ratio PSNR, the structural similarity SSIM, or the video evaluation criterion VMAF of multi-feature fusion.
34. The device according to claim 31 or 32, characterized in that, The video super-resolution network includes a convolutional neural network CNN, a deep neural network DNN, or a recurrent neural network RNN.
35. The device according to claim 34, characterized in that, The video super-resolution network includes a convolutional layer and an activation layer.
36. The device according to claim 35, characterized in that, The depth of the convolutional layer is 2, 3, 4, 5, 6, 16, 24, 32, 48, 64, or 128; the size of the convolutional kernel in the convolutional layer is 1×1, 3×3, 5×5, or 7×7.
37. The device according to any one of claims 35-36, characterized in that It further includes: A training module; The obtaining module is further configured to obtain a training data set, where the training data set includes the first-resolution images and the second-resolution images of multiple image frames respectively, and multiple quantization parameters, and the resolution of the first-resolution images is higher than that of the second-resolution images. The training module is configured to train multiple video super-resolution networks according to the training data set, and the multiple video super-resolution networks correspond to the multiple quantization parameters.
38. An image processing system, characterized in that, It includes: A sending-end device and a terminal device, where the sending-end device adopts the device described in any one of claims 31-37, and the terminal device adopts the device described in any one of claims 20-30.
39. A terminal, characterized in that, It includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1-11, or implement the method executed by the terminal in claim 19.
40. A transmitting end, characterized in that, It includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 12-18, or implement the method executed by the sending end in claim 19.
41. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program is executed on a computer, the computer executes the method described in any one of claims 1-19.
Citation Information
Patent Citations
Method for enhancing video service quality in wireless ad hoc network environment
CN110099280A
Video super-resolution method based on coding damage repair
CN110751597A
Image recovery method and coding end
CN110875906A